Open-Vocabulary Vision–Language Segmentation for Strawberry Harvesting Scene Understanding
📅 2025 🔖 RA Projects
This RA work is part of a strawberry harvesting efficiency evaluation study led by Dr. Leonardo Guevara at the Lincoln Institute for Agri-Food Technology, University of Lincoln. My primary contribution was to test and compare five SAM-centric and CLIP-centric open-vocabulary vision–language segmentation methods for strawberry harvesting scene understanding. These methods include SAM, FC-CLIP, CLIP-SAM, Grounded-SAM, and Semantic Segment Anything (SSA). The primary target objects for recognition include ripe strawberries, unripe strawberries, foliage, trolleys, punnets, and human hands.

Conclusions
SAM achieves fine-grained segmentation but lacks semantic understanding. FC-CLIP is relatively effective at identifying foliage, though it occasionally misclassifies strawberries as foliage and struggles with identifying punnets. SSA offers more detailed recognition, yet suffers from significant confusion among classes. CLIP-SAM demonstrates superior overall recognition performance, though ripe and unripe strawberries, as well as foliage and unripe strawberries, are prone to misclassification. Grounded-SAM exhibits the most balanced performance: despite performing poorly on distant strawberries and occasional failures in identifying punnets and trolleys, it accurately identifies and segments the majority of target objects, making it the most effective method. The following presents the results obtained using Grounded-SAM.


