Open-Vocabulary Vision–Language Segmentation for Strawberry Harvesting Scene Understanding


📅 2025 🔖 RA Projects

This RA work is part of a strawberry harvesting efficiency evaluation study led by Dr. Leonardo Guevara at the Lincoln Institute for Agri-Food Technology, University of Lincoln. My primary contribution was to test and compare five SAM-centric and CLIP-centric open-vocabulary vision–language segmentation methods for strawberry harvesting scene understanding. These methods include SAM, FC-CLIP, CLIP-SAM, Grounded-SAM, and Semantic Segment Anything (SSA). The primary target objects for recognition include ripe strawberries, unripe strawberries, foliage, trolleys, punnets, and human hands.

Strawberry Picking

Conclusions

SAM achieves fine-grained segmentation but lacks semantic understanding. FC-CLIP is relatively effective at identifying foliage, though it occasionally misclassifies strawberries as foliage and struggles with identifying punnets. SSA offers more detailed recognition, yet suffers from significant confusion among classes. CLIP-SAM demonstrates superior overall recognition performance, though ripe and unripe strawberries, as well as foliage and unripe strawberries, are prone to misclassification. Grounded-SAM exhibits the most balanced performance: despite performing poorly on distant strawberries and occasional failures in identifying punnets and trolleys, it accurately identifies and segments the majority of target objects, making it the most effective method. The following presents the results obtained using Grounded-SAM.

strawberry_picking_src
strawberry_picking_result