Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions A deep-learning perception framework using the YOLOv12-m-seg model achieved a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image, the most favorable balance of accuracy and speed among direct segmentation models evaluated for selective robotic cotton harvesting, according to an arXiv paper (2609.19592v1). The study trained on 1,008 annotated field images captured by three cameras under varying natural lighting and weather, and among detection models GELAN-s posted an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and F1-score of 79.0% at 42.3 ms per image. Field tests with a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera validated YOLOv12-m-seg for real-time cotton boll detection, segmentation, and selective picking, with the model reaching an R² of 0.966 against manually annotated masks versus 0.860 for GELAN-s + SAMv2.1 Tiny. arXiv:2609.19592v1 Announce Type: new Abstract: This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model SAM , SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model RAM . Among the detection models, GELAN-s achieved the most favorable balance between mean average precision mAP and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.