Object detection and Instance Segmentation are critical tasks for computer vision applications in precision agriculture. Though real-time object detectors have experienced significant development over the last few years, thanks to the introduction of the You Only Look Once (YOLO) family, these models often show weaker performance in more specific contexts, even when fine-tuned on a proper dataset. This work proposes an architectural modification of YOLOv11 aiming at increasing the localization accuracy in complex scenes which are typical of the agricultural field, often characterized by heterogeneous lighting, nonuniform terrain morphology, and high crop density. We compare a variety of configurations that introduce Simple Attention Module (SimAM) blocks and the Minimum Points Distance Intersection-over-Union (MPDIoU) cost function against the unaltered YOLOv11 model. To test how the models behave in real-world scenarios, we evaluate these variants (across all existing sizes of the original model) on a dataset of annotated images of lettuce crops in soil. Detection and segmentation results show that medium-sized models benefit the most from the proposed architectural changes, with strong potential for realtime deployment with little to no sacrifice in localization accuracy compared to larger variants.
Enhancing YOLOv11 with SimAM and MPDIoU for Lettuce Detection
Michele Paradiso;Angelo Cardellicchio;Vito Renò;Annalisa Milella
2026
Abstract
Object detection and Instance Segmentation are critical tasks for computer vision applications in precision agriculture. Though real-time object detectors have experienced significant development over the last few years, thanks to the introduction of the You Only Look Once (YOLO) family, these models often show weaker performance in more specific contexts, even when fine-tuned on a proper dataset. This work proposes an architectural modification of YOLOv11 aiming at increasing the localization accuracy in complex scenes which are typical of the agricultural field, often characterized by heterogeneous lighting, nonuniform terrain morphology, and high crop density. We compare a variety of configurations that introduce Simple Attention Module (SimAM) blocks and the Minimum Points Distance Intersection-over-Union (MPDIoU) cost function against the unaltered YOLOv11 model. To test how the models behave in real-world scenarios, we evaluate these variants (across all existing sizes of the original model) on a dataset of annotated images of lettuce crops in soil. Detection and segmentation results show that medium-sized models benefit the most from the proposed architectural changes, with strong potential for realtime deployment with little to no sacrifice in localization accuracy compared to larger variants.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


