Object detection and Instance Segmentation are critical tasks for computer vision applications in precision agriculture. Though real-time object detectors have experienced significant development over the last few years, thanks to the introduction of the You Only Look Once (YOLO) family, these models often show weaker performance in more specific contexts, even when fine-tuned on a proper dataset. This work proposes an architectural modification of YOLOv11 aiming at increasing the localization accuracy in complex scenes which are typical of the agricultural field, often characterized by heterogeneous lighting, nonuniform terrain morphology, and high crop density. We compare a variety of configurations that introduce Simple Attention Module (SimAM) blocks and the Minimum Points Distance Intersection-over-Union (MPDIoU) cost function against the unaltered YOLOv11 model. To test how the models behave in real-world scenarios, we evaluate these variants (across all existing sizes of the original model) on a dataset of annotated images of lettuce crops in soil. Detection and segmentation results show that medium-sized models benefit the most from the proposed architectural changes, with strong potential for realtime deployment with little to no sacrifice in localization accuracy compared to larger variants.

Enhancing YOLOv11 with SimAM and MPDIoU for Lettuce Detection

Michele Paradiso;Angelo Cardellicchio;Vito Renò;Annalisa Milella
2026

Abstract

Object detection and Instance Segmentation are critical tasks for computer vision applications in precision agriculture. Though real-time object detectors have experienced significant development over the last few years, thanks to the introduction of the You Only Look Once (YOLO) family, these models often show weaker performance in more specific contexts, even when fine-tuned on a proper dataset. This work proposes an architectural modification of YOLOv11 aiming at increasing the localization accuracy in complex scenes which are typical of the agricultural field, often characterized by heterogeneous lighting, nonuniform terrain morphology, and high crop density. We compare a variety of configurations that introduce Simple Attention Module (SimAM) blocks and the Minimum Points Distance Intersection-over-Union (MPDIoU) cost function against the unaltered YOLOv11 model. To test how the models behave in real-world scenarios, we evaluate these variants (across all existing sizes of the original model) on a dataset of annotated images of lettuce crops in soil. Detection and segmentation results show that medium-sized models benefit the most from the proposed architectural changes, with strong potential for realtime deployment with little to no sacrifice in localization accuracy compared to larger variants.
2026
Istituto di Sistemi e Tecnologie Industriali Intelligenti per il Manifatturiero Avanzato - STIIMA (ex ITIA) Sede Secondaria Bari
Precision Agriculture, Deep Learning, Computer Vision, Object Detection
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/592002
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ente

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact