The integration of multiple modalities, such as visual and textual, in Artificial Intelligence (AI) models improves efficiency by exploiting the complementary information in the data. This leads to significant advancements in high-impact domains including healthcare, education and mobility. However, the opacity of modern black box models raises trust and accountability concerns, limiting their adoption. Explainable AI (XAI) mitigates these issues by helping users and domain experts understand and audit the outputs produced by multimodal representations. In this systematic survey, we query major editorial databases, collecting and analyzing 155 publications and 54 datasets spanning 2018–2025, to provide a structured overview of the emerging field of multimodal explainability. Unlike other surveys, we organize the current literature by non-mutually-exclusive, multi-label dimensions comprising multimodal fusion strategies, XAI scopes and techniques, data modalities, model types and application domains. We derive novel conceptualizations specific to multimodal explainability, introducing the distinction between intrinsically and extrinsically multimodal processes. Our quantitative synthesis highlights a field dominated by healthcare applications (48%), with widespread use of intrinsic processes (87.7%), visual or quantitative-based techniques (88.4%), and Neural Network black boxes, while extrinsic processes, white box models, probabilistic or concept-based techniques remain under-explored. Moreover, the public availability of code and human evaluation remains limited. These findings identify critical open gaps and outline directions for developing reproducible and trustworthy multimodal intelligent systems.
A survey on multimodal explainable Artificial Intelligence
Metta C.;Monreale A.;Rinzivillo S.
2026
Abstract
The integration of multiple modalities, such as visual and textual, in Artificial Intelligence (AI) models improves efficiency by exploiting the complementary information in the data. This leads to significant advancements in high-impact domains including healthcare, education and mobility. However, the opacity of modern black box models raises trust and accountability concerns, limiting their adoption. Explainable AI (XAI) mitigates these issues by helping users and domain experts understand and audit the outputs produced by multimodal representations. In this systematic survey, we query major editorial databases, collecting and analyzing 155 publications and 54 datasets spanning 2018–2025, to provide a structured overview of the emerging field of multimodal explainability. Unlike other surveys, we organize the current literature by non-mutually-exclusive, multi-label dimensions comprising multimodal fusion strategies, XAI scopes and techniques, data modalities, model types and application domains. We derive novel conceptualizations specific to multimodal explainability, introducing the distinction between intrinsically and extrinsically multimodal processes. Our quantitative synthesis highlights a field dominated by healthcare applications (48%), with widespread use of intrinsic processes (87.7%), visual or quantitative-based techniques (88.4%), and Neural Network black boxes, while extrinsic processes, white box models, probabilistic or concept-based techniques remain under-explored. Moreover, the public availability of code and human evaluation remains limited. These findings identify critical open gaps and outline directions for developing reproducible and trustworthy multimodal intelligent systems.| File | Dimensione | Formato | |
|---|---|---|---|
|
Giovannoni et al_Intelligent Systems with Applications 2026.pdf
accesso aperto
Descrizione: A survey on multimodal explainable Artificial Intelligence
Tipologia:
Versione Editoriale (PDF)
Licenza:
Creative commons
Dimensione
2.41 MB
Formato
Adobe PDF
|
2.41 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


