This research focuses on the design and development of an automatic, AI-based, and open-science–oriented workflow that leverages Semantic Web technologies and a knowledge representation approach to transform unstructured textual data into a knowledge graph of formal narratives. The resulting knowledge graph is modelled according to the Narrative Ontology, a CIDOC CRM–based ontology developed by CNR-ISTI, ensuring semantic consistency, interoperability, and reusability of the represented knowledge. The workflow supports data exploration and inter-story correlation analysis across the formal narratives stored in the graph, thereby providing a structured and interpretable representation of data that can support evidence-based decision-making processes. Within an Open Science and open-source perspective, the workflow explores the use of freely available, open-weight, small Large Language Models (LLMs), both for knowledge graph construction and for supporting data analysis tasks. These models are evaluated as alternatives to traditional Natural Language Processing pipelines, with the aim of ensuring transparency, reproducibility, and sustainability while avoiding dependencies on proprietary systems. The workflow is validated through its application to the MOVING (MOuntain Valorisation through INterconnectedness and Green Growth) Horizon 2020 European project, which collected and released textual data describing 454 mountain value chains across 16 European countries. The dataset includes textual data stored in a CSV file as well as unstructured raw textual reports describing socio-economic, environmental, and cultural aspects of rural and mountain territories. The proposed workflow is organised into five modular components: (1) Data Pre-processing, which transforms the MOVING data into event-based narratives; (2) Story Structure Building, which extracts meaningful entities from the narratives and links them to a reference knowledge base (i.e. Wikidata); (3) Knowledge Graph Creation, which formalises the narratives as an ontology-based Linked Open Data knowledge graph modelled on the Narrative Ontology; (4) Story Map Visualization and Data Analysis, which enables accessible narrative visualisation and inter-story correlation analysis; and (5) Workflow Assessment, which evaluates the usefulness and relevance of the produced results through the validation of the MOVING experts.

An AI-based automatic workflow for transforming unstructured textual data into a narrative knowledge graph / Lenzi, E.. - ELETTRONICO. - (2026).

An AI-based automatic workflow for transforming unstructured textual data into a narrative knowledge graph

Lenzi Emanuele
2026

Abstract

This research focuses on the design and development of an automatic, AI-based, and open-science–oriented workflow that leverages Semantic Web technologies and a knowledge representation approach to transform unstructured textual data into a knowledge graph of formal narratives. The resulting knowledge graph is modelled according to the Narrative Ontology, a CIDOC CRM–based ontology developed by CNR-ISTI, ensuring semantic consistency, interoperability, and reusability of the represented knowledge. The workflow supports data exploration and inter-story correlation analysis across the formal narratives stored in the graph, thereby providing a structured and interpretable representation of data that can support evidence-based decision-making processes. Within an Open Science and open-source perspective, the workflow explores the use of freely available, open-weight, small Large Language Models (LLMs), both for knowledge graph construction and for supporting data analysis tasks. These models are evaluated as alternatives to traditional Natural Language Processing pipelines, with the aim of ensuring transparency, reproducibility, and sustainability while avoiding dependencies on proprietary systems. The workflow is validated through its application to the MOVING (MOuntain Valorisation through INterconnectedness and Green Growth) Horizon 2020 European project, which collected and released textual data describing 454 mountain value chains across 16 European countries. The dataset includes textual data stored in a CSV file as well as unstructured raw textual reports describing socio-economic, environmental, and cultural aspects of rural and mountain territories. The proposed workflow is organised into five modular components: (1) Data Pre-processing, which transforms the MOVING data into event-based narratives; (2) Story Structure Building, which extracts meaningful entities from the narratives and links them to a reference knowledge base (i.e. Wikidata); (3) Knowledge Graph Creation, which formalises the narratives as an ontology-based Linked Open Data knowledge graph modelled on the Narrative Ontology; (4) Story Map Visualization and Data Analysis, which enables accessible narrative visualisation and inter-story correlation analysis; and (5) Workflow Assessment, which evaluates the usefulness and relevance of the produced results through the validation of the MOVING experts.
2026
Istituto di Scienza e Tecnologie dell'Informazione "Alessandro Faedo" - ISTI
Dottorato
38
Corso 3
Knowledge representation, Knowledge graph, Narrative Ontology, Large Language Models
BARTALESI LENZI, VALENTINA
AMATO, GIUSEPPE
File in questo prodotto:
File Dimensione Formato  
PhD Thesis_Emanuele_Lenzi.pdf

accesso aperto

Descrizione: Tesi di dottorato
Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 3.56 MB
Formato Adobe PDF
3.56 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/600522
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact