Word Sense Disambiguation (WSD) remains a fundamental challenge in Natural Language Processing (NLP) because many words are polysemous, with meanings that vary across contexts. WSD contributes to the semantic interpretation by identifying the intended sense of a word in context. Although substantial progress has been achieved for English due to the availability of rich lexical resources and large sense-annotated corpora, WSD for Arabic dialects remains comparatively underexplored, largely because of the scarcity of annotated corpora and lexico-semantic resources. In parallel, Large Language Models (LLMs) offer promising opportunities to improve WSD in low-resource languages. This paper proposes an LLM-based approach to WSD in the Moroccan dialect (Darija) and evaluates its effectiveness at generating lexical definitions in two settings. In the first setting, the LLM generates a lexical definition based solely on the original context. In the second setting, the definition is generated from the original context augmented with additional contextual sentences previously produced by the LLM. WordNet serves as a reference sense inventory, providing candidate glosses that act as external semantic anchors for disambiguation. The candidate glosses are compared with the LLM-generated definitions in a shared embedding space to select the most appropriate sense. The proposed method is evaluated on both Modern Standard Arabic (MSA) and Darija. Results show that enriching the original context with generated sentences improves the quality of lexical definitions and enhances disambiguation performance in both varieties. Overall, aligning LLM-generated definitions with WordNet glosses via semantic matching provides a practical and effective solution for WSD in low-resource languages.
Word sense disambiguation approach for Moroccan dialect using large language models and a sense inventory
Nahli, OuafaeUltimo
Methodology
2026
Abstract
Word Sense Disambiguation (WSD) remains a fundamental challenge in Natural Language Processing (NLP) because many words are polysemous, with meanings that vary across contexts. WSD contributes to the semantic interpretation by identifying the intended sense of a word in context. Although substantial progress has been achieved for English due to the availability of rich lexical resources and large sense-annotated corpora, WSD for Arabic dialects remains comparatively underexplored, largely because of the scarcity of annotated corpora and lexico-semantic resources. In parallel, Large Language Models (LLMs) offer promising opportunities to improve WSD in low-resource languages. This paper proposes an LLM-based approach to WSD in the Moroccan dialect (Darija) and evaluates its effectiveness at generating lexical definitions in two settings. In the first setting, the LLM generates a lexical definition based solely on the original context. In the second setting, the definition is generated from the original context augmented with additional contextual sentences previously produced by the LLM. WordNet serves as a reference sense inventory, providing candidate glosses that act as external semantic anchors for disambiguation. The candidate glosses are compared with the LLM-generated definitions in a shared embedding space to select the most appropriate sense. The proposed method is evaluated on both Modern Standard Arabic (MSA) and Darija. Results show that enriching the original context with generated sentences improves the quality of lexical definitions and enhances disambiguation performance in both varieties. Overall, aligning LLM-generated definitions with WordNet glosses via semantic matching provides a practical and effective solution for WSD in low-resource languages.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


