Large Language Models (LLMs) are increasingly central to Alzheimer’s disease (AD) research, particularly for classifying speech-based transcripts. However, a significant factor affecting LLM performance is the choice of a prompting strategy. Despite the known sensitivity of these models, AD research currently lacks a formal framework for assessing prompt quality or stability. Using the ADReSS Challenge dataset, this study evaluates whether prompts used for AD speech-based classification are sensitive to minor linguistic variations and tests whether a prompt optimization framework can reduce this sensitivity. Our results demonstrate that various models are highly sensitive to small changes in prompts. Notably, optimization led to a significant 16.63% improvement in accuracy for the Mixtral 8x7B model. These findings highlight the necessity of validating and optimizing prompts to ensure reliable diagnostic outputs in clinical research.

Empirical Evaluation of Prompt Sensitivity and Optimization in Alzheimer Disease Speech Classification

Santopaolo A.;Giugliano S.;Sannino G.
2026

Abstract

Large Language Models (LLMs) are increasingly central to Alzheimer’s disease (AD) research, particularly for classifying speech-based transcripts. However, a significant factor affecting LLM performance is the choice of a prompting strategy. Despite the known sensitivity of these models, AD research currently lacks a formal framework for assessing prompt quality or stability. Using the ADReSS Challenge dataset, this study evaluates whether prompts used for AD speech-based classification are sensitive to minor linguistic variations and tests whether a prompt optimization framework can reduce this sensitivity. Our results demonstrate that various models are highly sensitive to small changes in prompts. Notably, optimization led to a significant 16.63% improvement in accuracy for the Mixtral 8x7B model. These findings highlight the necessity of validating and optimizing prompts to ensure reliable diagnostic outputs in clinical research.
2026
Istituto di Calcolo e Reti ad Alte Prestazioni - ICAR - Sede Secondaria Napoli
9783032292537
9783032292544
Alzheimer’s Disease
Large Language Models
Prompt Optimization
Prompt Sensitivity
Speech Classification
File in questo prodotto:
File Dimensione Formato  
Empirical Evaluation of Prompt Sensitivity and Optimization in Alzheimer Disease Speech Classification.pdf

solo utenti autorizzati

Tipologia: Documento in Pre-print
Licenza: Creative commons
Dimensione 775.22 kB
Formato Adobe PDF
775.22 kB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/592401
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact