AI models in critical sectors such as healthcare and finance must provide both data privacy and adversarial robustness. Differ-ential Privacy (DP) protects training data by injecting noise, but this noise smooths decision boundaries and leaves models opento adversarial evasion, a tension known as the Privacy-Robustness Trade-off. Although this trade-off is well documented, itsinternal mechanism remains underexplored: prior work does not reveal how the noise reshapes a model’s reasoning or whichfeatures become vulnerable. To close this gap, we propose a Privacy-Aware Adversarial Defense grounded in ExplainableAI. Specifically, we introduce the Attention Concentration Score (ACS), which measures how DP training shifts Transformerattention away from task-critical features toward non-functional ones. This attention drift correlates with adversarial vul-nerability, providing a mechanistic explanation of the trade-off. Building on this insight, we develop a Manifold-AlignedSemantic Attack that targets the most drifted features, and a TrustScore defense that fuses embedding-level anomaly detec-tion with attention-level consistency checks. We validate across two datasets (Adult Census, MIMIC-IV), two architectures(DeBERTa-V3-Large, LLaMA−3.1-8B), and seven experiments benchmarking five attacks against six defenses. Within therecommended range (epsilon ∈ [5, 10]), models retain 84.2% accuracy (96.2% of baseline), while TrustScore reaches an Area Underthe ROC Curve (AUC) of 0.87-−0.94, outperforming Isolation Forest (0.65) and supervised detection (0.58). Moreover, theseconclusions hold under feature-categorization variants (attribution- and PCA-based); privacy noise disproportionately desta-bilizes the minority class; the consistency signal adds sub-millisecond overhead; and a surrogate-attention variant preservesdetection under black-box deployment, establishing the approach’s dependability for reliable intelligent environments.

Privacy-aware adversarial defense with explainable AI for adversarial robustness in AI model

Stefano Silvestri;
2026

Abstract

AI models in critical sectors such as healthcare and finance must provide both data privacy and adversarial robustness. Differ-ential Privacy (DP) protects training data by injecting noise, but this noise smooths decision boundaries and leaves models opento adversarial evasion, a tension known as the Privacy-Robustness Trade-off. Although this trade-off is well documented, itsinternal mechanism remains underexplored: prior work does not reveal how the noise reshapes a model’s reasoning or whichfeatures become vulnerable. To close this gap, we propose a Privacy-Aware Adversarial Defense grounded in ExplainableAI. Specifically, we introduce the Attention Concentration Score (ACS), which measures how DP training shifts Transformerattention away from task-critical features toward non-functional ones. This attention drift correlates with adversarial vul-nerability, providing a mechanistic explanation of the trade-off. Building on this insight, we develop a Manifold-AlignedSemantic Attack that targets the most drifted features, and a TrustScore defense that fuses embedding-level anomaly detec-tion with attention-level consistency checks. We validate across two datasets (Adult Census, MIMIC-IV), two architectures(DeBERTa-V3-Large, LLaMA−3.1-8B), and seven experiments benchmarking five attacks against six defenses. Within therecommended range (epsilon ∈ [5, 10]), models retain 84.2% accuracy (96.2% of baseline), while TrustScore reaches an Area Underthe ROC Curve (AUC) of 0.87-−0.94, outperforming Isolation Forest (0.65) and supervised detection (0.58). Moreover, theseconclusions hold under feature-categorization variants (attribution- and PCA-based); privacy noise disproportionately desta-bilizes the minority class; the consistency signal adds sub-millisecond overhead; and a surrogate-attention variant preservesdetection under black-box deployment, establishing the approach’s dependability for reliable intelligent environments.
2026
Istituto di Calcolo e Reti ad Alte Prestazioni - ICAR - Sede Secondaria Napoli
Differential privacy
Adversarial robustness
Unsupervised anomaly detection
Explainable AI
Latent spacemonitoring
Reliable intelligent environments
Fairness
File in questo prodotto:
File Dimensione Formato  
s40860-026-00273-7.pdf

accesso aperto

Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 6.85 MB
Formato Adobe PDF
6.85 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14243/598907
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact