AI models in critical sectors such as healthcare and finance must provide both data privacy and adversarial robustness. Differ-ential Privacy (DP) protects training data by injecting noise, but this noise smooths decision boundaries and leaves models opento adversarial evasion, a tension known as the Privacy-Robustness Trade-off. Although this trade-off is well documented, itsinternal mechanism remains underexplored: prior work does not reveal how the noise reshapes a model’s reasoning or whichfeatures become vulnerable. To close this gap, we propose a Privacy-Aware Adversarial Defense grounded in ExplainableAI. Specifically, we introduce the Attention Concentration Score (ACS), which measures how DP training shifts Transformerattention away from task-critical features toward non-functional ones. This attention drift correlates with adversarial vul-nerability, providing a mechanistic explanation of the trade-off. Building on this insight, we develop a Manifold-AlignedSemantic Attack that targets the most drifted features, and a TrustScore defense that fuses embedding-level anomaly detec-tion with attention-level consistency checks. We validate across two datasets (Adult Census, MIMIC-IV), two architectures(DeBERTa-V3-Large, LLaMA−3.1-8B), and seven experiments benchmarking five attacks against six defenses. Within therecommended range (epsilon ∈ [5, 10]), models retain 84.2% accuracy (96.2% of baseline), while TrustScore reaches an Area Underthe ROC Curve (AUC) of 0.87-−0.94, outperforming Isolation Forest (0.65) and supervised detection (0.58). Moreover, theseconclusions hold under feature-categorization variants (attribution- and PCA-based); privacy noise disproportionately desta-bilizes the minority class; the consistency signal adds sub-millisecond overhead; and a surrogate-attention variant preservesdetection under black-box deployment, establishing the approach’s dependability for reliable intelligent environments.
Privacy-aware adversarial defense with explainable AI for adversarial robustness in AI model
Stefano Silvestri;
2026
Abstract
AI models in critical sectors such as healthcare and finance must provide both data privacy and adversarial robustness. Differ-ential Privacy (DP) protects training data by injecting noise, but this noise smooths decision boundaries and leaves models opento adversarial evasion, a tension known as the Privacy-Robustness Trade-off. Although this trade-off is well documented, itsinternal mechanism remains underexplored: prior work does not reveal how the noise reshapes a model’s reasoning or whichfeatures become vulnerable. To close this gap, we propose a Privacy-Aware Adversarial Defense grounded in ExplainableAI. Specifically, we introduce the Attention Concentration Score (ACS), which measures how DP training shifts Transformerattention away from task-critical features toward non-functional ones. This attention drift correlates with adversarial vul-nerability, providing a mechanistic explanation of the trade-off. Building on this insight, we develop a Manifold-AlignedSemantic Attack that targets the most drifted features, and a TrustScore defense that fuses embedding-level anomaly detec-tion with attention-level consistency checks. We validate across two datasets (Adult Census, MIMIC-IV), two architectures(DeBERTa-V3-Large, LLaMA−3.1-8B), and seven experiments benchmarking five attacks against six defenses. Within therecommended range (epsilon ∈ [5, 10]), models retain 84.2% accuracy (96.2% of baseline), while TrustScore reaches an Area Underthe ROC Curve (AUC) of 0.87-−0.94, outperforming Isolation Forest (0.65) and supervised detection (0.58). Moreover, theseconclusions hold under feature-categorization variants (attribution- and PCA-based); privacy noise disproportionately desta-bilizes the minority class; the consistency signal adds sub-millisecond overhead; and a surrogate-attention variant preservesdetection under black-box deployment, establishing the approach’s dependability for reliable intelligent environments.| File | Dimensione | Formato | |
|---|---|---|---|
|
s40860-026-00273-7.pdf
accesso aperto
Tipologia:
Versione Editoriale (PDF)
Licenza:
Creative commons
Dimensione
6.85 MB
Formato
Adobe PDF
|
6.85 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


