article وصول مفتوح

Balancing explainability strength and privacy in counterfactual explanations for insider threat detection

  • Greenwich Academic Literature Archive (University of Greenwich)
  • University of Greenwich
Research footprint

At a glance

الاستشهادات
0
المراجع
0
Comments
0
Paper overview

Abstract

Explainable AI (XAI) methods, such as counterfactual explanations, help make model decisions more transparent and actionable. However, they can unintentionally reveal sensitive information, which is a critical concern in security settings such as insider threat detection, where both transparency and privacy are vital, particularly when explanations are shared with external analysts or decision-makers. We propose a privacy-aware counterfactual framework that integrates Differential Privacy (DP) to add only the minimum adjusted noise required to satisfy user-defined privacy caps, reducing information leakage while preserving explanation fidelity. Two task-specific metrics are introduced: XAIStrength, which measures how faithfully a counterfactual follows the model’s decision boundary, and XAILeakage, which quantifies privacy risk based on movement into sparse feature regions. Experiments on the CERT insider threat dataset show that initially generated counterfactuals were highly faithful for most users but frequently violated privacy requirements. After optimisation, leakage was reduced to meet strict, moderate, and lenient privacy caps (τ = 0.05/0.10/0.20) while maintaining practical fidelity. More than 60% of users retained Sopt ≥ 0.70 under strict privacy constraints, while over 60% maintained Sopt ≥ 0.85 under moderate privacy settings. Comparative evaluation against standard and privacy-aware baselines shows that the proposed framework achieves a stronger balance between explanation utility and disclosure risk. To the best of our knowledge, this is the first study to formally quantify and optimise the XAIStrength–XAILeakage trade-off in counterfactual explanations using statistically grounded risk measures and DP-calibrated optimisation.

Record transparency

Publication details

OpenAlex
W7171910555
Document type
article
Language
EN
Source
Greenwich Academic Literature Archive (University of Greenwich)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.