Enhancing Adversarial Robustness With Energy-Based Models: a Novel Training Approach for Input Purification
At a glance
- Citations
- 0
- References
- 38
- Comments
- 0
Öz
Deep learning has advanced in fields like computer vision and natural language processing but still faces reliability challenges, including adversarial attacks that mislead models and threaten security-sensitive applications. Generative models help counter such threats by purifying adversarial images. Among them, energy-based models (EBMs) are effective for adversarial purification. Our approach utilizes EBMs for purification of inputs into a classifier while incorporating adversarial examples of that classifier into training procedure of the EBM. The proposed EBM training technique outperforms vanilla training procedure by 10.52 % and 20.48 % on MNIST and FashionMNIST, respectively, for the purification task. In our experiment setting, robust accuracy, defined as the accuracy of adversarial examples in the test set after the purification process, achieves 48 % against benchmark AutoAttack ($L_{\infty}, \epsilon=8 / 255$) on CIFAR10, improving the state-of-the-art purification method using EBM by 1.1 %, with a clean accuracy of 70 %.
Publication details
- DOI
- 10.1109/iwcit67595.2025.11073010
- OpenAlex
- W4412129802
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.