preprint Open access

Gradient-based Adversarial Attacks against Text Transformers

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
2
References
36
Comments
0
Paper overview

Abstract

We propose the first general-purpose gradient-based attack against transformer models. Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization. We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks. Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs.

Record transparency

Publication details

DOI
10.48550/arxiv.2104.13733
OpenAlex
W3213493070
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.