preprint Open access

Vulnerability of Large Language Models to Output Prefix Jailbreaks: Impact of Positions on Safety

Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Previous research on jailbreak attacks has mainly focused on optimizing the adversarial snippet content injected into input prompts to expose LLM security vulnerabilities. A significant portion of this research focuses on developing more complex, less readable adversarial snippets that can achieve higher attack success rates. In contrast to this trend, our research investigates the impact of the adversarial snippet's position on the effectiveness of jailbreak attacks. We find that placing a simple and readable adversarial snippet at the beginning of the output effectively exposes LLM safety vulnerabilities, leading to much higher attack success rates than the input suffix attack or promptbased output jailbreaks. Precisely speaking, we discover that directly enforcing the user's target embedded output prefix is an effective method to expose LLMs' safety vulnerabilities.

Record transparency

Publication details

DOI
10.36227/techrxiv.173161123.33560859/v1
OpenAlex
W4404349720
Document type
preprint
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.