ملف الباحث
Sander Schulhoff
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
2025 · arXiv (Cornell University)
How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful knowledge or remotely triggering malicious actions, respectively) are typically …