ملف الباحث
Yilin Zhang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
2024 · arXiv (Cornell University)
Large Language Models (LLMs) have excelled in various tasks but are still vulnerable to jailbreaking attacks, where attackers create jailbreak prompts to mislead the model to produce harmful or offensive content. Current jailbreak methods either …