ملف الباحث

Yige Li

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

    2024 · arXiv (Cornell University)

    Large language models (LLMs) are increasingly being adopted in a wide range of real-world applications. Despite their impressive performance, recent studies have shown that LLMs are vulnerable to deliberately crafted adversarial prompts even when aligned …