ملف الباحث
Dhanya Sridhar
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
In-Context Learning Can Re-learn Forbidden Tasks
2024 · arXiv (Cornell University)
Despite significant investment into safety training, large language models (LLMs) deployed in the real world still suffer from numerous vulnerabilities. One perspective on LLM safety training is that it algorithmically forbids the model from answering …