ملف الباحث

Dhanya Sridhar

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. In-Context Learning Can Re-learn Forbidden Tasks

    2024 · arXiv (Cornell University)

    Despite significant investment into safety training, large language models (LLMs) deployed in the real world still suffer from numerous vulnerabilities. One perspective on LLM safety training is that it algorithmically forbids the model from answering …