Researcher profile

David Dobre

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. In-Context Learning Can Re-learn Forbidden Tasks

    2024 · arXiv (Cornell University)

    Despite significant investment into safety training, large language models (LLMs) deployed in the real world still suffer from numerous vulnerabilities. One perspective on LLM safety training is that it algorithmically forbids the model from answering …