ملف الباحث

Niels Warncke

ورقتان في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time

    2025 · arXiv (Cornell University)

    Language model finetuning often results in learning undesirable traits in combination with desired ones. To address this, we propose inoculation prompting: modifying finetuning data by prepending a short system-prompt instruction that deliberately elicits the undesirable …

  2. Training large language models on narrow tasks can lead to broad misalignment

    2026 · Nature

    Abstract The widespread adoption of large language models (LLMs) raises important questions about their safety and alignment 1 . Previous safety research has largely focused on isolated undesirable behaviours, such as reinforcing harmful stereotypes or …