ملف الباحث
A. W. Woodruff
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
2025 · arXiv (Cornell University)
Language model finetuning often results in learning undesirable traits in combination with desired ones. To address this, we propose inoculation prompting: modifying finetuning data by prepending a short system-prompt instruction that deliberately elicits the undesirable …