ملف الباحث
Xuehang Cang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Selective Component Ablation for Refusal Removal
2026 · Zenodo (CERN European Organization for Nuclear Research)
Safety alignment in large language models (LLMs) is predominantly achieved through training models to refuse harmful requests. However, prior work has demonstrated that refusal behavior is mediated by a single direction in the residual stream …