ملف الباحث

Xuehang Cang

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Selective Component Ablation for Refusal Removal

    2026 · Zenodo (CERN European Organization for Nuclear Research)

    Safety alignment in large language models (LLMs) is predominantly achieved through training models to refuse harmful requests. However, prior work has demonstrated that refusal behavior is mediated by a single direction in the residual stream …