ملف الباحث
Kunpeng Ning
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin
2025 · arXiv (Cornell University)
Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely compromise safety measures. We observe that perturbations orthogonal to the alignment direction - defined by weight differences …