Researcher profile
H. L. Dai
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin
2025 · arXiv (Cornell University)
Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely compromise safety measures. We observe that perturbations orthogonal to the alignment direction - defined by weight differences …