Researcher profile
Yulin Hu
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
2024 · arXiv (Cornell University)
Safety alignment of large language models (LLMs) has been gaining increasing attention. However, current safety-aligned LLMs suffer from the fragile and imbalanced safety mechanisms, which can still be induced to generate unsafe responses, exhibit over-safety …