Researcher profile
Rongwu Xu
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
On the Role of Attention Heads in Large Language Model Safety
2024 · arXiv (Cornell University)
Large language models (LLMs) achieve state-of-the-art performance on multiple language tasks, yet their safety guardrails can be circumvented, leading to harmful generations. In light of this, recent research on safety mechanisms has emerged, revealing that …