Researcher profile
Wang, Zongqi
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Probing the Robustness of Large Language Models Safety to Latent Perturbations
2025 · arXiv (Cornell University)
Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that …