Researcher profile

Wang, Yixu

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Probing the Robustness of Large Language Models Safety to Latent Perturbations

    2025 · arXiv (Cornell University)

    Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that …