ملف الباحث
Teng, Yan
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Probing the Robustness of Large Language Models Safety to Latent Perturbations
2025 · arXiv (Cornell University)
Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that …