ملف الباحث

Sui, Xingyu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

    2024 · arXiv (Cornell University)

    Safety alignment of large language models (LLMs) has been gaining increasing attention. However, current safety-aligned LLMs suffer from the fragile and imbalanced safety mechanisms, which can still be induced to generate unsafe responses, exhibit over-safety …