ملف الباحث
Di Huang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions
2026 · Proceedings of the AAAI Conference on Artificial Intelligence
With the widespread application of Large Language Models (LLMs), it has become a significant concern to ensure their safety and prevent harmful responses. While current safe-alignment methods based on instruction fine-tuning and Reinforcement Learning from …