ملف الباحث

Di Huang

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions

    2026 · Proceedings of the AAAI Conference on Artificial Intelligence

    With the widespread application of Large Language Models (LLMs), it has become a significant concern to ensure their safety and prevent harmful responses. While current safe-alignment methods based on instruction fine-tuning and Reinforcement Learning from …