ملف الباحث
Q. He
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
2025
Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful technique for aligning large language models (LLMs) with human preferences.However, effectively aligning LLMs with diverse human preferences remains a significant challenge, particularly when they …