Researcher profile
Yunkun Xu
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
2025
Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful technique for aligning large language models (LLMs) with human preferences.However, effectively aligning LLMs with diverse human preferences remains a significant challenge, particularly when they …