Researcher profile
Du Su
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Indirect Online Preference Optimization via Reinforcement Learning
2025
Human preference alignment (HPA) aims to ensure Large Language Models (LLMs) responding appropriately to meet human moral and ethical requirements. Existing methods, such as RLHF and DPO, rely heavily on high-quality human annotation, which restrict …