Researcher profile

Du Su

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Indirect Online Preference Optimization via Reinforcement Learning

    2025

    Human preference alignment (HPA) aims to ensure Large Language Models (LLMs) responding appropriately to meet human moral and ethical requirements. Existing methods, such as RLHF and DPO, rely heavily on high-quality human annotation, which restrict …