Researcher profile
Danqing Shi
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Interactive Groupwise Comparison for Reinforcement Learning from Human Feedback
2025 · arXiv (Cornell University)
Reinforcement learning from human feedback (RLHF) has emerged as a key enabling technology for aligning AI behaviour with human preferences. The traditional way to collect data in RLHF is via pairwise comparisons: human raters are …