Researcher profile
Guoxi Zhang
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback
2024 · arXiv (Cornell University)
This paper addresses the cost-efficiency aspect of Reinforcement Learning from Human Feedback (RLHF). RLHF leverages datasets of human preferences over outputs of large language models (LLM)s to instill human expectations into LLMs. Although preference annotation …