ملف الباحث
Guoxi Zhang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback
2024 · arXiv (Cornell University)
This paper addresses the cost-efficiency aspect of Reinforcement Learning from Human Feedback (RLHF). RLHF leverages datasets of human preferences over outputs of large language models (LLM)s to instill human expectations into LLMs. Although preference annotation …