ملف الباحث

Duan, Jiuding

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback

    2024 · arXiv (Cornell University)

    This paper addresses the cost-efficiency aspect of Reinforcement Learning from Human Feedback (RLHF). RLHF leverages datasets of human preferences over outputs of large language models (LLM)s to instill human expectations into LLMs. Although preference annotation …