ملف الباحث

Yang, Cheng

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

    2024 · arXiv (Cornell University)

    Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained or supplied to act as a proxy for humans in …