ملف الباحث

Nisan Stiennon

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Fine-Tuning Language Models from Human Preferences

    2019 · arXiv (Cornell University)

    Reward learning enables the application of reinforcement learning (RL) to tasks where reward is defined by human judgment, building a model of reward by asking humans questions. Most work on reward learning has used simulated …