ملف الباحث
John Schulman
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Scaling Laws for Reward Model Overoptimization
2022 · arXiv (Cornell University)
In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences. Because the reward model is an imperfect proxy, optimizing its value too much can hinder …
-
Training language models to follow instructions with human feedback
2022 · arXiv (Cornell University)
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In …