ملف الباحث
Madhavan, Rahul
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
2024 · arXiv (Cornell University)
Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily long responses. In reinforcement learning from human feedback …