ملف الباحث
Qifei Wang
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models
2024 · arXiv (Cornell University)
Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using reinforcement learning (RL) to maximize rewards that reflect human …