ملف الباحث

Qifei Wang

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models

    2024 · arXiv (Cornell University)

    Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using reinforcement learning (RL) to maximize rewards that reflect human …