ملف الباحث
Bowen Qin
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
2024 · arXiv (Cornell University)
Direct Preference Optimization (DPO), which derives reward signals directly from pairwise preference data, has shown its effectiveness on aligning Large Language Models (LLMs) with human preferences. Despite its widespread use across various tasks, DPO has …