ملف الباحث

Duanyu Feng

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

    2024 · arXiv (Cornell University)

    Direct Preference Optimization (DPO), which derives reward signals directly from pairwise preference data, has shown its effectiveness on aligning Large Language Models (LLMs) with human preferences. Despite its widespread use across various tasks, DPO has …