ملف الباحث

Yongqian Xu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF

    2025 · arXiv (Cornell University)

    Quality of datasets plays an important role in large language model (LLM) alignment. In collecting human feedback, however, preference flipping is ubiquitous and causes corruption in data annotation; the issue necessitates the alignment algorithms with …