ملف الباحث
Xianzhou Zeng
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
NGRPO: Negative-enhanced Group Relative Policy Optimization
2025 · arXiv (Cornell University)
RLVR has enhanced the reasoning capabilities of Large Language Models (LLMs) across various tasks. However, GRPO, a representative RLVR algorithm, suffers from a critical limitation: when all responses within a group are either entirely correct …