ملف الباحث

Zhongzhi Li

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR

    2025 · arXiv (Cornell University)

    Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), particularly for complex reasoning tasks. However, vanilla RLVR training has been shown to improve Pass@1 performance …