Researcher profile
Zhongzhi Li
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
2025 · arXiv (Cornell University)
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), particularly for complex reasoning tasks. However, vanilla RLVR training has been shown to improve Pass@1 performance …