Researcher profile

Zhongzhi Li

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR

    2025 · arXiv (Cornell University)

    Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), particularly for complex reasoning tasks. However, vanilla RLVR training has been shown to improve Pass@1 performance …