Researcher profile
Chen, Weizhu
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
2025 · arXiv (Cornell University)
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), particularly for complex reasoning tasks. However, vanilla RLVR training has been shown to improve Pass@1 performance …
-
LoRA Fine-Tuning of a 3B Code LLM for Algorithmic Efficiency
2021 · arXiv (Cornell University)
An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larger models, full fine-tuning, which retrains all model parameters, becomes …