Researcher profile
Zhuokai Zhao
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
2025 · arXiv (Cornell University)
Reinforcement learning with verifiable rewards (RLVR) is a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs), yet its success hinges on effective exploration. An ideal exploration strategy must navigate two fundamental …