ملف الباحث
Zhuokai Zhao
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
2025 · arXiv (Cornell University)
Reinforcement learning with verifiable rewards (RLVR) is a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs), yet its success hinges on effective exploration. An ideal exploration strategy must navigate two fundamental …