ملف الباحث
Wenlong Deng
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
2025 · arXiv (Cornell University)
Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward exploration or exploitation remains an open problem. We introduce Token Hidden Reward (THR), …