ملف الباحث

Wenlong Deng

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning

    2025 · arXiv (Cornell University)

    Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward exploration or exploitation remains an open problem. We introduce Token Hidden Reward (THR), …