ملف الباحث
X. Su
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
2025 · arXiv (Cornell University)
Inference scaling further accelerates Large Language Models (LLMs) toward Artificial General Intelligence (AGI), with large-scale Reinforcement Learning (RL) to unleash long Chain-of-Thought reasoning. Most contemporary reasoning approaches usually rely on handcrafted rule-based reward functions. However, …