ملف الباحث

X. Su

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

    2025 · arXiv (Cornell University)

    Inference scaling further accelerates Large Language Models (LLMs) toward Artificial General Intelligence (AGI), with large-scale Reinforcement Learning (RL) to unleash long Chain-of-Thought reasoning. Most contemporary reasoning approaches usually rely on handcrafted rule-based reward functions. However, …