Researcher profile

X. Su

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

    2025 · arXiv (Cornell University)

    Inference scaling further accelerates Large Language Models (LLMs) toward Artificial General Intelligence (AGI), with large-scale Reinforcement Learning (RL) to unleash long Chain-of-Thought reasoning. Most contemporary reasoning approaches usually rely on handcrafted rule-based reward functions. However, …