ملف الباحث

Wang, Maolin

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration

    2025 · arXiv (Cornell University)

    Reinforcement Learning (RL) has become a key approach for enhancing the reasoning capabilities of large language models. However, prevalent RL approaches like proximal policy optimization and group relative policy optimization suffer from sparse, outcome-based rewards …