ملف الباحث

Zhou, Ruochen

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

    2025 · arXiv (Cornell University)

    Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, …