ملف الباحث
Zhou, Ruochen
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
2025 · arXiv (Cornell University)
Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, …