ملف الباحث
Keyu Duan
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Efficient Process Reward Model Training via Active Learning
2025 · arXiv (Cornell University)
Process Reward Models (PRMs) provide step-level supervision to large language models (LLMs), but scaling up training data annotation remains challenging for both humans and LLMs. To address this limitation, we propose an active learning approach, …