ملف الباحث
Qipeng Guo
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-Thoughts
2023
As large language models (LLMs) have shown effectiveness with different prompting methods, such as Chain of Thought, Program of Thought, we find that these methods have formed a great complementarity to each other on math …
-
Pre-Trained Policy Discriminators are General Reward Models
2025
We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy …