ملف الباحث

Qipeng Guo

ورقتان في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-Thoughts

    2023

    As large language models (LLMs) have shown effectiveness with different prompting methods, such as Chain of Thought, Program of Thought, we find that these methods have formed a great complementarity to each other on math …

  2. Pre-Trained Policy Discriminators are General Reward Models

    2025

    We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy …