ملف الباحث

David C. Parkes

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Chain-of-Thought Reasoning is a Policy Improvement Operator

    2023 · arXiv (Cornell University)

    Large language models have astounded the world with fascinating new capabilities. However, they currently lack the ability to teach themselves new skills, relying instead on large amounts of human-generated training data. We introduce SECToR (Self-Education …