ملف الباحث
David C. Parkes
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Chain-of-Thought Reasoning is a Policy Improvement Operator
2023 · arXiv (Cornell University)
Large language models have astounded the world with fascinating new capabilities. However, they currently lack the ability to teach themselves new skills, relying instead on large amounts of human-generated training data. We introduce SECToR (Self-Education …