ملف الباحث
Runlong Zhou
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning
2023 · arXiv (Cornell University)
Off-policy dynamic programming (DP) techniques such as $Q$-learning have proven to be important in sequential decision-making problems. In the presence of function approximation, however, these techniques often diverge due to the absence of Bellman completeness …