ملف الباحث
Tianheng Song
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Online Sparse Temporal Difference Learning Based on Nested Optimization and Regularized Dual Averaging
2021 · IEEE Transactions on Systems Man and Cybernetics Systems
In policy evaluation of reinforcement learning tasks, the temporal difference (TD) learning with value function approximation has been widely studied. However, feature representation has a decisive influence on both accuracy of value function approximation and …