ملف الباحث

Tianheng Song

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Online Sparse Temporal Difference Learning Based on Nested Optimization and Regularized Dual Averaging

    2021 · IEEE Transactions on Systems Man and Cybernetics Systems

    In policy evaluation of reinforcement learning tasks, the temporal difference (TD) learning with value function approximation has been widely studied. However, feature representation has a decisive influence on both accuracy of value function approximation and …