Researcher profile

Yunhao Tang

2 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Self-Imitation Learning via Generalized Lower Bound Q-learning

    2020 · arXiv (Cornell University)

    Self-imitation learning motivated by lower-bound Q-learning is a novel and effective approach for off-policy learning. In this work, we propose a n-step lower bound which generalizes the original return-based lower-bound Q-learning, and introduce a new …

  2. The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement Learning

    2022 · arXiv (Cornell University)

    We study the multi-step off-policy learning approach to distributional RL. Despite the apparent similarity between value-based RL and distributional RL, our study reveals intriguing and fundamental differences between the two cases in the multi-step setting. …