ملف الباحث

Mark Rowland

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement Learning

    2022 · arXiv (Cornell University)

    We study the multi-step off-policy learning approach to distributional RL. Despite the apparent similarity between value-based RL and distributional RL, our study reveals intriguing and fundamental differences between the two cases in the multi-step setting. …