ملف الباحث

Richard S. Sutton

3 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Multi-Step Reinforcement Learning: A Unifying Algorithm

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    Unifying seemingly disparate algorithmic ideas to produce better performing algorithms has been a longstanding goal in reinforcement learning. As a primary example, TD(λ) elegantly unifies one-step TD prediction with Monte Carlo methods through the use …

  2. Per-decision Multi-step Temporal Difference Learning with Control Variates

    2018 · arXiv (Cornell University)

    Multi-step temporal difference (TD) learning is an important approach in reinforcement learning, as it unifies one-step TD learning with Monte Carlo methods in a way where intermediate algorithms can outperform either extreme. They address a …

  3. Average-Reward Learning and Planning with Options

    2021 · arXiv (Cornell University)

    We extend the options framework for temporal abstraction in reinforcement learning from discounted Markov decision processes (MDPs) to average-reward MDPs. Our contributions include general convergent off-policy inter-option learning algorithms, intra-option algorithms for learning values and …