preprint Open access

Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
12
References
22
Comments
0
Paper overview

Abstract

We study model-based reinforcement learning in an unknown finite communicating Markov decision process. We propose a simple algorithm that leverages a variance based confidence interval. We show that the proposed algorithm, UCRL-V, achieves the optimal regret $\tilde{\mathcal{O}}(\sqrt{DSAT})$ up to logarithmic factors, and so our work closes a gap with the lower bound without additional assumptions on the MDP. We perform experiments in a variety of environments that validates the theoretical bounds as well as prove UCRL-V to be better than the state-of-the-art algorithms.

Record transparency

Publication details

DOI
10.48550/arxiv.1905.12425
OpenAlex
W2947782105
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.