preprint
Open access
Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities
Research footprint
At a glance
- Citations
- 12
- References
- 22
- Comments
- 0
Paper overview
Abstract
We study model-based reinforcement learning in an unknown finite communicating Markov decision process. We propose a simple algorithm that leverages a variance based confidence interval. We show that the proposed algorithm, UCRL-V, achieves the optimal regret $\tilde{\mathcal{O}}(\sqrt{DSAT})$ up to logarithmic factors, and so our work closes a gap with the lower bound without additional assumptions on the MDP. We perform experiments in a variety of environments that validates the theoretical bounds as well as prove UCRL-V to be better than the state-of-the-art algorithms.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.1905.12425
- OpenAlex
- W2947782105
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.