conference-paper
A Sublinear-Regret Reinforcement Learning Algorithm on Constrained Markov Decision Processes with reset action
Research footprint
At a glance
- Citations
- 0
- References
- 6
- Comments
- 0
Paper overview
Abstract
In this paper, we study model-based reinforcement learning in an unknown constrained Markov Decision Processes (CMDPs) with reset action. We propose an algorithm, Constrained-UCRL, which uses confidence interval like UCRL2, and solves linear programming problem to compute policy at the start of each episode. We show that Constrained-UCRL achieves sublinear regret bounds Õ(SA1/2T3/4) up to logarithmic factors with high probability for both the gain and the constraint violations.
Record transparency
Publication details
- DOI
- 10.1145/3380688.3380706
- OpenAlex
- W3009923408
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.