conference-paper

A Sublinear-Regret Reinforcement Learning Algorithm on Constrained Markov Decision Processes with reset action

Research footprint

At a glance

Citations
0
References
6
Comments
0
Paper overview

Abstract

In this paper, we study model-based reinforcement learning in an unknown constrained Markov Decision Processes (CMDPs) with reset action. We propose an algorithm, Constrained-UCRL, which uses confidence interval like UCRL2, and solves linear programming problem to compute policy at the start of each episode. We show that Constrained-UCRL achieves sublinear regret bounds Õ(SA1/2T3/4) up to logarithmic factors with high probability for both the gain and the constraint violations.

Record transparency

Publication details

DOI
10.1145/3380688.3380706
OpenAlex
W3009923408
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.