conference-paper
وصول مفتوح
A Provably-Efficient Model-Free Algorithm for Infinite-Horizon Average-Reward Constrained Markov Decision Processes
Research footprint
At a glance
- الاستشهادات
- 10
- المراجع
- 46
- Comments
- 0
Paper overview
Abstract
This paper presents a model-free reinforcement learning (RL) algorithm for infinite-horizon average-reward Constrained Markov Decision Processes (CMDPs). Considering a learning horizon K, which is sufficiently large, the proposed algorithm achieves sublinear regret and zero constraint violation. The bounds depend on the number of states S, the number of actions A, and two constants which are independent of the learning horizon K.
Record transparency
Publication details
- DOI
- 10.1609/aaai.v36i4.20302
- OpenAlex
- W4283805055
- Document type
- conference-paper
- Language
- EN
- Source
- Proceedings of the AAAI Conference on Artificial Intelligence
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.