ملف الباحث
Honghao Wei
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
A Provably-Efficient Model-Free Algorithm for Infinite-Horizon Average-Reward Constrained Markov Decision Processes
2022 · Proceedings of the AAAI Conference on Artificial Intelligence
This paper presents a model-free reinforcement learning (RL) algorithm for infinite-horizon average-reward Constrained Markov Decision Processes (CMDPs). Considering a learning horizon K, which is sufficiently large, the proposed algorithm achieves sublinear regret and zero constraint …