ملف الباحث

Maurits Kaptein

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Provably Efficient Exploration in Constrained Reinforcement Learning:Posterior Sampling Is All You Need

    2023 · arXiv (Cornell University)

    We present a new algorithm based on posterior sampling for learning in constrained Markov decision processes (CMDP) in the infinite-horizon undiscounted setting. The algorithm achieves near-optimal regret bounds while being advantageous empirically compared to the …