ملف الباحث
Chen-Yu Wei
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
2023 · arXiv (Cornell University)
We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-based policy search methods used in practice, the oscillation of policy iterates …