ملف الباحث

Chen-Yu Wei

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs

    2023 · arXiv (Cornell University)

    We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-based policy search methods used in practice, the oscillation of policy iterates …