ملف الباحث

Qingkai Liang

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning

    2018 · arXiv (Cornell University)

    Constrained Markov Decision Process (CMDP) is a natural framework for reinforcement learning tasks with safety constraints, where agents learn a policy that maximizes the long-term reward while satisfying the constraints on the long-term cost. A …