ملف الباحث

Siqiao Mu

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. On the Second-Order Convergence of Biased Policy Gradient Algorithms

    2023 · arXiv (Cornell University)

    Since the objective functions of reinforcement learning problems are typically highly nonconvex, it is desirable that policy gradient, the most popular algorithm, escapes saddle points and arrives at second-order stationary points. Existing results only consider …