ملف الباحث
Siqiao Mu
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
On the Second-Order Convergence of Biased Policy Gradient Algorithms
2023 · arXiv (Cornell University)
Since the objective functions of reinforcement learning problems are typically highly nonconvex, it is desirable that policy gradient, the most popular algorithm, escapes saddle points and arrives at second-order stationary points. Existing results only consider …