ملف الباحث
Wouter Jongeneel
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
A Large Deviations Perspective on Policy Gradient Algorithms
2023 · arXiv (Cornell University)
Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent for possibly non-convex objectives satisfying a Polyak-Łojasiewicz condition. Leveraging …