ملف الباحث

Wouter Jongeneel

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. A Large Deviations Perspective on Policy Gradient Algorithms

    2023 · arXiv (Cornell University)

    Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent for possibly non-convex objectives satisfying a Polyak-Łojasiewicz condition. Leveraging …