preprint
وصول مفتوح
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
Research footprint
At a glance
- الاستشهادات
- 1
- المراجع
- 0
- Comments
- 0
Paper overview
Abstract
We study the convergence of stochastic gradient descent (SGD) for non-convex objective functions. We establish the local convergence with positive probability under the local Łojasiewicz condition introduced by Chatterjee in \cite{chatterjee2022convergence} and an additional local structural assumption of the loss function landscape. A key component of our proof is to ensure that the whole trajectories of SGD stay inside the local region with a positive probability. We also provide examples of neural networks with finite widths such that our assumptions hold.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.2304.09221
- OpenAlex
- W4366549661
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.