preprint Open access

On the convergence of gradient descent for two layer neural networks

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
0
References
14
Comments
0
Paper overview

Abstract

It has been shown that gradient descent can yield the zero training loss in the over-parametrized regime (the width of the neural networks is much larger than the number of data points). In this work, combining the ideas of some existing works, we investigate the gradient descent method for training two-layer neural networks for approximating some target continuous functions. By making use the generic chaining technique from probability theory, we show that gradient descent can yield an exponential convergence rate, while the width of the neural networks needed is independent of the size of the training data. The result also implies some strong approximation ability of the two-layer neural networks without curse of dimensionality.

Record transparency

Publication details

DOI
10.48550/arxiv.1909.13671
OpenAlex
W2976213001
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.