Analysis of the Rate of Convergence of an Over-Parametrized Deep Neural Network Estimate Learned by Gradient Descent
At a glance
- الاستشهادات
- 3
- المراجع
- 66
- Comments
- 0
Abstract
Estimation of a regression function from independent and identically distributed random variables is considered. The$L_{2}$error with integration with respect to the design measure is used as an error criterion. Over-parametrized deep neural network estimates are defined which are based on a special network topology, which use a special random initialization and where all the weights are learned by the gradient descent. It is shown that the expected$L_{2}$error of these estimates converges to zero with the rate close to$n^{-1/(1+d)}$in case that the regression function is Hölder smooth with Hölder exponent$p \in [{1/2,1}]$. In case of an interaction model where the regression function is assumed to be a sum of Hölder smooth functions where each of the functions depends only on$d^{*}$of ofdcomponents of the design variable, it is shown that these estimates achieve the corresponding$d^{*}$-dimensional rate of convergence.
Publication details
- DOI
- 10.1109/tit.2025.3541181
- OpenAlex
- W4407639540
- Document type
- article
- Language
- EN
- Source
- IEEE Transactions on Information Theory
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.