article

Analysis of the Rate of Convergence of an Over-Parametrized Deep Neural Network Estimate Learned by Gradient Descent

  • IEEE Transactions on Information Theory
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

الاستشهادات
3
المراجع
66
Comments
0
Paper overview

Abstract

Estimation of a regression function from independent and identically distributed random variables is considered. The$L_{2}$error with integration with respect to the design measure is used as an error criterion. Over-parametrized deep neural network estimates are defined which are based on a special network topology, which use a special random initialization and where all the weights are learned by the gradient descent. It is shown that the expected$L_{2}$error of these estimates converges to zero with the rate close to$n^{-1/(1+d)}$in case that the regression function is Hölder smooth with Hölder exponent$p \in [{1/2,1}]$. In case of an interaction model where the regression function is assumed to be a sum of Hölder smooth functions where each of the functions depends only on$d^{*}$of ofdcomponents of the design variable, it is shown that these estimates achieve the corresponding$d^{*}$-dimensional rate of convergence.

Record transparency

Publication details

DOI
10.1109/tit.2025.3541181
OpenAlex
W4407639540
Document type
article
Language
EN
Source
IEEE Transactions on Information Theory
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.