preprint وصول مفتوح

Stochastic Training of Residual Networks: a Differential Equation\n Viewpoint

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
6
المراجع
0
Comments
0
Paper overview

Abstract

During the last few years, significant attention has been paid to the\nstochastic training of artificial neural networks, which is known as an\neffective regularization approach that helps improve the generalization\ncapability of trained models. In this work, the method of modified equations is\napplied to show that the residual network and its variants with noise injection\ncan be regarded as weak approximations of stochastic differential equations.\nSuch observations enable us to bridge the stochastic training processes with\nthe optimal control of backward Kolmogorov's equations. This not only offers a\nnovel perspective on the effects of regularization from the loss landscape\nviewpoint but also sheds light on the design of more reliable and efficient\nstochastic training strategies. As an example, we propose a new way to utilize\nBernoulli dropout within the plain residual network architecture and conduct\nexperiments on a real-world image classification task to substantiate our\ntheoretical findings.\n

Record transparency

Publication details

DOI
10.48550/arxiv.1812.00174
OpenAlex
W4289220573
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.