Stochastic Training of Residual Networks: a Differential Equation\n Viewpoint
At a glance
- Citations
- 6
- References
- 0
- Comments
- 0
Abstract
During the last few years, significant attention has been paid to the\nstochastic training of artificial neural networks, which is known as an\neffective regularization approach that helps improve the generalization\ncapability of trained models. In this work, the method of modified equations is\napplied to show that the residual network and its variants with noise injection\ncan be regarded as weak approximations of stochastic differential equations.\nSuch observations enable us to bridge the stochastic training processes with\nthe optimal control of backward Kolmogorov's equations. This not only offers a\nnovel perspective on the effects of regularization from the loss landscape\nviewpoint but also sheds light on the design of more reliable and efficient\nstochastic training strategies. As an example, we propose a new way to utilize\nBernoulli dropout within the plain residual network architecture and conduct\nexperiments on a real-world image classification task to substantiate our\ntheoretical findings.\n
Publication details
- DOI
- 10.48550/arxiv.1812.00174
- OpenAlex
- W4289220573
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.