preprint Open access

Growing Neural Networks have Flat Optima and Generalize Better

  • HAL (Le Centre pour la Communication Scientifique Directe)
  • Centre National de la Recherche Scientifique
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

In this work, we study the loss landscape of growing neural networks and show that they have flatter minima than when trained with all of their parameters from random initialization. Then, we further evaluate and compare the generalization properties of both growing and non-growing models using, along with standard measures such as the training loss and the validation accuracy, an uncommon approximation of the population risk. The results we find suggest that growing models have better generalization properties. This supports the argument that flatness of the loss positively correlates with generalization in the current debate in the scientific community about flatness. We validate our approach on a wide range of binary Natural Language Processing tasks with large state-of-the-art deep learning models. Our theoretical and experimental results open new perspectives to study these questions through the prism of growing neural networks and risk approximations.

Record transparency

Publication details

OpenAlex
W4402659277
Document type
preprint
Language
EN
Source
HAL (Le Centre pour la Communication Scientifique Directe)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.