The Infinite Model
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Overparameterized neural networks exhibit a persistent simplicity bias: among the many functions that interpolate the data, the induced function law preferentially selects simpler solutions. We give a unified theory of this bias by identifying the geometric quantity that governs selection at finite width and its explicit limit at infinite width. At finite width, the induced prior on functions is controlled by the minimum parameter cost required to realize a function, yielding an exact selection law and a refined description of relative preference among interpolants. At infinite width, whenever the pushforward prior converges to a Gaussian process, the governing potential admits a closed form determined by the reproducing-kernel geometry of the limiting model. A bridge theorem connects the finite-width and infinite-width theories, showing that infinite-width simplicity bias arises as the limit of finite-width geometry rather than as a separate principle. In the lazy regime, this yields a single equilibrium framework for minimum-norm interpolation, PAC-Bayes complexity, barrier suppression, and scaling behavior. Beyond the fixed-kernel setting, we formulate the kernel-adaptive potential for feature learning, prove its cylindrical GPS description unconditionally, establish the full function-space lift for compact kernel families and for single-hidden-layer networks, and show that the non-lazy trajectory is not closed in kernel - function variables; the remaining architecture-specific task is the coercivity verification for general deep architectures.
Publication details
- DOI
- 10.5281/zenodo.20395889
- OpenAlex
- W7162428191
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.