Eective Regularisation from Loss-Landscape Geometry: A Unied Derivation of Direction-Dependent Grokking Dynamics
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
Grokking is the phenomenon in which a neural network memorises its training data and then, after a prolonged delay, suddenly generalises. Two problems remain open: grokking can occur even at weight-decay strength β=0, and a uniform penalty β‖θ‖² somehow produces direction-selective compression. We prove a main theorem that directly handles discrete SGD and non-quadratic loss surfaces. By analysing constrained minimisation of F(θ) = L_data(θ) + β‖θ‖² on the zero-loss manifold M_0 in the Hessian eigenbasis, we show that the effective decay rate γ_k in eigendirection v_k satisfies: −log(1 − η(h_k + 2β))/η − C'_k·ε ≤ γ_k ≤ −log(1 − η(h_k + 2β))/η + C'_k·ε Three corollaries—discrete linear (ε→0), continuous nonlinear (η→0), and continuous linear (η,ε→0)—are derived as special cases, unifying four theoretical levels. This single theorem resolves both open problems and shows that SGD discreteness accelerates grokking. Numerical verification on two-layer MLPs for modular addition (mod 7 + mod 5) confirms the main theorem in 374/374 conditions (100%). Changes from v1: Main theorem extended from continuous×linear (γ_k = h_k + 2β) to discrete×nonlinear. Verification upgraded from quadratic surrogate to actual neural networks.
Publication details
- DOI
- 10.5281/zenodo.18860534
- OpenAlex
- W7133482351
- Document type
- preprint
- Language
- EN
- Source
- Zenodo (CERN European Organization for Nuclear Research)
- Last metadata update
Comments
Log in to join the discussion.