preprint Open access

Eective Regularisation from Loss-Landscape Geometry: A Unied Derivation of Direction-Dependent Grokking Dynamics

  • Zenodo (CERN European Organization for Nuclear Research)
  • European Organization for Nuclear Research
Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

Grokking is the phenomenon in which a neural network memorises its training data and then, after a prolonged delay, suddenly generalises. Two problems remain open: grokking can occur even at weight-decay strength β=0, and a uniform penalty β‖θ‖² somehow produces direction-selective compression. We prove a main theorem that directly handles discrete SGD and non-quadratic loss surfaces. By analysing constrained minimisation of F(θ) = L_data(θ) + β‖θ‖² on the zero-loss manifold M_0 in the Hessian eigenbasis, we show that the effective decay rate γ_k in eigendirection v_k satisfies: −log(1 − η(h_k + 2β))/η − C'_k·ε ≤ γ_k ≤ −log(1 − η(h_k + 2β))/η + C'_k·ε Three corollaries—discrete linear (ε→0), continuous nonlinear (η→0), and continuous linear (η,ε→0)—are derived as special cases, unifying four theoretical levels. This single theorem resolves both open problems and shows that SGD discreteness accelerates grokking. Numerical verification on two-layer MLPs for modular addition (mod 7 + mod 5) confirms the main theorem in 374/374 conditions (100%). Changes from v1: Main theorem extended from continuous×linear (γ_k = h_k + 2β) to discrete×nonlinear. Verification upgraded from quadratic surrogate to actual neural networks.

Record transparency

Publication details

DOI
10.5281/zenodo.18860534
OpenAlex
W7133482351
Document type
preprint
Language
EN
Source
Zenodo (CERN European Organization for Nuclear Research)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.