Gradient flow on a two-layer network trained to compose finite-group elements provably pushes each neuron to a single irreducible representation with rank-one cross-layer alignment; for Abelian groups it yields a uniformly diversified, Haar-phase majority-vote predictor.
Provable scaling laws of feature emergence from learning dynamics of grokking,
7 Pith papers cite this work. Polarity classification is still indexing.
years
2026 7representative citing papers
Weight decay controls distinct learning regimes in grokking transformers on modular arithmetic, tracked by new cheap attention-based diagnostics with empirical critical value and exponent fits.
Grokking in linear DNNs is explained as hysteresis in L2 phase transitions where SGD noise enables escape from low-accuracy metastable phases with Arrhenius scaling; the same mechanism is suggested for nonlinear networks.
Experiments on modular arithmetic with heavy label noise show that over-parameterized networks form a distributed internal generalization structure that can be extracted via frequency methods to achieve high accuracy despite 80% noise.
Gradient-based SVD diagnostic uncovers hidden SED-LCH coupling in single and multitask settings and shows rank-3 subspace constraints speed up grokking by 2.3x.
For exponential family data, optimal FC-DNN parameters equal the RG fixed points of the input characteristic parameters, making training equivalent to RG coarse-graining.
Empirical tests confirm robust feature repulsion signs but reveal activation-dependent spectral lock-in in grokking, with x^2 yielding rank-2 updates at epoch ~174 and ReLU remaining rank-1.
citing papers explorer
-
Neural Networks Provably Learn Spectral Representations for Group Composition
Gradient flow on a two-layer network trained to compose finite-group elements provably pushes each neuron to a single irreducible representation with rank-one cross-layer alignment; for Abelian groups it yields a uniformly diversified, Haar-phase majority-vote predictor.
-
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
Weight decay controls distinct learning regimes in grokking transformers on modular arithmetic, tracked by new cheap attention-based diagnostics with empirical critical value and exponent fits.
-
Noise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks
Grokking in linear DNNs is explained as hysteresis in L2 phase transitions where SGD noise enables escape from low-accuracy metastable phases with Arrhenius scaling; the same mechanism is suggested for nonlinear networks.
-
Unveiling Memorization-Generalization Coexistence: A Case Study on Arithmetic Tasks with Label Noise
Experiments on modular arithmetic with heavy label noise show that over-parameterized networks form a distributed internal generalization structure that can be extracted via frequency methods to achieve high accuracy despite 80% noise.
-
Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories
Gradient-based SVD diagnostic uncovers hidden SED-LCH coupling in single and multitask settings and shows rank-3 subspace constraints speed up grokking by 2.3x.
-
Interpreting FCDNNs via RG on Exponential Family
For exponential family data, optimal FC-DNN parameters equal the RG fixed points of the input characteristic parameters, making training equivalent to RG coarse-graining.
-
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
Empirical tests confirm robust feature repulsion signs but reveal activation-dependent spectral lock-in in grokking, with x^2 yielding rank-2 updates at epoch ~174 and ReLU remaining rank-1.