pith. machine review for the scientific record. sign in

arxiv: 2602.16967 · v3 · submitted 2026-02-19 · 💻 cs.LG · cs.AI

Recognition: unknown

Early-Warning Signals of Grokking via Loss-Landscape Geometry

Authors on Pith no claims yet
classification 💻 cs.LG cs.AI
keywords arithmeticgeneralizationgrokkingscancommutatordefectdyckmodular
0
0 comments X
read the original abstract

Grokking -- the abrupt transition from memorization to generalization after prolonged training -- has been linked to confinement on low-dimensional execution manifolds in modular arithmetic. Whether this mechanism extends beyond arithmetic remains open. We study two sequence-learning benchmarks: SCAN compositional generalization and Dyck-1 depth prediction. Across both tasks and a wide range of learning rates, the commutator defect -- a curvature measure derived from non-commuting gradient updates -- rises well before generalization, with lead times following a superlinear power law (alpha approximately 1.18 for SCAN, approximately 1.13 for Dyck), consistent with prior results on modular arithmetic. Weight-space PCA reveals that spectral concentration is not a universal precursor; the commutator defect is. Causal interventions demonstrate a mechanistic role: amplifying non-commutativity accelerates grokking (roughly 32% on SCAN, roughly 50% on Dyck), while suppressing orthogonal gradient flow delays or prevents it. The three task families form a spectrum of causal sensitivity -- modular arithmetic is rigid, Dyck is responsive, SCAN is intermediate -- yet suppression delays or prevents grokking in all cases, establishing necessity as a universal finding. These results identify the commutator defect as a robust, architecture-agnostic, causally implicated early-warning signal for delayed generalization in transformers.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Lifecycle of the Spectral Edge: From Gradient Learning to Weight-Decay Compression

    cs.LG 2026-04 unverdicted novelty 7.0

    The spectral edge transitions from a gradient-driven functional direction before grokking to a perturbation-flat, ablation-critical compression axis at grokking, forming three universality classes predicted by a gap f...

  2. Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training

    cs.LG 2026-03 unverdicted novelty 6.0

    Spectral gaps in the Gram matrix of parameter updates control phase transitions such as grokking in neural network training.