Embedding effective rank at grokking is a transient that overstates the converged floor by 3–5× (MLP) / 1.3–1.5× (transformer), and compression lags generalization by order T_grok, modulated by LayerNorm.
Topological Signatures of Grokking
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
We study the grokking phenomenon through the lens of topology. Using persistent homology on point clouds derived from the embedding matrices of a range of models trained on modular arithmetic with varying primes, we identify a clear and consistent topological signature of grokking: a sharp increase in both the maximum and total persistence of first homology ($H_1$). Persistence diagrams reveal the emergence of a dominant long-lived topological feature together with increasingly structured secondary features, reflecting the underlying cyclic structure of the task. Compared to existing spectral and geometric diagnostics -- specifically, Fourier analysis and local intrinsic dimension -- persistent homology provides a unified geometric and topological characterization of representation learning, capturing both local and global multi-scale structure. Ablations across data regimes and control settings show that these topological transitions are tied to generalization rather than memorization. Our results suggest that persistent homology offers a principled and interpretable framework for analyzing how neural networks internalize latent structure during training.
fields
cs.LG 3years
2026 3representative citing papers
Persistent homology analysis of LLM activations shows most topological reorganization occurs early in fine-tuning, with a transient peak followed by stabilization and distinct trajectories for different alignment objectives.
Weight decay controls distinct learning regimes in grokking transformers on modular arithmetic, tracked by new cheap attention-based diagnostics with empirical critical value and exponent fits.
citing papers explorer
-
At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics
Embedding effective rank at grokking is a transient that overstates the converged floor by 3–5× (MLP) / 1.3–1.5× (transformer), and compression lags generalization by order T_grok, modulated by LayerNorm.
-
Tracking Representation Dynamics in Large Language Models with Persistent Homology
Persistent homology analysis of LLM activations shows most topological reorganization occurs early in fine-tuning, with a transient peak followed by stabilization and distinct trajectories for different alignment objectives.
-
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
Weight decay controls distinct learning regimes in grokking transformers on modular arithmetic, tracked by new cheap attention-based diagnostics with empirical critical value and exponent fits.