REVIEW 12 cited by
Loss Landscape Degeneracy and Stagewise Development in Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep learning involves navigating a high-dimensional loss landscape over the neural network parameter space. Over the course of training, complex computational structures form and re-form inside the neural network, leading to shifts in input/output behavior. It is a priority for the science of deep learning to uncover principles governing the development of neural network structure and behavior. Drawing on the framework of singular learning theory, we propose that model development is deeply linked to degeneracy in the local geometry of the loss landscape. We investigate this link by monitoring loss landscape degeneracy throughout training, as quantified by the local learning coefficient, for a transformer language model and an in-context linear regression transformer. We show that training can be divided into distinct periods of change in loss landscape degeneracy, and that these changes in degeneracy coincide with significant changes in the internal computational structure and the input/output behavior of the transformers. This finding provides suggestive evidence that degeneracy and development are linked in transformers, underscoring the potential of a degeneracy-based perspective for understanding modern deep learning.
Forward citations
Cited by 12 Pith papers
-
Emergent Misalignment Recruits a Pre-existing Persona Subspace
Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...
-
Specialization of softmax attention heads: insights from the high-dimensional single-location model
In a high-dimensional toy task, multi-head softmax attention first aligns all heads with the mean signal, then sequentially specializes to latent directions; the paper introduces Bayes-softmax, which attains the Bayes...
-
Influence Dynamics and Stagewise Data Attribution
Using Bayesian influence functions and singular learning theory, the authors show that a sample's influence on a model varies non-monotonically over training, peaking and flipping sign at phase transitions.
-
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
Capability formation in small transformers is claimed to obey a driven-nucleation rate law J = Nνσ(c)e^{−βK} − D, read forward as emergence, backward as plasticity loss, and completed as circuit control.
-
A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks
Even for wide three-layer nets, hierarchical learning of both layers yields a smaller asymptotic generalization error than single-layer learning with a fixed random kernel, because of singularities.
-
Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization
An information-theoretic 'local redundancy' is lower-bounded, via an entropy-cancellation argument, by the expected squared gradient norm on synthetic probe data, and this proxy modestly out-predicts existing plastici...
-
Embryology of a Language Model
UMAP projections of susceptibility vectors reveal a reproducible 'body plan' in a small transformer, including a newly identified 'spacing fin' that distinguishes tokens by the number of preceding spaces.
-
From Global to Local: A Scalable Benchmark for Local Posterior Sampling
A scalable benchmark using deep linear networks shows RMSProp-preconditioned SGLD most accurately estimates the local learning coefficient, a degeneracy-aware measure of posterior geometry, up to 100M parameters.
-
Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis
Representation smoothness can be used to regularize training, stop early without validation labels, and guide active learning combined with parameter-efficient fine-tuning, reducing data and compute.
-
Model Organisms for Emergent Misalignment
Emergent misalignment can be induced in small models via a single rank-1 LoRA adapter, and its onset coincides with a phase transition in the adapter's weight direction.
-
TRACE for Tracking the Emergence of Semantic Representations in Transformers
Using Hessian curvature, intrinsic dimensionality, and linguistic probes on a synthetic frame-semantic corpus, the paper claims a coordinated intersection-based phase transition in small transformers, though the marke...
-
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
A two-layer transformer solving an in-context meta-learning task acquires skill in three abrupt phases, each corresponding to a distinct attention circuit: bigram, label attention, then chunking plus label attention.
Discussion (0). Sign in to comment.