Pith. sign in

REVIEW 12 cited by

Loss Landscape Degeneracy and Stagewise Development in Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02364 v3 pith:BS5J75NG submitted 2024-02-04 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords degeneracylandscapelearninglossdevelopmentbehaviordeepnetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning involves navigating a high-dimensional loss landscape over the neural network parameter space. Over the course of training, complex computational structures form and re-form inside the neural network, leading to shifts in input/output behavior. It is a priority for the science of deep learning to uncover principles governing the development of neural network structure and behavior. Drawing on the framework of singular learning theory, we propose that model development is deeply linked to degeneracy in the local geometry of the loss landscape. We investigate this link by monitoring loss landscape degeneracy throughout training, as quantified by the local learning coefficient, for a transformer language model and an in-context linear regression transformer. We show that training can be divided into distinct periods of change in loss landscape degeneracy, and that these changes in degeneracy coincide with significant changes in the internal computational structure and the input/output behavior of the transformers. This finding provides suggestive evidence that degeneracy and development are linked in transformers, underscoring the potential of a degeneracy-based perspective for understanding modern deep learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Emergent Misalignment Recruits a Pre-existing Persona Subspace

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...

  2. Specialization of softmax attention heads: insights from the high-dimensional single-location model

    cs.LG 2026-03 conditional novelty 7.0 of 10

    In a high-dimensional toy task, multi-head softmax attention first aligns all heads with the mean signal, then sequentially specializes to latent directions; the paper introduces Bayes-softmax, which attains the Bayes...

  3. Influence Dynamics and Stagewise Data Attribution

    cs.LG 2025-10 conditional novelty 7.0 of 10

    Using Bayesian influence functions and singular learning theory, the authors show that a sample's influence on a model varies non-monotonically over training, peaking and flipping sign at phase transitions.

  4. The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Capability formation in small transformers is claimed to obey a driven-nucleation rate law J = Nνσ(c)e^{−βK} − D, read forward as emergence, backward as plasticity loss, and completed as circuit control.

  5. A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Even for wide three-layer nets, hierarchical learning of both layers yields a smaller asymptotic generalization error than single-layer learning with a fixed random kernel, because of singularities.

  6. Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An information-theoretic 'local redundancy' is lower-bounded, via an entropy-cancellation argument, by the expected squared gradient norm on synthetic probe data, and this proxy modestly out-predicts existing plastici...

  7. Embryology of a Language Model

    cs.LG 2025-08 conditional novelty 6.0 of 10

    UMAP projections of susceptibility vectors reveal a reproducible 'body plan' in a small transformer, including a newly identified 'spacing fin' that distinguishes tokens by the number of preceding spaces.

  8. From Global to Local: A Scalable Benchmark for Local Posterior Sampling

    stat.ML 2025-07 conditional novelty 6.0 of 10

    A scalable benchmark using deep linear networks shows RMSProp-preconditioned SGLD most accurately estimates the local learning coefficient, a degeneracy-aware measure of posterior geometry, up to 100M parameters.

  9. Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Representation smoothness can be used to regularize training, stop early without validation labels, and guide active learning combined with parameter-efficient fine-tuning, reducing data and compute.

  10. Model Organisms for Emergent Misalignment

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Emergent misalignment can be induced in small models via a single rank-1 LoRA adapter, and its onset coincides with a phase transition in the adapter's weight direction.

  11. TRACE for Tracking the Emergence of Semantic Representations in Transformers

    cs.CL 2025-05 reject novelty 6.0 of 10

    Using Hessian curvature, intrinsic dimensionality, and linguistic probes on a synthetic frame-semantic corpus, the paper claims a coordinated intersection-based phase transition in small transformers, though the marke...

  12. Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A two-layer transformer solving an in-context meta-learning task acquires skill in three abrupt phases, each corresponding to a distinct attention circuit: bigram, label attention, then chunking plus label attention.

Pith tools