Pith. sign in

REVIEW 5 cited by

Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02984 v1 pith:JHBTHPER submitted 2024-10-03 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningheadsattentionrefinedcoefficientdatadifferentiationlocal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce refined variants of the Local Learning Coefficient (LLC), a measure of model complexity grounded in singular learning theory, to study the development of internal structure in transformer language models during training. By applying these \textit{refined LLCs} (rLLCs) to individual components of a two-layer attention-only transformer, we gain novel insights into the progressive differentiation and specialization of attention heads. Our methodology reveals how attention heads differentiate into distinct functional roles over the course of training, analyzes the types of data these heads specialize to process, and discovers a previously unidentified multigram circuit. These findings demonstrate that rLLCs provide a principled, quantitative toolkit for \textit{developmental interpretability}, which aims to understand models through their evolution across the learning process. More broadly, this work takes a step towards establishing the correspondence between data distributional structure, geometric properties of the loss landscape, learning dynamics, and emergent computational structures in neural networks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Emergent Misalignment Recruits a Pre-existing Persona Subspace

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...

  2. Specialization of softmax attention heads: insights from the high-dimensional single-location model

    cs.LG 2026-03 conditional novelty 7.0 of 10

    In a high-dimensional toy task, multi-head softmax attention first aligns all heads with the mean signal, then sequentially specializes to latent directions; the paper introduces Bayes-softmax, which attains the Bayes...

  3. Embryology of a Language Model

    cs.LG 2025-08 conditional novelty 6.0 of 10

    UMAP projections of susceptibility vectors reveal a reproducible 'body plan' in a small transformer, including a newly identified 'spacing fin' that distinguishes tokens by the number of preceding spaces.

  4. From Global to Local: A Scalable Benchmark for Local Posterior Sampling

    stat.ML 2025-07 conditional novelty 6.0 of 10

    A scalable benchmark using deep linear networks shows RMSProp-preconditioned SGLD most accurately estimates the local learning coefficient, a degeneracy-aware measure of posterior geometry, up to 100M parameters.

  5. Stochastic Parameter Decomposition

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SPD uses stochastic masking and a learned causal importance function to decompose neural network parameters into sparsely active rank-one subcomponents, recovering ground-truth mechanisms in toy models where APD struggled.

Pith tools