Pith. sign in

REVIEW 4 major objections 6 minor 19 references

Within-class variance in language models is stored context, not unfinished neural collapse, and it obeys an information floor.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 02:40 UTC pith:KRPIWLB5

load-bearing objection Serious, carefully scoped paper: binary floor and type-count reduction are real; the identity-tracking law is strong empirics that the floor does not yet force. the 4 major comments →

arxiv 2607.09487 v1 pith:KRPIWLB5 submitted 2026-07-10 cs.LG cs.CLstat.ML

Neural Collapse Is Forbidden: Information Floors in Language Models

classification cs.LG cs.CLstat.ML
keywords neural collapselanguage modelsinformation floorwithin-class varianceconditional mutual informationweight decayrepresentation geometrypretraining dynamics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Language-model representations keep a large within-class variance that classification theory would call incomplete neural collapse. This paper argues the residual is not noise but allocated storage for context, and that the allocation obeys a law. Across fourteen models spanning a hundredfold parameter range, dimensionless variance shares show that within-token context carries most of the mass (roughly 79–91 percent) while coarse category structure is only a thin slice (4–12 percent). Theory shows that token-level weight decay turns next-token prediction into an imbalanced category problem ordered by type count, and a proved binary floor forces within-category dispersion to be at least proportional to the conditional mutual information between token and context given the category. Empirically, the spread of token identity means inside each category tracks that information in every model and partition tested—even when one model’s information is used to predict another model’s dispersion—while total variance does not. Over pretraining the category share rises, overshoots, decays, and partially recovers, because the information that category structure must carry never left.

Core claim

Residual within-class variance in language models is allocated information storage, not unfinished collapse. A centering identity voids mean-cosine simplex-ETF claims; a reduction shows token-level weight decay penalizes categories by type count and yields type-count-ordered norms; and a binary information floor forces within-category dispersion to scale with I(token; context | category). The component that tracks the information is identity dispersion of token means inside a category, not total variance, across every tested model and partition.

What carries the argument

The information floor (Proposition 4, binary case): within-category feature dispersion is at least proportional to the model-realized conditional mutual information I(token; context | category). It is proved by chaining a chi-square bound on binary KL, the Lipschitz constant of the logistic, and Cauchy–Schwarz on the readout direction, and it explains why collapse cannot finish when within-category choice is context-dependent.

Load-bearing premise

The reduction of next-token prediction to a size-weighted category problem assumes that within-category preferences do not vary with context—an idealization the paper itself shows is false in trained models and prices with the floor.

What would settle it

Find a model and partition in which within-category identity dispersion of token means fails to track model-realized or corpus-count conditional information after controlling for marginal entropy, mass, and type count, or a binary within-category pair whose readout-margin variance falls below the proved floor 16 q̄(1−q̄) I / R².

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper argues that residual within-class variance in language-model representations is allocated information storage rather than unfinished neural collapse, and that the allocation obeys a law. Methodologically it introduces a dimensionless three-level variance decomposition (within-token context, token identity within category, between category) and a centering identity that voids mean-cosine simplex-ETF evidence. Across 14 models, context carries 79–91% of variance and macro-category structure only 4–12%. Theoretically, token-level weight decay reduces next-token prediction to a type-count-weighted imbalanced K-class problem (Theorem 1; exact K=2 solution and closed-form twisted SELI geometry for general K), and a proved binary information floor (Proposition 4) lower-bounds within-category dispersion by conditional mutual information I(token; context | category). Empirically, identity dispersion—not total variance—tracks model-realized conditional information in every tested model–partition cell, including under model-free and cross-model checks; over pretraining, category share overshoots, decays, and partially recovers on Pythia and OLMo-2.

Significance. If the results hold, the paper reframes a central geometric narrative in representation learning for language models: residual within-class variance is not incomplete collapse but a necessary storage cost of context-dependent choice. Strengths that raise the bar for the field include: complete proofs for the centering identity, exact loss/decay split, K=2 solution, and general-K twisted SELI (Theorem 2) with a closed-form sign condition; explicit tagging of proved vs conjectured claims; pre-registered OLMo-2 dynamics criteria; extensive negative controls (total variance, marginal entropy, frequency-matched nulls, Gaussian surrogates); model-free and cross-model tests that break self-reference in the information estimator; and an honest audit of the authors’ own earlier ETF and dimensional-metric claims. The empirical law (identity dispersion tracks Îk) is unusually thorough for this literature. These contributions would matter for NC theory, LM representation geometry, and training-dynamics interpretation even if some packaging claims are scoped more carefully.

major comments (4)
  1. [Abstract; §3.4 Prop. 4 / Remark 2; §4] Title/abstract claim vs Proposition 4 and Remark 2: Proposition 4 lower-bounds total within-category dispersion Dk = tr(Σ) (and, in direct checks, readout-margin variance) by a multiple of Ik in the binary case. Remark 2 states explicitly that the floor “does not by itself say which component of that dispersion carries the requirement,” and Section 4 finds that total variance Vk does not track Îk while identity dispersion Bk does. The storage narrative and the abstract’s “identity dispersion, not total variance, tracks this information” therefore rest on an empirical localization motivated by head-row/feature-mean co-movement (Section 5), not on a corollary of the proved inequality. Until a component-level bound exists (flagged as “the natural next theorem”), the abstract and title should not present the identity law as following from the floor. Please separate (i) “full collapse is forb
  2. [§3.2 Assumption 1, Theorem 1, Proposition 2; §5] Theorem 1 and the geometry pillar rest on Assumption 1 (context-free within-category readout δc⊤ hi = βc), which the paper states is quantitatively false in trained models and treats the floor as the price of its failure. Proposition 2 exhibits a coupling term C and estimates it as 10−6–10−3 of residual-norm budget, but the between-category geometry claims (type-count ordering of norms, twisted SELI) still inherit the reduced problem as the operative skeleton. The main text should state more sharply what survives without Assumption 1: which ordering predictions are robust to nonzero C, and which are only first-order under the idealization. A short sensitivity statement or explicit “geometry under the reduced problem” scoping in §3 and §5 would make the load-bearing dependence transparent rather than deferred to Appendix B.2.
  3. [Abstract; Figure 1(c); §6] Dynamics recovery is presented as the expected ending under the floor, but is not universal: Pythia-1B declines after its peak (to 0.046), Pythia-2.8B is excluded pending re-extraction, and OLMo-2 fails pre-registered R4 (CDNV U-shape ratio 1.21 < 1.3 bar). The token-axis alignment of the minimum and the partial recovery in most sizes are interesting, but the causal gloss (“because the information it must carry never left”) is consistency, not demonstration—as the paper itself notes. Please either (a) demote recovery from a headline law-like claim to a documented non-monotonic pattern with named exceptions, or (b) add a quantitative link (e.g., identity-share or per-category Îk trajectories) that predicts which sizes recover. As written, Figure 1(c) and the abstract oversell uniformity relative to §6.
  4. [§5; Proposition 3; Theorem 2] Section 5 attributes type-count ordering of category norms to the decay/offset channel of Proposition 3, with mass as a secondary/frequency-inherited confound controlled by frequency-matched shuffles. The head-row test (24/24 cells, type-count partial negative, mass partial unstable) is the cleanest evidence and should be foregrounded earlier. The centroid-frame residual beyond the null is modest (Stouffer z ≈ −2.9; composition explains ~70% of the coupling), and no decay-ablated training run separates AdamW adaptive scaling from explicit λ|Sk| penalization—limitations the paper notes. For the geometry pillar to carry equal weight with the information law, either strengthen identification (decay ablation or functional-form fit of Theorem 2’s Z̄*) or scope §5 more clearly as “ordering consistent with first-order theory, residual mass channel controlled but not fully identified.”
minor comments (6)
  1. [Abstract; Appendix B.5] Conjecture 1 (general-K floor) is appropriately labeled, but the abstract’s “converse floor, proved for binary categories” could briefly note that quantitative tests use the proved binary case on within-category pairs (286 pairs), so readers do not infer a multi-class theorem from the main empirical program.
  2. [Table 1; Figure 2] Table 1 and Figure 2: GPT-2 XL’s between-category share (~28%) is repeatedly flagged as degenerate; consider a one-line note in the table caption on the near-collinear-centroid diagnosis so the outlier is self-explanatory.
  3. [§3.1; Appendix B.1] Lemma 1 is excellent; the weighted version (Lemma 2) is in the appendix. A forward pointer in §3.1 when empirical sections use mass-weighted grand means would help readers who skip appendices.
  4. [§4] Pooled partial correlations (r = 0.755, etc.) are accompanied by appropriate caveats on non-independence; the permutation and cross-model checks are the right response. Consider putting the cross-model POS result (one model’s I predicts another’s Bk) in the main figure or a small panel, since it is among the strongest anti-circularity controls.
  5. [§6; §7] Related work: Zhao et al. and Li et al. are engaged fairly; a sentence on how the three-level shares map onto Li et al.’s spectral phases (even if deferred quantitatively) would help readers of both papers.
  6. [Figure 3; notation throughout] Typos/clarity: “voids a family of simplex equiangular-tight-frame claims” is clear; ensure consistent notation for Îk vs Ik and Bk vs identity share Bk/Vk across figures and tables. Figure 3 log-scale y-axis ranges differ by model—state that explicitly in the caption.

Circularity Check

1 steps flagged

No load-bearing circularity: floor is independent math; identity law is empirical and cross-checked; self-reference of model-realized I is broken by model-free and cross-model controls.

specific steps
  1. other [Section 4, Estimating per-category conditional information; Robustness and scope]
    "The model-realized conditional information is Îk = H(¯qk) − Eh H(q(·|Sk,h)) ... The estimator remains model-based by design in the main analysis (the floor concerns information the model actually realizes); these checks show the correlation is not an artifact of that choice."

    Mild residual self-reference risk only: Îk is read from the same model’s head whose hidden-state identity dispersion Bk is correlated. Not circular by construction—the paper’s own placebos (Vk uncorrelated; marginal entropy weaker) and externalizing checks (corpus-count Î, cross-model POS prediction, permutations) show the association is not forced. Flagged at score 1 as design proximity, not a definitional collapse of prediction into input.

full rationale

Walked the chain: (i) Lemma 1 is a pure centering identity, not a fitted claim; (ii) Theorem 1 / Prop. 3 / Theorem 2 are closed-form reductions under stated assumptions (Assumption 1 flagged as false in trained models; floor prices the failure); (iii) Prop. 4 is a proved binary inequality from KL/Lipschitz/Cauchy–Schwarz, independent of the empirical law; (iv) Remark 2 explicitly refuses to derive the identity-channel localization from the total-dispersion floor, so the title narrative glues parallel supports rather than equating them by construction; (v) the empirical law (Bk tracks Îk) is not forced—total variance Vk and marginal entropy are negative controls with mixed/weaker signs, and the paper reports model-free corpus-count Î (partial r≈0.68), cross-model POS transfer (median Spearman 0.76), and within-model permutations (p<5e-5). Self-citation of the authors’ earlier version is used only to retire mean-cosine ETF and dimensional-CDNV claims, not as a uniqueness premise. One residual mild concern: main Îk is model-realized by design (same head that shapes geometry), which can look self-referential until the controls are applied—hence score 1 rather than 0—but that is not a by-construction reduction of prediction to input. No fitted-parameter-as-prediction, no uniqueness imported from overlapping authors, no ansatz smuggled via self-citation.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

The central law and floor rest on standard information inequalities and the UFM-style regularized next-token objective, plus one idealization (context-free within-category readout) that the floor itself prices when it fails. Free choices are mostly analysis hyperparameters (K, coverage thresholds, layer L−1) that the paper sweeps or holds fixed across models. Invented constructs are measurement instruments (three-level variance shares, identity dispersion Bk) and the twisted-SELI geometry of the reduced problem, not new physical entities. The general-K floor is explicitly conjectural.

free parameters (5)
  • Illustrative partition K (k-means) = 10 (illustrative); sweeps 10/20/50
    Primary tables use K=10; results are swept at K∈{10,20,50} and POS, but K remains an analyst choice that defines categories for shares and the law.
  • Coverage threshold n_c ≥ 50 = 50 occurrences
    Tokens need ≥50 next-token occurrences to enter class statistics; affects which types enter validation (~500–560) vs train (~24k–26k) regimes.
  • Category position filter (≥200 realized positions) = 200
    Categories with fewer than 200 realized positions are excluded from the information-law correlations.
  • Weight-decay coefficient λ in UFM analysis
    Theory uses λ in the regularized UFM; finite-λ approach rate 1/log(1/λ) is discussed; not fitted to the empirical law but shapes the geometric limit.
  • OLMo-2 CDNV recovery threshold (R4 bar 1.3) = 1.3
    Pre-registered pass/fail bar for cluster-level CDNV U-shape; R4 fails at 1.21 < 1.3 by construction of the bar.
axioms (5)
  • domain assumption Unconstrained features model (UFM) with token-level weight decay on head rows and features is an adequate skeleton for between-category geometry of next-token prediction.
    Section 3 setup and Theorems 1–2; standard in NC theory but not identical to full transformer training dynamics.
  • ad hoc to paper Assumption 1: context-free within-category readout δ_c^⊤ h_i = β_c for all training contexts (idealization for the reduction).
    Stated in §3.2; paper admits it is quantitatively false and uses Prop. 4 as the cost of failure; still load-bearing for the exact reduction skeleton.
  • standard math Binary KL ≤ chi-square bound, 1/4-Lipschitz sigmoid, and Cauchy–Schwarz on readout variance for the information floor.
    Proof of Proposition 4 chains standard inequalities; constants are deliberately loose.
  • domain assumption Empirical within-category conditionals may be treated as uniform for the exact reduction (Theorem 1); nonuniform case adds coupling C estimated small.
    Theorem 1 and Proposition 2; C estimated 10^{-6}–10^{-3} of residual-norm budget but not sharply controlled analytically.
  • ad hoc to paper General-K information floor holds with constants degrading polynomially in category size (Conjecture 1).
    Appendix B.5; only binary case is proved; quantitative program uses pairwise binary tests.
invented entities (3)
  • Three-level dimensionless variance shares (within-token context, token identity within category, between category) independent evidence
    purpose: Artifact-resistant allocation measurement replacing dimensional CDNV trends and mean-cosine ETF scores.
    Defined via law of total variance in §2; instrument rather than physical entity; independent of any single model once class means are fixed.
  • Identity dispersion B_k as the information-carrying component independent evidence
    purpose: Localize the floor’s requirement to spread of token means within a category rather than total within-category variance.
    Remark 2 and §4; empirically tracks I while total variance does not; component-level bound not proved.
  • Twisted SELI geometry (oblique projection Z̄* = I − 1 a^⊤/tr(A)) no independent evidence
    purpose: Closed-form optimum of the type-count-weighted reduced K-class problem under heterogeneous decay and offsets.
    Theorem 2; extends SELI to per-class decay; numerical certificates shipped; not an extra particle but a new geometric object for this objective.

pith-pipeline@v1.1.0-grok45 · 29262 in / 4236 out tokens · 47492 ms · 2026-07-13T02:40:01.808791+00:00 · methodology

0 comments
read the original abstract

Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a family of simplex equiangular-tight-frame claims, including our own earlier ones; in dimensionless variance shares across 14 models, macro-category structure carries only 4-12% of representational variance and within-token context carries 79-91%, stable across a 100x parameter range. On the theory side, token-level weight decay penalizes a category in proportion to its type count, not its occurrence mass, reducing next-token prediction to an imbalanced K-class problem whose optimum orders category norms by type count. A converse floor, proved for binary categories, forces within-category dispersion to be at least proportional to the conditional mutual information I(token; context | category). The law holds: identity dispersion, not total variance, tracks this information across every tested model and partition, under a model-free estimate and even across models, where one model's information predicts another's dispersion; and over pretraining the category share overshoots, decays, and partially recovers, because the information it must carry never left.

Figures

Figures reproduced from arXiv: 2607.09487 by Bruno Abrahao.

Figure 1
Figure 1. Figure 1: (a) Across 13 models, variance is dominated by within-token context; macro-category structure is a thin slice, stable across a 100x parameter range. (b) Within-category identity dispersion tracks model-realized conditional information ˆI(c; ctx | Sk) (pooled partial r = 0.755), the dispersion channel the information floor of Section 3 makes necessary. (c) Over pretraining the between-category share oversho… view at source ↗
Figure 2
Figure 2. Figure 2: The allocation of representational variance (count-weighted shares, [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Identity dispersion Bk (log scale) against model-realized conditional information ˆIk, per category, for four models and three partitions. The relationship is positive in every model-partition combination tested (16/16 including K=10, not shown for clarity); pooled within-model-rank partial correlation r = 0.755 (p = 3 × 10−30 , n = 157 categories at K=50) controlling for marginal entropy, category mass, a… view at source ↗
Figure 4
Figure 4. Figure 4: Left: per-model Spearman correlation of centered centroid norm with log category mass [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Variance-share trajectories under fixed final-checkpoint partitions (count-weighted, common [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

19 extracted references · 1 canonical work pages · 1 internal anchor

  1. [1]

    Neural networks learn statistics of increasing complexity

    Nora Belrose, Quintin Pope, Lucia Quirke, Alex Mallen, and Xiaoli Fern. Neural networks learn statistics of increasing complexity. InInternational Conference on Machine Learning,

  2. [2]

    Pythia: A suite for analyzing large language models across training and scaling

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pages 2397–2430. PMLR, 2023. doi: 10.485...

  3. [3]

    Cong Fang, Hangfeng He, Qi Long, and Weijie J. Su. Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training.Proceedings of the National Academy of Sciences, 118(43), 2021. doi: 10.1073/pnas.2103091118

  4. [4]

    Representation degeneration problem in training natural language generation models

    Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. Representation degeneration problem in training natural language generation models. InInternational Conference on Learning Representations, 2019

  5. [5]

    Neural collapse for unconstrained feature model under cross-entropy loss with imbalanced data.Journal of Machine Learning Research, 25, 2024

    Wanli Hong and Shuyang Ling. Neural collapse for unconstrained feature model under cross-entropy loss with imbalanced data.Journal of Machine Learning Research, 25, 2024. arXiv:2309.09725

  6. [6]

    Generalized neural collapse for a large number of classes

    Jiachen Jiang, Jinxin Zhou, Peng Wang, Qing Qu, Dustin Mixon, Chong You, and Zhihui Zhu. Generalized neural collapse for a large number of classes. InInternational Conference on Machine Learning. PMLR, 2024. doi: 10.48550/arXiv.2310.05351

  7. [7]

    Richards

    Melody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh, Komal Kumar Teru, Adam Santoro, Guillaume Lajoie, and Blake A. Richards. Tracing the representation geometry of language models from pretraining to post-training.arXiv preprint arXiv:2509.23024, 2025

  8. [9]

    Mixon, Hans Parshall, and Jianzong Pi

    Dustin G. Mixon, Hans Parshall, and Jianzong Pi. Neural collapse with unconstrained features.Sampling Theory, Signal Processing, and Data Analysis, 20(11), 2022. doi: 10.1007/ s43670-022-00027-5

  9. [10]

    2 OLMo 2 Furious.arXiv preprint arXiv:2501.00656, 2025

    OLMo Team, Pete Walsh, Luca Soldaini, Dirk Groeneveld, et al. 2 OLMo 2 Furious.arXiv preprint arXiv:2501.00656, 2025

  10. [11]

    In-context learning and induction heads.Transformer Circuits Thread, 2022

    Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, et al. In-context learning and induction heads.Transformer Circuits Thread, 2022. arXiv:2209.11895

  11. [12]

    Han, and David L

    Vardan Papyan, X.Y. Han, and David L. Donoho. Prevalence of neural collapse during the terminal phase of deep learning training.Proceedings of the National Academy of Sciences, 117 (40):24652–24663, 2020. doi: 10.1073/pnas.2015509117

  12. [13]

    The geometry of categorical and hierarchical concepts in large language models

    Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch. The geometry of categorical and hierarchical concepts in large language models. InInternational Conference on Learning Representations, 2025. arXiv:2406.01506. 20

  13. [14]

    A universal part-of-speech tagset

    Slav Petrov, Dipanjan Das, and Ryan McDonald. A universal part-of-speech tagset. In Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC), pages 2089–2096, 2012

  14. [15]

    Explaining grokking and information bottleneck through neural collapse emergence.arXiv preprint arXiv:2509.20829, 2025

    Keitaro Sakamoto and Issei Sato. Explaining grokking and information bottleneck through neural collapse emergence.arXiv preprint arXiv:2509.20829, 2025

  15. [16]

    Imbalance trouble: Revisiting neural-collapse geometry

    Christos Thrampoulidis, Ganesh Ramachandra Kini, Vala Vakilian, and Tina Behnia. Imbalance trouble: Revisiting neural-collapse geometry. InAdvances in Neural Information Processing Systems, volume 35, 2022. arXiv:2208.05512

  16. [17]

    Linguistic collapse: Neural collapse in (large) language models

    Robert Wu and Vardan Papyan. Linguistic collapse: Neural collapse in (large) language models. InAdvances in Neural Information Processing Systems, volume 37, 2024. doi: 10.48550/arXiv. 2405.17767

  17. [18]

    Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. Breaking the softmax bottleneck: A high-rank RNN language model. InInternational Conference on Learning Representations, 2018. arXiv:1711.03953

  18. [19]

    Structure before collapse: Transient semantic geometry in next-token prediction.arXiv preprint arXiv:2606.26749, 2026

    Yize Zhao, Isabel Papadimitriou, and Christos Thrampoulidis. Structure before collapse: Transient semantic geometry in next-token prediction.arXiv preprint arXiv:2606.26749, 2026

  19. [20]

    A Geometric Analysis of Neural Collapse with Unconstrained Features

    Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu. A geometric analysis of neural collapse with unconstrained features. InAdvances in Neural Information Processing Systems, volume 34, 2021. doi: 10.48550/arXiv.2105.02375. A Measurement pitfalls in collapse-style analyses of language mod- els This appendix documents four...