REVIEW 4 major objections 6 minor 19 references
Within-class variance in language models is stored context, not unfinished neural collapse, and it obeys an information floor.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 02:40 UTC pith:KRPIWLB5
load-bearing objection Serious, carefully scoped paper: binary floor and type-count reduction are real; the identity-tracking law is strong empirics that the floor does not yet force. the 4 major comments →
Neural Collapse Is Forbidden: Information Floors in Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Residual within-class variance in language models is allocated information storage, not unfinished collapse. A centering identity voids mean-cosine simplex-ETF claims; a reduction shows token-level weight decay penalizes categories by type count and yields type-count-ordered norms; and a binary information floor forces within-category dispersion to scale with I(token; context | category). The component that tracks the information is identity dispersion of token means inside a category, not total variance, across every tested model and partition.
What carries the argument
The information floor (Proposition 4, binary case): within-category feature dispersion is at least proportional to the model-realized conditional mutual information I(token; context | category). It is proved by chaining a chi-square bound on binary KL, the Lipschitz constant of the logistic, and Cauchy–Schwarz on the readout direction, and it explains why collapse cannot finish when within-category choice is context-dependent.
Load-bearing premise
The reduction of next-token prediction to a size-weighted category problem assumes that within-category preferences do not vary with context—an idealization the paper itself shows is false in trained models and prices with the floor.
What would settle it
Find a model and partition in which within-category identity dispersion of token means fails to track model-realized or corpus-count conditional information after controlling for marginal entropy, mass, and type count, or a binary within-category pair whose readout-margin variance falls below the proved floor 16 q̄(1−q̄) I / R².
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that residual within-class variance in language-model representations is allocated information storage rather than unfinished neural collapse, and that the allocation obeys a law. Methodologically it introduces a dimensionless three-level variance decomposition (within-token context, token identity within category, between category) and a centering identity that voids mean-cosine simplex-ETF evidence. Across 14 models, context carries 79–91% of variance and macro-category structure only 4–12%. Theoretically, token-level weight decay reduces next-token prediction to a type-count-weighted imbalanced K-class problem (Theorem 1; exact K=2 solution and closed-form twisted SELI geometry for general K), and a proved binary information floor (Proposition 4) lower-bounds within-category dispersion by conditional mutual information I(token; context | category). Empirically, identity dispersion—not total variance—tracks model-realized conditional information in every tested model–partition cell, including under model-free and cross-model checks; over pretraining, category share overshoots, decays, and partially recovers on Pythia and OLMo-2.
Significance. If the results hold, the paper reframes a central geometric narrative in representation learning for language models: residual within-class variance is not incomplete collapse but a necessary storage cost of context-dependent choice. Strengths that raise the bar for the field include: complete proofs for the centering identity, exact loss/decay split, K=2 solution, and general-K twisted SELI (Theorem 2) with a closed-form sign condition; explicit tagging of proved vs conjectured claims; pre-registered OLMo-2 dynamics criteria; extensive negative controls (total variance, marginal entropy, frequency-matched nulls, Gaussian surrogates); model-free and cross-model tests that break self-reference in the information estimator; and an honest audit of the authors’ own earlier ETF and dimensional-metric claims. The empirical law (identity dispersion tracks Îk) is unusually thorough for this literature. These contributions would matter for NC theory, LM representation geometry, and training-dynamics interpretation even if some packaging claims are scoped more carefully.
major comments (4)
- [Abstract; §3.4 Prop. 4 / Remark 2; §4] Title/abstract claim vs Proposition 4 and Remark 2: Proposition 4 lower-bounds total within-category dispersion Dk = tr(Σ) (and, in direct checks, readout-margin variance) by a multiple of Ik in the binary case. Remark 2 states explicitly that the floor “does not by itself say which component of that dispersion carries the requirement,” and Section 4 finds that total variance Vk does not track Îk while identity dispersion Bk does. The storage narrative and the abstract’s “identity dispersion, not total variance, tracks this information” therefore rest on an empirical localization motivated by head-row/feature-mean co-movement (Section 5), not on a corollary of the proved inequality. Until a component-level bound exists (flagged as “the natural next theorem”), the abstract and title should not present the identity law as following from the floor. Please separate (i) “full collapse is forb
- [§3.2 Assumption 1, Theorem 1, Proposition 2; §5] Theorem 1 and the geometry pillar rest on Assumption 1 (context-free within-category readout δc⊤ hi = βc), which the paper states is quantitatively false in trained models and treats the floor as the price of its failure. Proposition 2 exhibits a coupling term C and estimates it as 10−6–10−3 of residual-norm budget, but the between-category geometry claims (type-count ordering of norms, twisted SELI) still inherit the reduced problem as the operative skeleton. The main text should state more sharply what survives without Assumption 1: which ordering predictions are robust to nonzero C, and which are only first-order under the idealization. A short sensitivity statement or explicit “geometry under the reduced problem” scoping in §3 and §5 would make the load-bearing dependence transparent rather than deferred to Appendix B.2.
- [Abstract; Figure 1(c); §6] Dynamics recovery is presented as the expected ending under the floor, but is not universal: Pythia-1B declines after its peak (to 0.046), Pythia-2.8B is excluded pending re-extraction, and OLMo-2 fails pre-registered R4 (CDNV U-shape ratio 1.21 < 1.3 bar). The token-axis alignment of the minimum and the partial recovery in most sizes are interesting, but the causal gloss (“because the information it must carry never left”) is consistency, not demonstration—as the paper itself notes. Please either (a) demote recovery from a headline law-like claim to a documented non-monotonic pattern with named exceptions, or (b) add a quantitative link (e.g., identity-share or per-category Îk trajectories) that predicts which sizes recover. As written, Figure 1(c) and the abstract oversell uniformity relative to §6.
- [§5; Proposition 3; Theorem 2] Section 5 attributes type-count ordering of category norms to the decay/offset channel of Proposition 3, with mass as a secondary/frequency-inherited confound controlled by frequency-matched shuffles. The head-row test (24/24 cells, type-count partial negative, mass partial unstable) is the cleanest evidence and should be foregrounded earlier. The centroid-frame residual beyond the null is modest (Stouffer z ≈ −2.9; composition explains ~70% of the coupling), and no decay-ablated training run separates AdamW adaptive scaling from explicit λ|Sk| penalization—limitations the paper notes. For the geometry pillar to carry equal weight with the information law, either strengthen identification (decay ablation or functional-form fit of Theorem 2’s Z̄*) or scope §5 more clearly as “ordering consistent with first-order theory, residual mass channel controlled but not fully identified.”
minor comments (6)
- [Abstract; Appendix B.5] Conjecture 1 (general-K floor) is appropriately labeled, but the abstract’s “converse floor, proved for binary categories” could briefly note that quantitative tests use the proved binary case on within-category pairs (286 pairs), so readers do not infer a multi-class theorem from the main empirical program.
- [Table 1; Figure 2] Table 1 and Figure 2: GPT-2 XL’s between-category share (~28%) is repeatedly flagged as degenerate; consider a one-line note in the table caption on the near-collinear-centroid diagnosis so the outlier is self-explanatory.
- [§3.1; Appendix B.1] Lemma 1 is excellent; the weighted version (Lemma 2) is in the appendix. A forward pointer in §3.1 when empirical sections use mass-weighted grand means would help readers who skip appendices.
- [§4] Pooled partial correlations (r = 0.755, etc.) are accompanied by appropriate caveats on non-independence; the permutation and cross-model checks are the right response. Consider putting the cross-model POS result (one model’s I predicts another’s Bk) in the main figure or a small panel, since it is among the strongest anti-circularity controls.
- [§6; §7] Related work: Zhao et al. and Li et al. are engaged fairly; a sentence on how the three-level shares map onto Li et al.’s spectral phases (even if deferred quantitatively) would help readers of both papers.
- [Figure 3; notation throughout] Typos/clarity: “voids a family of simplex equiangular-tight-frame claims” is clear; ensure consistent notation for Îk vs Ik and Bk vs identity share Bk/Vk across figures and tables. Figure 3 log-scale y-axis ranges differ by model—state that explicitly in the caption.
Circularity Check
No load-bearing circularity: floor is independent math; identity law is empirical and cross-checked; self-reference of model-realized I is broken by model-free and cross-model controls.
specific steps
-
other
[Section 4, Estimating per-category conditional information; Robustness and scope]
"The model-realized conditional information is Îk = H(¯qk) − Eh H(q(·|Sk,h)) ... The estimator remains model-based by design in the main analysis (the floor concerns information the model actually realizes); these checks show the correlation is not an artifact of that choice."
Mild residual self-reference risk only: Îk is read from the same model’s head whose hidden-state identity dispersion Bk is correlated. Not circular by construction—the paper’s own placebos (Vk uncorrelated; marginal entropy weaker) and externalizing checks (corpus-count Î, cross-model POS prediction, permutations) show the association is not forced. Flagged at score 1 as design proximity, not a definitional collapse of prediction into input.
full rationale
Walked the chain: (i) Lemma 1 is a pure centering identity, not a fitted claim; (ii) Theorem 1 / Prop. 3 / Theorem 2 are closed-form reductions under stated assumptions (Assumption 1 flagged as false in trained models; floor prices the failure); (iii) Prop. 4 is a proved binary inequality from KL/Lipschitz/Cauchy–Schwarz, independent of the empirical law; (iv) Remark 2 explicitly refuses to derive the identity-channel localization from the total-dispersion floor, so the title narrative glues parallel supports rather than equating them by construction; (v) the empirical law (Bk tracks Îk) is not forced—total variance Vk and marginal entropy are negative controls with mixed/weaker signs, and the paper reports model-free corpus-count Î (partial r≈0.68), cross-model POS transfer (median Spearman 0.76), and within-model permutations (p<5e-5). Self-citation of the authors’ earlier version is used only to retire mean-cosine ETF and dimensional-CDNV claims, not as a uniqueness premise. One residual mild concern: main Îk is model-realized by design (same head that shapes geometry), which can look self-referential until the controls are applied—hence score 1 rather than 0—but that is not a by-construction reduction of prediction to input. No fitted-parameter-as-prediction, no uniqueness imported from overlapping authors, no ansatz smuggled via self-citation.
Axiom & Free-Parameter Ledger
free parameters (5)
- Illustrative partition K (k-means) =
10 (illustrative); sweeps 10/20/50
- Coverage threshold n_c ≥ 50 =
50 occurrences
- Category position filter (≥200 realized positions) =
200
- Weight-decay coefficient λ in UFM analysis
- OLMo-2 CDNV recovery threshold (R4 bar 1.3) =
1.3
axioms (5)
- domain assumption Unconstrained features model (UFM) with token-level weight decay on head rows and features is an adequate skeleton for between-category geometry of next-token prediction.
- ad hoc to paper Assumption 1: context-free within-category readout δ_c^⊤ h_i = β_c for all training contexts (idealization for the reduction).
- standard math Binary KL ≤ chi-square bound, 1/4-Lipschitz sigmoid, and Cauchy–Schwarz on readout variance for the information floor.
- domain assumption Empirical within-category conditionals may be treated as uniform for the exact reduction (Theorem 1); nonuniform case adds coupling C estimated small.
- ad hoc to paper General-K information floor holds with constants degrading polynomially in category size (Conjecture 1).
invented entities (3)
-
Three-level dimensionless variance shares (within-token context, token identity within category, between category)
independent evidence
-
Identity dispersion B_k as the information-carrying component
independent evidence
-
Twisted SELI geometry (oblique projection Z̄* = I − 1 a^⊤/tr(A))
no independent evidence
read the original abstract
Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a family of simplex equiangular-tight-frame claims, including our own earlier ones; in dimensionless variance shares across 14 models, macro-category structure carries only 4-12% of representational variance and within-token context carries 79-91%, stable across a 100x parameter range. On the theory side, token-level weight decay penalizes a category in proportion to its type count, not its occurrence mass, reducing next-token prediction to an imbalanced K-class problem whose optimum orders category norms by type count. A converse floor, proved for binary categories, forces within-category dispersion to be at least proportional to the conditional mutual information I(token; context | category). The law holds: identity dispersion, not total variance, tracks this information across every tested model and partition, under a model-free estimate and even across models, where one model's information predicts another's dispersion; and over pretraining the category share overshoots, decays, and partially recovers, because the information it must carry never left.
Figures
Reference graph
Works this paper leans on
-
[1]
Neural networks learn statistics of increasing complexity
Nora Belrose, Quintin Pope, Lucia Quirke, Alex Mallen, and Xiaoli Fern. Neural networks learn statistics of increasing complexity. InInternational Conference on Machine Learning,
-
[2]
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pages 2397–2430. PMLR, 2023. doi: 10.485...
Pith/arXiv arXiv 2023
-
[3]
Cong Fang, Hangfeng He, Qi Long, and Weijie J. Su. Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training.Proceedings of the National Academy of Sciences, 118(43), 2021. doi: 10.1073/pnas.2103091118
-
[4]
Representation degeneration problem in training natural language generation models
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. Representation degeneration problem in training natural language generation models. InInternational Conference on Learning Representations, 2019
2019
-
[5]
Wanli Hong and Shuyang Ling. Neural collapse for unconstrained feature model under cross-entropy loss with imbalanced data.Journal of Machine Learning Research, 25, 2024. arXiv:2309.09725
Pith/arXiv arXiv 2024
-
[6]
Generalized neural collapse for a large number of classes
Jiachen Jiang, Jinxin Zhou, Peng Wang, Qing Qu, Dustin Mixon, Chong You, and Zhihui Zhu. Generalized neural collapse for a large number of classes. InInternational Conference on Machine Learning. PMLR, 2024. doi: 10.48550/arXiv.2310.05351
- [7]
-
[9]
Mixon, Hans Parshall, and Jianzong Pi
Dustin G. Mixon, Hans Parshall, and Jianzong Pi. Neural collapse with unconstrained features.Sampling Theory, Signal Processing, and Data Analysis, 20(11), 2022. doi: 10.1007/ s43670-022-00027-5
2022
-
[10]
2 OLMo 2 Furious.arXiv preprint arXiv:2501.00656, 2025
OLMo Team, Pete Walsh, Luca Soldaini, Dirk Groeneveld, et al. 2 OLMo 2 Furious.arXiv preprint arXiv:2501.00656, 2025
Pith/arXiv arXiv 2025
-
[11]
In-context learning and induction heads.Transformer Circuits Thread, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, et al. In-context learning and induction heads.Transformer Circuits Thread, 2022. arXiv:2209.11895
Pith/arXiv arXiv 2022
-
[12]
Vardan Papyan, X.Y. Han, and David L. Donoho. Prevalence of neural collapse during the terminal phase of deep learning training.Proceedings of the National Academy of Sciences, 117 (40):24652–24663, 2020. doi: 10.1073/pnas.2015509117
-
[13]
The geometry of categorical and hierarchical concepts in large language models
Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch. The geometry of categorical and hierarchical concepts in large language models. InInternational Conference on Learning Representations, 2025. arXiv:2406.01506. 20
Pith/arXiv arXiv 2025
-
[14]
A universal part-of-speech tagset
Slav Petrov, Dipanjan Das, and Ryan McDonald. A universal part-of-speech tagset. In Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC), pages 2089–2096, 2012
2089
-
[15]
Keitaro Sakamoto and Issei Sato. Explaining grokking and information bottleneck through neural collapse emergence.arXiv preprint arXiv:2509.20829, 2025
arXiv 2025
-
[16]
Imbalance trouble: Revisiting neural-collapse geometry
Christos Thrampoulidis, Ganesh Ramachandra Kini, Vala Vakilian, and Tina Behnia. Imbalance trouble: Revisiting neural-collapse geometry. InAdvances in Neural Information Processing Systems, volume 35, 2022. arXiv:2208.05512
Pith/arXiv arXiv 2022
-
[17]
Linguistic collapse: Neural collapse in (large) language models
Robert Wu and Vardan Papyan. Linguistic collapse: Neural collapse in (large) language models. InAdvances in Neural Information Processing Systems, volume 37, 2024. doi: 10.48550/arXiv. 2405.17767
-
[18]
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. Breaking the softmax bottleneck: A high-rank RNN language model. InInternational Conference on Learning Representations, 2018. arXiv:1711.03953
Pith/arXiv arXiv 2018
-
[19]
Yize Zhao, Isabel Papadimitriou, and Christos Thrampoulidis. Structure before collapse: Transient semantic geometry in next-token prediction.arXiv preprint arXiv:2606.26749, 2026
Pith/arXiv arXiv 2026
-
[20]
A Geometric Analysis of Neural Collapse with Unconstrained Features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu. A geometric analysis of neural collapse with unconstrained features. InAdvances in Neural Information Processing Systems, volume 34, 2021. doi: 10.48550/arXiv.2105.02375. A Measurement pitfalls in collapse-style analyses of language mod- els This appendix documents four...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2105.02375 2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.