Pith. sign in

REVIEW 6 cited by

Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central Path

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.02073 v4 pith:XRLATOKU submitted 2021-06-03 cs.LG cs.AImath.DGmath.OCstat.ML

classification cs.LGcs.AImath.DGmath.OCstat.ML
keywords lossclassifiercollapsecentraldeepdynamicsneuralcollapsepath
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recently discovered Neural Collapse (NC) phenomenon occurs pervasively in today's deep net training paradigm of driving cross-entropy (CE) loss towards zero. During NC, last-layer features collapse to their class-means, both classifiers and class-means collapse to the same Simplex Equiangular Tight Frame, and classifier behavior collapses to the nearest-class-mean decision rule. Recent works demonstrated that deep nets trained with mean squared error (MSE) loss perform comparably to those trained with CE. As a preliminary, we empirically establish that NC emerges in such MSE-trained deep nets as well through experiments on three canonical networks and five benchmark datasets. We provide, in a Google Colab notebook, PyTorch code for reproducing MSE-NC and CE-NC: at https://colab.research.google.com/github/neuralcollapse/neuralcollapse/blob/main/neuralcollapse.ipynb. The analytically-tractable MSE loss offers more mathematical opportunities than the hard-to-analyze CE loss, inspiring us to leverage MSE loss towards the theoretical investigation of NC. We develop three main contributions: (I) We show a new decomposition of the MSE loss into (A) terms directly interpretable through the lens of NC and which assume the last-layer classifier is exactly the least-squares classifier; and (B) a term capturing the deviation from this least-squares classifier. (II) We exhibit experiments on canonical datasets and networks demonstrating that term-(B) is negligible during training. This motivates us to introduce a new theoretical construct: the central path, where the linear classifier stays MSE-optimal for feature activations throughout the dynamics. (III) By studying renormalized gradient flow along the central path, we derive exact dynamics that predict NC.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MaxSketch: Robust Distinct Counting in Streams via Random Projections

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    MaxSketch achieves O~(log n / ε²) memory for (1+ε)-approximate distinct counting in streams with geometric structure via max-linear random projections.

  2. Residual Feature Integration is Sufficient to Prevent Negative Transfer

    cs.LG 2025-05 unverdicted novelty 7.0 of 10

    Residual feature integration with a trainable target-side encoder provably prevents negative transfer, achieving convergence rates no worse than training from scratch under informative target distributions.

  3. Representation Collapse in Sequential Post-Training of Large Language Models

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Sequential post-training of LLMs induces representation collapse that correlates with reduced plasticity, weaker generalization, and poorer calibration, with lightweight interventions tested to mitigate it.

  4. Geometric Analysis of Neural Regression Collapse via Intrinsic Dimension

    cs.LG 2025-10 unverdicted novelty 5.0 of 10

    Neural regression collapse occurs when last-layer feature intrinsic dimension falls below target intrinsic dimension, creating over-compressed and under-compressed regimes that govern generalization based on data quan...

  5. The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations

    cs.LG 2025-07 conditional novelty 5.0 of 10

    FACT is a first-order stationarity identity for weight matrices that matches or beats the Neural Feature Ansatz as a description of learned features at convergence.

  6. There Will Be a Scientific Theory of Deep Learning

    stat.ML 2026-04 unverdicted novelty 2.0 of 10

    A mechanics of the learning process is emerging in deep learning theory, characterized by dynamics, coarse statistics, and falsifiable predictions across idealized settings, limits, laws, hyperparameters, and universa...

Pith tools