Pith. sign in

REVIEW 2 cited by

A Geometric Analysis of Neural Collapse with Unconstrained Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.02375 v1 pith:FO74R6OU submitted 2021-05-06 cs.LG cs.AIcs.ITmath.ITmath.OCstat.ML

A Geometric Analysis of Neural Collapse with Unconstrained Features

classification cs.LG cs.AIcs.ITmath.ITmath.OCstat.ML
keywords neuralanalysislast-layercollapsefeaturesgloballandscapenetwork
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We provide the first global optimization landscape analysis of $Neural\;Collapse$ -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training. As recently reported by Papyan et al., this phenomenon implies that ($i$) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and ($ii$) cross-example within-class variability of last-layer activations collapses to zero. We study the problem based on a simplified $unconstrained\;feature\;model$, which isolates the topmost layers from the classifier of the neural network. In this context, we show that the classical cross-entropy loss with weight decay has a benign global landscape, in the sense that the only global minimizers are the Simplex ETFs while all other critical points are strict saddles whose Hessian exhibit negative curvature directions. In contrast to existing landscape analysis for deep neural networks which is often disconnected from practice, our analysis of the simplified model not only does it explain what kind of features are learned in the last layer, but it also shows why they can be efficiently optimized in the simplified settings, matching the empirical observations in practical deep network architectures. These findings could have profound implications for optimization, generalization, and robustness of broad interests. For example, our experiments demonstrate that one may set the feature dimension equal to the number of classes and fix the last-layer classifier to be a Simplex ETF for network training, which reduces memory cost by over $20\%$ on ResNet18 without sacrificing the generalization performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Neural Collapse Is Forbidden: Information Floors in Language Models

    cs.LG 2026-07 conditional novelty 8.0

    Within-category identity dispersion in language models tracks conditional mutual information I(token; context|category) and is forced by a proved information floor that forbids full neural collapse.

  2. Test Case Prioritization for DNNs via Neural Collapse Instability

    cs.LG 2026-07 conditional novelty 6.0

    DNN test inputs ranked by prediction instability across late-training checkpoints find faults earlier than confidence-based ranking in most benchmark settings.