Pith. sign in

REVIEW 1 cited by

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.01562 v1 pith:OAYRY7AP submitted 2025-06-02 cs.LG stat.ML

classification cs.LGstat.ML
keywords softmaxdeepfunctionmodelnetworksrepresentationtemperaturearchitectures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The softmax function is a fundamental building block of deep neural networks, commonly used to define output distributions in classification tasks or attention weights in transformer architectures. Despite its widespread use and proven effectiveness, its influence on learning dynamics and learned representations remains poorly understood, limiting our ability to optimize model behavior. In this paper, we study the pivotal role of the softmax function in shaping the model's representation. We introduce the concept of rank deficit bias - a phenomenon in which softmax-based deep networks find solutions of rank much lower than the number of classes. This bias depends on the softmax function's logits norm, which is implicitly influenced by hyperparameters or directly modified by softmax temperature. Furthermore, we demonstrate how to exploit the softmax dynamics to learn compressed representations or to enhance their performance on out-of-distribution data. We validate our findings across diverse architectures and real-world datasets, highlighting the broad applicability of temperature tuning in improving model performance. Our work provides new insights into the mechanisms of softmax, enabling better control over representation learning in deep neural networks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Training neural networks with an intermediate feature learning strength generalizes best; the paper derives this from a trade-off between over-alignment to the empirical class mean and over-fitting from a large hypoth...

Pith tools