Pith. sign in

REVIEW 2 cited by

Exploring the Impact of Temperature Scaling in Softmax for Classification and Adversarial Robustness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.20604 v1 pith:NORNHE7G submitted 2025-02-28 cs.LG cs.AI

Exploring the Impact of Temperature Scaling in Softmax for Classification and Adversarial Robustness

classification cs.LG cs.AI
keywords temperatureadversarialmodelsoftmaxfunctionlearningscalingtemperatures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The softmax function is a fundamental component in deep learning. This study delves into the often-overlooked parameter within the softmax function, known as "temperature," providing novel insights into the practical and theoretical aspects of temperature scaling for image classification. Our empirical studies, adopting convolutional neural networks and transformers on multiple benchmark datasets, reveal that moderate temperatures generally introduce better overall performance. Through extensive experiments and rigorous theoretical analysis, we explore the role of temperature scaling in model training and unveil that temperature not only influences learning step size but also shapes the model's optimization direction. Moreover, for the first time, we discover a surprising benefit of elevated temperatures: enhanced model robustness against common corruption, natural perturbation, and non-targeted adversarial attacks like Projected Gradient Descent. We extend our discoveries to adversarial training, demonstrating that, compared to the standard softmax function with the default temperature value, higher temperatures have the potential to enhance adversarial training. The insights of this work open new avenues for improving model performance and security in deep learning applications.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Deep Learning Framework for Joint Channel Acquisition and Communication Optimization in Movable Antenna Systems

    cs.IT 2025-08 conditional novelty 6.0

    An end-to-end neural network jointly designs pilots, quantized feedback, movable-antenna positions, and precoding, achieving near-perfect-CSI sum rates with limited feedback in simulated MA downlink systems.

  2. Multi-User Localization via Active Sensing with Electromagnetically Reconfigurable Antennas

    eess.SP 2026-07 conditional novelty 5.0

    A learned LSTM-GNN policy that sequentially reconfigures a shared reconfigurable antenna aperture improves multi-user wireless localization accuracy over fixed-pattern arrays in simulation.