Unsupervised Temperature Scaling: An Unsupervised Post-Processing Calibration Method of Deep Networks

Azadeh Sadat Mozafari , Hugo Siqueira Gomes , Wilson Le\~ao , Christian Gagn\'e

Authors on Pith no claims yet

classification 💻 cs.CV cs.LG

keywords calibratecalibrationdeepmodelsscalingtemperatureunsupervisedconfidence

read the original abstract

The great performances of deep learning are undeniable, with impressive results over a wide range of tasks. However, the output confidence of these models is usually not well-calibrated, which can be an issue for applications where confidence on the decisions is central to providing trust and reliability (e.g., autonomous driving or medical diagnosis). For models using softmax at the last layer, Temperature Scaling (TS) is a state-of-the-art calibration method, with low time and memory complexity as well as demonstrated effectiveness. TS relies on a T parameter to rescale and calibrate values of the softmax layer, whose parameter value is computed from a labelled dataset. We are proposing an Unsupervised Temperature Scaling (UTS) approach, which does not depend on labelled samples to calibrate the model, which allows, for example, the use of a part of a test samples to calibrate the pre-trained model before going into inference mode. We provide theoretical justifications for UTS and assess its effectiveness on a wide range of deep models and datasets. We also demonstrate calibration results of UTS on skin lesion detection, a problem where a well-calibrated output can play an important role for accurate decision-making.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

VOLTA: The Surprising Ineffectiveness of Auxiliary Losses for Calibrated Deep Learning
cs.LG 2026-04 unverdicted novelty 5.0

VOLTA, consisting of a deep encoder with learnable prototypes plus cross-entropy and post-hoc temperature scaling, matches or exceeds ten UQ baselines in accuracy, achieves lower expected calibration error, and perfor...