REVIEW 12 cited by
Calibration in Deep Learning: A Survey of the State-of-the-Art
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Calibration in Deep Learning: A Survey of the State-of-the-Art
read the original abstract
Calibrating deep neural models plays an important role in building reliable, robust AI systems in safety-critical applications. Recent work has shown that modern neural networks that possess high predictive capability are poorly calibrated and produce unreliable model predictions. Though deep learning models achieve remarkable performance on various benchmarks, the study of model calibration and reliability is relatively under-explored. Ideal deep models should have not only high predictive performance but also be well calibrated. There have been some recent advances in calibrating deep models. In this survey, we review the state-of-the-art calibration methods and their principles for performing model calibration. First, we start with the definition of model calibration and explain the root causes of model miscalibration. Then we introduce the key metrics that can measure this aspect. It is followed by a summary of calibration methods that we roughly classify into four categories: post-hoc calibration, regularization methods, uncertainty estimation, and composition methods. We also cover recent advancements in calibrating large models, particularly large language models (LLMs). Finally, we discuss some open issues, challenges, and potential directions.
Forward citations
Cited by 12 Pith papers
-
Divide et Calibra: Multiclass Local Calibration via Vector Quantization
Vector quantization induces a structured partition of the representation space for composing heterogeneous multiclass calibration maps via shared codeword-dependent Dirichlet factors.
-
Inducing Artificial Uncertainty in Language Models
Inducing artificial uncertainty on trivial tasks allows training probes that achieve higher calibration on hard data than standard approaches while retaining performance on easy data.
-
Calibeating Prediction-Powered Inference
Post-hoc calibration of miscalibrated black-box predictions on a labeled sample improves efficiency of prediction-powered inference for semisupervised mean estimation.
-
PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation
PUF replaces deterministic 2D-to-3D scene graph fusion with probabilistic node association and Dirichlet evidence accumulation, yielding substantial accuracy gains on 3DSSG and ReplicaSSG at 15ms/frame.
-
The Score Granularity Gap in Black-Box LLM Classification: A Comparative Study of Confidence Constructions
Comparative evaluation of seven confidence constructions across 25 LLM-dataset pairs reveals that verbalized scores provide good ranking but coarse granularity for thresholding, while multi-query aggregation helps wea...
-
Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis
Proposes NDCG-optimized survival models for organ allocation with allocation guarantees and 50-100% empirical NDCG gains on historical data.
-
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
A structured Zoom-then-Diagnose paradigm with uncertainty-aware GRPO rewards improves lesion localization by 39.3% on liver, breast, and thyroid ultrasound VQA datasets by encouraging caution under ambiguity.
-
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
Introduces Zoom-then-Diagnose paradigm and uncertainty-aware reward in GRPO for confidence-aware ultrasound VQA, reporting 39.3% improvement in lesion localization across liver, breast, and thyroid datasets.
-
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
Unsupervised single-generation confidence calibration for reasoning LLMs via offline self-consistency proxy distillation outperforms baselines on math and QA tasks and improves selective prediction.
-
Same Target, Different Basins: Hard vs. Soft Labels for Annotator Distributions
Hard-label delivery via multipass or SLS matches or beats soft-label training on annotator disagreement data when annotations are sparse and leads to flatter minima.
-
SINAPSE: A lightweight deep learning framework for accurate and explainable neutron-$\gamma$ discrimination
SINAPSE uses a dual-branch neural network with a 1D convolutional autoencoder for denoising and a classifier for neutron-gamma discrimination, trained via random augmentations on high-SNR data and validated with SHAP ...
-
VOLTA: The Surprising Ineffectiveness of Auxiliary Losses for Calibrated Deep Learning
VOLTA, consisting of a deep encoder with learnable prototypes plus cross-entropy and post-hoc temperature scaling, matches or exceeds ten UQ baselines in accuracy, achieves lower expected calibration error, and perfor...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.