Pith. sign in

REVIEW 12 cited by

Calibration in Deep Learning: A Survey of the State-of-the-Art

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01222 v4 pith:IGZYJHKQ submitted 2023-08-02 cs.LG cs.AI

Calibration in Deep Learning: A Survey of the State-of-the-Art

classification cs.LG cs.AI
keywords calibrationmodelsdeepmodelmethodscalibratingrecentcalibrated
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Calibrating deep neural models plays an important role in building reliable, robust AI systems in safety-critical applications. Recent work has shown that modern neural networks that possess high predictive capability are poorly calibrated and produce unreliable model predictions. Though deep learning models achieve remarkable performance on various benchmarks, the study of model calibration and reliability is relatively under-explored. Ideal deep models should have not only high predictive performance but also be well calibrated. There have been some recent advances in calibrating deep models. In this survey, we review the state-of-the-art calibration methods and their principles for performing model calibration. First, we start with the definition of model calibration and explain the root causes of model miscalibration. Then we introduce the key metrics that can measure this aspect. It is followed by a summary of calibration methods that we roughly classify into four categories: post-hoc calibration, regularization methods, uncertainty estimation, and composition methods. We also cover recent advancements in calibrating large models, particularly large language models (LLMs). Finally, we discuss some open issues, challenges, and potential directions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Divide et Calibra: Multiclass Local Calibration via Vector Quantization

    cs.LG 2026-05 unverdicted novelty 7.0

    Vector quantization induces a structured partition of the representation space for composing heterogeneous multiclass calibration maps via shared codeword-dependent Dirichlet factors.

  2. Inducing Artificial Uncertainty in Language Models

    cs.CL 2026-05 unverdicted novelty 7.0

    Inducing artificial uncertainty on trivial tasks allows training probes that achieve higher calibration on hard data than standard approaches while retaining performance on easy data.

  3. Calibeating Prediction-Powered Inference

    stat.ML 2026-04 unverdicted novelty 7.0

    Post-hoc calibration of miscalibrated black-box predictions on a labeled sample improves efficiency of prediction-powered inference for semisupervised mean estimation.

  4. PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation

    cs.CV 2026-07 conditional novelty 6.0

    PUF replaces deterministic 2D-to-3D scene graph fusion with probabilistic node association and Dirichlet evidence accumulation, yielding substantial accuracy gains on 3DSSG and ReplicaSSG at 15ms/frame.

  5. The Score Granularity Gap in Black-Box LLM Classification: A Comparative Study of Confidence Constructions

    cs.CL 2026-06 unverdicted novelty 6.0

    Comparative evaluation of seven confidence constructions across 25 LLM-dataset pairs reveals that verbalized scores provide good ranking but coarse granularity for thresholding, while multi-query aggregation helps wea...

  6. Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis

    cs.LG 2026-06 unverdicted novelty 6.0

    Proposes NDCG-optimized survival models for organ allocation with allocation guarantees and 50-100% empirical NDCG gains on historical data.

  7. Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming

    cs.CV 2026-05 unverdicted novelty 6.0

    A structured Zoom-then-Diagnose paradigm with uncertainty-aware GRPO rewards improves lesion localization by 39.3% on liver, breast, and thyroid ultrasound VQA datasets by encouraging caution under ambiguity.

  8. Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming

    cs.CV 2026-05 unverdicted novelty 6.0

    Introduces Zoom-then-Diagnose paradigm and uncertainty-aware reward in GRPO for confidence-aware ultrasound VQA, reporting 39.3% improvement in lesion localization across liver, breast, and thyroid datasets.

  9. Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation

    cs.LG 2026-04 unverdicted novelty 6.0

    Unsupervised single-generation confidence calibration for reasoning LLMs via offline self-consistency proxy distillation outperforms baselines on math and QA tasks and improves selective prediction.

  10. Same Target, Different Basins: Hard vs. Soft Labels for Annotator Distributions

    cs.LG 2026-05 conditional novelty 5.0

    Hard-label delivery via multipass or SLS matches or beats soft-label training on annotator disagreement data when annotations are sparse and leads to flatter minima.

  11. SINAPSE: A lightweight deep learning framework for accurate and explainable neutron-$\gamma$ discrimination

    physics.ins-det 2026-05 unverdicted novelty 5.0

    SINAPSE uses a dual-branch neural network with a 1D convolutional autoencoder for denoising and a classifier for neutron-gamma discrimination, trained via random augmentations on high-SNR data and validated with SHAP ...

  12. VOLTA: The Surprising Ineffectiveness of Auxiliary Losses for Calibrated Deep Learning

    cs.LG 2026-04 unverdicted novelty 5.0

    VOLTA, consisting of a deep encoder with learnable prototypes plus cross-entropy and post-hoc temperature scaling, matches or exceeds ten UQ baselines in accuracy, achieves lower expected calibration error, and perfor...