Pith. sign in

REVIEW 2 cited by

Calibration of Neural Networks using Splines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.12800 v2 pith:2PCUGVD4 submitted 2020-06-23 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords calibrationfunctionrecalibrationcumulativedistributionsempiricalerrorexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Calibrating neural networks is of utmost importance when employing them in safety-critical applications where the downstream decision making depends on the predicted probabilities. Measuring calibration error amounts to comparing two empirical distributions. In this work, we introduce a binning-free calibration measure inspired by the classical Kolmogorov-Smirnov (KS) statistical test in which the main idea is to compare the respective cumulative probability distributions. From this, by approximating the empirical cumulative distribution using a differentiable function via splines, we obtain a recalibration function, which maps the network outputs to actual (calibrated) class assignment probabilities. The spine-fitting is performed using a held-out calibration set and the obtained recalibration function is evaluated on an unseen test set. We tested our method against existing calibration approaches on various image classification datasets and our spline-based recalibration approach consistently outperforms existing methods on KS error as well as other commonly used calibration measures. Our Code is available at https://github.com/kartikgupta-at-anu/spline-calibration.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Eigenvalue Calibration for Semantic Embeddings of Large Language Models

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Temperature scaling of density-matrix eigenvalues from LLM semantic embeddings optimizes proper-score calibration and corrects systematic overconfidence so entropy equals risk.

  2. Instance-Wise Monotonic Calibration by Constrained Transformation

    cs.LG 2025-07 reject novelty 6.0 of 10

    MCCT and MCCT-I fit monotone per-rank scale and bias parameters on sorted logits for calibration, but the claimed monotonicity theorem fails for logits with negative values.

Pith tools