Pith. sign in

REVIEW 1 cited by

Properties of the ENCE and other MAD-based calibration metrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11905 v1 pith:4TXDEIUO submitted 2023-05-17 cs.LG physics.chem-phphysics.data-anstat.ME

classification cs.LGphysics.chem-phphysics.data-anstat.ME
keywords calibrationencebinscalibratederrornumberdatasetserrors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Expected Normalized Calibration Error (ENCE) is a popular calibration statistic used in Machine Learning to assess the quality of prediction uncertainties for regression problems. Estimation of the ENCE is based on the binning of calibration data. In this short note, I illustrate an annoying property of the ENCE, i.e. its proportionality to the square root of the number of bins for well calibrated or nearly calibrated datasets. A similar behavior affects the calibration error based on the variance of z-scores (ZVE), and in both cases this property is a consequence of the use of a Mean Absolute Deviation (MAD) statistic to estimate calibration errors. Hence, the question arises of which number of bins to choose for a reliable estimation of calibration error statistics. A solution is proposed to infer ENCE and ZVE values that do not depend on the number of bins for datasets assumed to be calibrated, providing simultaneously a statistical calibration test. It is also shown that the ZVE is less sensitive than the ENCE to outstanding errors or uncertainties.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Good Practice Guide for quantifying uncertainties for machine learning models applied to photoplethysmography signals

    cs.LG 2026-07 conditional novelty 3.0 of 10

    A consortium guide that standardizes how to quantify and validate uncertainty for ML models applied to wearable PPG signals, with benchmarks, datasets, and software.

Pith tools