REVIEW 2 major objections 4 minor 22 references
Post-hoc calibration strength is condition-dependent and metric-specific inside a single dataset, not a fixed property of the method.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 04:51 UTC pith:JADALDMD
load-bearing objection Useful cautionary demo that aggregate calibration numbers can hide condition-level sign flips, but the four strata that carry the whole argument are never defined. the 2 major comments →
Condition-Stratified Robustness Analysis of Post-Hoc Calibration Methods for Probabilistic Classifiers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In-dataset robustness of post-hoc calibration is condition-dependent and metric-specific. Across four pre-defined strata, temperature scaling and isotonic regression never separate significantly on discrimination after Holm correction, yet temperature scaling produces consistently negative Brier differences and slopes closer to one, while isotonic regression shows Brier sign reversals and severely compressed slopes; AUROC differences also grow with condition severity.
What carries the argument
Condition-stratified evaluation of TEMP versus ISO under four locked hypothesis groups (H1–H4) with registry-locked numbers and Holm multiplicity control on the discrimination deltas; the stratification itself is the device that exposes the metric-specific instability hidden by pooling.
Load-bearing premise
The four pre-defined strata truly represent distinct, controlled perturbations of the data-generating process whose contrasts can be read as evidence of robustness or its absence.
What would settle it
Re-running the identical TEMP and ISO pipeline on the same corpus after redefining or collapsing the four strata so that ISO Brier differences no longer reverse sign and TEMP slopes lose their consistent advantage over ISO would falsify the claim that robustness is condition-dependent.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a pre-registered, in-dataset comparison of temperature scaling (TEMP) and isotonic regression (ISO) under four controlled condition strata (C1–C4) drawn from a single corpus. Four hypothesis groups are examined: H1 discrimination deltas with Holm-adjusted multiplicity control, H2 Brier-score differences, H3 calibration slopes, and H4 AUROC differences under best-condition setups. Registry-locked results show small non-significant TEMP–ISO discrimination deltas (Holm p = 0.9895 everywhere), consistently negative TEMP Brier differences versus ISO sign reversals, TEMP slopes nearer unity (0.76–0.95) than ISO (0.14–0.27), and AUROC differences rising from near zero in C1 to +0.0264 in C4. The authors conclude that in-dataset robustness is condition-dependent and metric-specific, while explicitly disclaiming external transportability.
Significance. If the four strata are genuine controlled perturbations, the work supplies a useful methodological corrective: aggregate calibration metrics can mask sign reversals and condition-specific behavior, and the paper demonstrates this with locked registry values, Holm control on H1, and transparent disclosure of incomplete inferential support for H2–H4. The explicit integrity protocol and refusal to over-claim transportability are strengths relative to much of the calibration literature. The practical implication—that calibrator choice should be monitored per operating regime rather than fixed once—is of interest to applied ML reliability. Significance is currently limited by the missing operational definitions of C1–C4; once those are supplied the contribution becomes a clear, scoped empirical demonstration rather than an uninterpretable contrast.
major comments (2)
- [§II, §III.C] §II and §III.C introduce C1–C4 as “controlled perturbations of the data-generating process” and reject pooling because of an ISO Brier sign reversal in C3, yet nowhere does the manuscript define how the strata were constructed, which variables differ between them, the sample sizes or class balances per stratum, or the partitioning procedure. The central claim that “in-dataset robustness is condition-dependent” rests entirely on contrasts across these four labels (Tables I–II, H1–H4). Without operational definitions the observed patterns (TEMP Brier always negative, ISO slopes ≤0.28, AUROC Δ rising from −0.0004 to 0.0264) cannot be interpreted as robustness evidence versus arbitrary or confounded splits. This is a load-bearing omission that must be remedied before the results can support the stated conclusion.
- [§IV.B–D, Tables I–II] H2–H4 are reported only as directional point estimates (Tables I–II) with no confidence intervals, effect sizes, or multiplicity-adjusted tests, while H1 alone receives Holm correction. The Discussion and Conclusion treat the directional TEMP advantages on Brier and slope, and the monotonic AUROC pattern, as substantive findings. Given that the registry is said to lack complete inferential tuples, either the missing statistics should be supplied or the language should be further softened so that only H1 carries inferential weight. As written, the asymmetry between H1 and H2–H4 undercuts the strength of the multi-metric robustness claim.
minor comments (4)
- [Figs. 4–5] Figure captions for the reliability diagrams (Figs. 4 and 5) are labeled “Fig. 4.(1)” and “Fig. 4.(2)”; renumber consistently with the main text references.
- [§III.B, Fig. 1] The pipeline diagram (Fig. 1) is described but the actual figure is not rendered in the supplied text; ensure it appears in the camera-ready version and that the Holm step is visually clear.
- [§II] AUPRC is mentioned in §II as a complementary metric under imbalance but is never reported; either drop the mention or add the corresponding numbers.
- [Abstract, §IV] Minor typographic inconsistencies appear in the abstract and results (e.g., spacing around minus signs and “through C4:−0.0074”); standardize numeric formatting.
Circularity Check
No circularity: empirical head-to-head of standard calibrators on pre-defined strata; metrics are independent of the fitted maps.
full rationale
The paper is a pre-registered empirical comparison of two standard post-hoc methods (temperature scaling and isotonic regression) under four in-dataset conditions. Temperature T and the isotonic map are fitted on held-out calibration partitions and then scored with independent metrics (Brier, calibration slope, AUROC, discrimination deltas). No equation equates a claimed prediction to a fitted free parameter by construction; the reported deltas and slopes are ordinary evaluation outputs, not rearrangements of the fit. Citations are to external foundational work (Guo et al., Zadrozny & Elkan, Brier, etc.) and do not form a self-citation chain that forces the conclusions. The design deliberately withholds external transportability claims and discloses incomplete inferential tuples for H2–H4. The reader’s concern about underspecified strata C1–C4 is a validity/definition issue, not a circularity reduction. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (2)
- temperature T (TEMP)
- isotonic regression stepwise map (ISO)
axioms (4)
- domain assumption Perfect calibration is defined by P(Y=1 | g(f(x))=q) = q for all q in [0,1]
- ad hoc to paper C1–C4 are distinct controlled perturbations of the data-generating process inside one corpus
- standard math Holm step-down procedure controls family-wise error for the four H1 tests
- domain assumption Brier score, logistic calibration slope, and AUROC measure distinct, non-interchangeable aspects of probabilistic quality
read the original abstract
Post-hoc calibration is widely adopted to correct probability estimates from trained classifiers, yet most evaluations report aggregate performance without testing whether that performance holds across distinct operating conditions within a single dataset. We present a pre-registered, condition-stratified robustness analysis comparing temperature scaling (TEMP) and isotonic regression (ISO) across four controlled conditions (C1--C4). Four hypothesis groups are evaluated: discrimination deltas with Holm-corrected multiplicity control (H1), Brier score differences (H2), calibration slope outcomes (H3), and AUROC differences under best-condition setups (H4). TEMP-minus-ISO discrimination deltas remain small across all conditions (-0.0155 to 0.0139), with Holm-adjusted p-values of 0.9895 everywhere. TEMP Brier differences are consistently negative (C1: -0.0002 through C4: -0.0074), while ISO shows sign reversals. TEMP calibration slopes stay closer to unity in every condition (range 0.7597--0.9493) than ISO slopes (0.1364--0.2726). AUROC differences shift from near zero in C1 (-0.0004) to positive in C4 (0.0264). These results establish that in-dataset robustness is condition-dependent and metric-specific. No claim of external transportability is made.
Figures
Reference graph
Works this paper leans on
-
[1]
An introduction to ROC analysis,
T. Fawcett, “An introduction to ROC analysis,”Pattern Recognition Letters, vol. 27, no. 8, pp. 861–874, 2006
2006
-
[2]
The relationship between precision-recall and ROC curves,
J. Davis and M. Goadrich, “The relationship between precision-recall and ROC curves,” inProceedings of the 23rd International Conference on Machine Learning, 2006, pp. 233–240
2006
-
[3]
Predicting good probabilities with supervised learning,
A. Niculescu-Mizil and R. Caruana, “Predicting good probabilities with supervised learning,” inProceedings of the 22nd International Conference on Machine Learning, 2005, pp. 625–632
2005
-
[4]
On calibration of modern neural networks,
C. Guo, G. Pleiss, Y . Sun, and K. Q. Weinberger, “On calibration of modern neural networks,”arXiv preprint arXiv:1706.04599, 2017
Pith/arXiv arXiv 2017
-
[5]
Transforming classifier scores into accurate multiclass probability estimates,
B. Zadrozny and C. Elkan, “Transforming classifier scores into accurate multiclass probability estimates,” inProceedings of the 8th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2002, pp. 694–699
2002
-
[6]
Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift,
Y . Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. V . Dillon, B. Lakshminarayanan, and J. Snoek, “Can you trust your model’s uncertainty? Evaluating predictive uncertainty under dataset shift,”arXiv preprint arXiv:1906.02530, 2019
Pith/arXiv arXiv 1906
-
[7]
Registered reports: Realigning incentives in scientific publishing,
C. D. Chambers, Z. Dienes, R. D. McIntosh, P. Rotshtein, and K. Willmes, “Registered reports: Realigning incentives in scientific publishing,”Cortex, vol. 66, pp. A1–A2, 2015
2015
-
[8]
Measuring calibration in deep learning,
J. Nixon, M. Dusenberry, G. Jerfel, T. Nguyen, J. Liu, L. Zhang, and D. Tran, “Measuring calibration in deep learning,”arXiv preprint arXiv:1904.01685, 2019
Pith/arXiv arXiv 1904
-
[9]
Verification of forecasts expressed in terms of probability,
G. W. Brier, “Verification of forecasts expressed in terms of probability,” Monthly Weather Review, vol. 78, no. 1, pp. 1–3, 1950
1950
-
[10]
A new vector partition of the probability score,
A. H. Murphy, “A new vector partition of the probability score,”Journal of Applied Meteorology, vol. 12, no. 4, pp. 595–600, 1973
1973
-
[11]
The meaning and use of the area under a receiver operating characteristic (ROC) curve,
J. A. Hanley and B. J. McNeil, “The meaning and use of the area under a receiver operating characteristic (ROC) curve,”Radiology, vol. 143, no. 1, pp. 29–36, 1982
1982
-
[12]
The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,
T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,”PLoS ONE, vol. 10, no. 3, p. e0118432, 2015
2015
-
[13]
Revisiting the calibration of modern neural networks,
M. Minderer, J. Djolonga, R. Romijnders, F. Hubis, X. Zhai, N. Houlsby, D. Tran, and M. Lucic, “Revisiting the calibration of modern neural networks,”arXiv preprint arXiv:2106.07998, 2021
Pith/arXiv arXiv 2021
-
[14]
M. Kull, M. Perello-Nieto, M. Kangsepp, T. S. Filho, H. Song, and P. Flach, “Beyond temperature scaling: Obtaining well-calibrated multiclass probabilities with Dirichlet calibration,”arXiv preprint arXiv:1910.12656, 2019
Pith/arXiv arXiv 1910
-
[15]
A note on Platt’s probabilistic outputs for support vector machines,
H.-T. Lin, C.-J. Lin, and R. C. Weng, “A note on Platt’s probabilistic outputs for support vector machines,”Machine Learning, vol. 68, no. 3, pp. 267–276, 2007
2007
-
[16]
Registered reports: A new publishing initiative at Cortex,
C. D. Chambers, “Registered reports: A new publishing initiative at Cortex,”Cortex, vol. 49, no. 3, pp. 609–610, 2013
2013
-
[17]
Evaluating model calibration in classification,
J. Vaicenavicius, D. Widmann, C. Andersson, F. Lindsten, J. Roll, and T. B. Sch ¨on, “Evaluating model calibration in classification,”arXiv preprint arXiv:1902.06977, 2019
Pith/arXiv arXiv 1902
-
[18]
Simple and scalable predictive uncertainty estimation using deep ensembles,
B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,”arXiv preprint arXiv:1612.01474, 2016
Pith/arXiv arXiv 2016
-
[19]
Deep ensembles: A loss landscape perspective,
S. Fort, H. Hu, and B. Lakshminarayanan, “Deep ensembles: A loss landscape perspective,”arXiv preprint arXiv:1912.02757, 2019
Pith/arXiv arXiv 1912
-
[20]
Dropout as a bayesian approximation: Representing model uncertainty in deep learning,
Y . Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,”arXiv preprint arXiv:1506.02142, 2015
Pith/arXiv arXiv 2015
-
[21]
Verified uncertainty calibration,
A. Kumar, P. Liang, and T. Ma, “Verified uncertainty calibration,”arXiv preprint arXiv:1909.10155, 2019
Pith/arXiv arXiv 1909
-
[22]
What uncertainties do we need in bayesian deep learning for computer vision?,
A. Kendall and Y . Gal, “What uncertainties do we need in bayesian deep learning for computer vision?,”arXiv preprint arXiv:1703.04977, 2017
Pith/arXiv arXiv 2017
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.