REVIEW 3 major objections 5 minor 143 references
Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Five tests stay reliable for sparse-data logistic regression
desk verdict A broad, transparent benchmark with plausible winners, but the recommended multi-test workflow is missing multiple-testing control. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The simulation template from Hosmer et al. (1997) is the centerpiece: covariates are drawn from uniform, normal, chi-square, and multi-independent distributions; the null distribution of each test is examined under correct specification; and power is assessed under two controlled misspecifications, an omitted quadratic term and an omitted interaction, whose true coefficients are fixed by solving small systems of logit equations based on chosen fixed points. Each test is evaluated by its empirical Type I error at α=0.05 and empirical power over 10,000 replications, with the 'good' tests required to do well on both criteria. The central identity tying the thesis together is the definition of the five best tests working as standardized or calibration-based statistics rather than as raw chi-square counts.
What would settle it
Re-run the same 30 tests with a misspecified link function (e.g., Stukel's generalized logistic with nonzero shape parameters) or with clustered outcomes: if any of the five champion tests shows a grossly inflated Type I error, or if Farrington suddenly acquires power, the empirical ranking is scenario-dependent rather than general.
Extended reading notes
Core claim
The paper establishes an empirical ranking: when logistic regression is fit to sparse data, the tests that best balance size and power are the GiViTI calibration test (internal validation version), McCullagh's conditional-moment standardization of the Pearson statistic, Osius-Rojek's normal approximation, le Cessie-van Houwelingen's smoothed residual score test, and the Stute-Zhu cumulative-residual process test. These five maintain rejection rates close to the nominal 5% under correct specification and detect omitted quadratic and interaction terms that the classical tests miss. The classical Pearson and deviance tests are liberal in sparse settings, Farrington's test has essentially zero power, and the Unreliability (U) index and Spiegelhalter's z-test are insensitive to non-linear misspecification. Consequently, the thesis argues model assessment requires a battery of powerful formal tests together with visual inspection of calibration.
Load-bearing premise
The whole ranking rests on the assumption that the Hosmer et al. simulation scenarios, with IID data, four covariate distributions, and misspecification limited to an omitted quadratic or interaction term, are representative enough of real sparse-data practice that a test winning there will also win in applications.
Editorial extensions
If this is right
- Applied researchers should prefer the GiViTI, McCullagh, Osius-Rojek, le Cessie-van Houwelingen, and Stute-Zhu tests when fitting logistic models with continuous predictors.
- Classical Pearson, deviance, and related chi-square-based tests should not be used alone for sparse data, since they reject good models too often or fail to detect bad ones.
- Calibration plots should be used routinely alongside formal tests, because formal methods alone miss patterns the eye can catch.
- For moderate-to-large samples up to 5,000, the large-sample HL modification behaves similarly to the traditional HL test, so the added complexity is only justified for very large datasets.
Reading between the lines
- The ranking is established within the Hosmer benchmark's IID, low-dimensional scenarios; it may not transfer to clustered, high-dimensional, or link-misspecified settings, where a different set of tests could win.
- A practical battery could be formed from the five best tests, with a decision rule based on agreement across tests rather than any single p-value; the thesis's own results support this but stop short of defining such a rule.
- Test developers now have a concrete benchmark: a new GOF test should beat or match the five best on both Type I error and power across these exact scenarios.
- Because the Unreliability index and Spiegelhalter's z-test are insensitive to non-linear misspecification, they are best used only as secondary checks of average calibration, not as omnibus GOF tests.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This M.Sc. thesis manuscript compares approximately 30 goodness-of-fit and calibration tests for binary logistic regression under sparse data, following the simulation design of Hosmer et al. (1997). For sample sizes 200 to 5,000 and 10,000 replications per scenario, it estimates empirical Type I error under five covariate distributions and empirical power for omitted quadratic and omitted interaction misspecifications. It identifies five tests—GiViTi calibration, McCullagh, Osius–Rojek, le Cessie, and Stute–Zhu—as empirically powerful while maintaining correct Type I error, and it concludes that formal tests should be combined with visual calibration inspection. A real-data application to the Low Birth Weight dataset illustrates that many tests give conflicting conclusions.
Significance. If the empirical ranking is reliable, the paper provides a practically useful resource for applied statisticians choosing goodness-of-fit tests for logistic regression with continuous covariates. The study has notable strengths: the simulation uses an established benchmark, the number of replications is large, the exclusion of non-applicable or degenerate tests is transparent, the appendices include detailed algorithms and code, and the real-data example adds a concrete illustration. However, the paper is an empirical comparison, not a methodological advance, and its central recommendation depends on a statistical property—the behavior of the recommended multi-test workflow—that the simulation never actually evaluates. The external-validity limits of the Hosmer et al. scenario (IID, low-dimensional, two misspecification types) are acknowledged in Section 5.1.1, so those limits are not the main concern.
major comments (3)
- [6.4 / Abstract] The recommendation to use 'a combination of several powerful statistical tests' (Section 6.4) is not supported by the simulation evidence, which evaluates each test at alpha=0.05 individually (Sections 5.2 and 5.8). The abstract's claim that the five tests balance power with 'not raising false alarms on good models' is per-test; if a practitioner treats rejection by any of the five as evidence of poor fit, the chance of at least one false rejection under a correct model is approximately 1-(0.95)^5 = 0.226 under independence, which is far above the nominal 5%. No Bonferroni, Benjamini-Hochberg, or any other familywise control is discussed, and no combined decision rule is defined or validated. Please either specify and validate a multi-test decision rule with multiplicity adjustment, or restrict the recommendation to per-test screening and state that the visual calibration step, rather than the data-driven combination of p-values, is what synthesizes the evidence.
- [5.4 / Figures 5.2–5.4] Rejection rates are reported as raw proportions from 10,000 replications without Monte Carlo standard errors or confidence intervals. For a true rejection rate of 0.05, the binomial standard error is sqrt(0.05*0.95/10000) = 0.0022, so observed rates of 4.8% and 5.5% are statistically indistinguishable. Several passages in Section 5.8 describe tests as showing 'excellent control' or being 'robust' based on small absolute differences in these heatmaps. Reporting Monte Carlo standard errors or Wilson confidence intervals would let the reader judge which differences are meaningful and would make the ranking of 'best' tests more credible.
- [5.7 / 4.3.8 / Table 5.3] The classification of the 'Hosmer Bootstrap' procedure as a liberal test appears inconsistent with the method described in Section 4.3.8, where Lai and Liu's standardized-power procedure is designed to prevent the ordinary Hosmer-Lemeshow test from rejecting large but practically well-fitting models. Table 5.3 records the Lai-Liu decision as a 0/1 p-value, which is not an ordinary p-value, and Section 5.7 then groups this test with Pearson and deviance as liberal. This suggests a possible implementation or interpretation mismatch. Please clarify how the test was implemented, what threshold was used, and whether the exclusion reflects a property of the original method or of the current implementation.
minor comments (5)
- [Abstract and Chapter 5] There are numerous typos and grammatical errors throughout, including 'emperical power', 'thoes', 'assypmtotically', and 'model access requires'. A careful proofread is needed before publication.
- [3.16.3 / 4.2.3] The naming of the Copas unweighted sum-of-squares test is inconsistent: it is called 'Copas USS test with Osius and Rojek normal approximation' in Section 3.16.3 but appears as 'le Cessie-van Houwelingen-Copas-Hosmer unweighted sum of squares test' in the text. Please standardize the terminology.
- [5.3.3 / Table 5.3] Section 5.3.3 states that link-function misspecification is beyond the scope of the research, yet the Stukel score test is listed in Table 5.3 and described in detail in Section 3.13. Clarify whether the Stukel test is included in the main simulation or only in the excluded-test analysis.
- [5.9 / Figures 5.9-5.12] The real-data figures plot p-values for many tests on a single axis, and the range of p-values makes small values hard to read. Consider using a log-scale or separate panels, and include the sample size and number of events in the figure captions.
- [Table 5.2] Table 5.2 reproduces values from Hosmer et al. (1997), but the source is not fully identified in the table footnote. Please add the original source and confirm permission/reproduction rights.
Circularity Check
No circular derivation: the test ranking is an empirical simulation finding with externally imposed truth, not an equivalence between inputs and outputs.
full rationale
The paper contains no derivation in which a predicted quantity is algebraically or definitionally identical to a fitted input. The load-bearing claim—that GiViTI, McCullagh, Osius-Rojek, le Cessie, and Stute-Zhu balance power and Type I error—is read directly from simulation rejection rates under data-generating processes specified in Sections 5.1 through 5.3, where the null and alternative hypotheses are defined by the true model, not by the tests under evaluation. The simulation design is imported from Hosmer et al. (1997), an external benchmark, and the algorithms are implemented from the cited primary literature; no load-bearing step is justified solely by a self-citation. The only self-referential element is that the tests recommended in Section 6.4 are selected from the same simulation that produced the ranking, which is a model-selection and generalizability limitation rather than a circularity: it concerns whether the empirical ranking generalizes to new settings, not whether the ranking is identical to its inputs. Concerns such as multiplicity control for the recommended combination of tests or extrapolation beyond the Hosmer scenarios are correctness and validity risks, not circular-derivation defects. Accordingly, the paper is not circular in the sense of the review schema.
Assumptions & free parameters
free parameters (5)
- Misspecification level J (quadratic scenario) =
0.01 (slight), 0.4 (pronounced)
- Interaction level I =
0.1 (slight), 0.7 (pronounced)
- Bootstrap replications B =
200
- Standard sample size n0 for Lai-Liu power standardization =
min(n, max(100, 0.2*n))
- Number of HL groups G =
10
assumptions (5)
- domain assumption Asymptotic chi-square distributions of the test statistics are valid under the simulated null scenarios where regularity conditions hold
- domain assumption Model-based bootstrap with B=200 approximates the null distribution of Stute-Zhu and projection statistics
- ad hoc to paper The Hosmer et al. (1997) scenarios (IID, low-dimensional, omitted quadratic/interaction) are representative of sparse-data logistic regression problems
- domain assumption Observations are independent and identically distributed under both null and alternative models
- standard math MLE regularity conditions hold, including consistency and asymptotic normality
Cite this review
Pith. "Pith review of Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data." pith.science (2026). https://pith.science/paper/6L4IOWP4
@misc{pith2026260811140,
author = {Pith},
title = {Pith review of: Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/6L4IOWP4}},
note = {Machine review of arXiv:2608.11140}
}
read the original abstract
Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when the data are "sparse" -- a common issue with continuous predictors like age or weight, where the asymptotic distributional assumptions are not satisfied. This thesis studies classical GOF tests for binary logistic regression under both grouped and sparse data, comparing about 30 statistical tests and machine-learning calibration algorithms. These span the classical chi-square and Hosmer-Lemeshow variants, standardized Pearson statistics, covariate-space partitioning, smoothing-based methods, and contemporary calibration machine-learning and bootstrap procedures. At a fixed size, the GiViTI calibration test (2016), McCullagh (1989), Osius-Rojek (1992), le Cessie (1995) and Stute-Zhu (2002) proved empirically powerful, balancing correct identification of bad models (high empirical power) against not raising false alarms on good models (correct empirical Type I error). Relying on formal methods alone is insufficient: visual diagnostics such as calibration plots are a vital exploratory step for detecting model deficiencies that formal tests often overlook. An application to real data (the Low Birth Weight dataset) shows that many of these tests fail to give valid conclusions when exposed to the complexities of actual datasets. The main conclusion is that model assessment requires a combination of several powerful statistical tests alongside careful visual inspection of model calibration.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Hauck, W. W. and Neuhaus, J. M. and Kalbfleisch, J. D. and Anderson, S. , year =. A consequence of omitted covariates when estimating odds ratios , journal =
-
[2]
Cox, D. R. and Snell, E. J. , year =. Analysis of binary data , edition =
-
[3]
, year =
Agresti, A. , year =. Foundations of linear and generalized linear models , publisher =
-
[4]
and Lemeshow, S
Hosmer, D. and Lemeshow, S. , year =. Applied logistic regression , edition =
-
[5]
Nelder, J. A. and Wedderburn, R. W. M. , year =. Generalized linear models , journal =
-
[6]
and Nelder, J
McCullagh, P. and Nelder, J. A. , title =. 1989 , edition =
1989
-
[7]
and Pennell, M
Paul, P. and Pennell, M. L. and Lemeshow, S. , year =. Standardizing the power of the. Statistics in Medicine , volume =
-
[8]
and Hosmer, T
Hosmer, D. and Hosmer, T. and Le Cessie, S. and Lemeshow, S. , year =. A comparison of goodness-of-fit tests for the logistic regression model , journal =
Show all 143 references
-
[9]
Jennings, D. E. , year =. Outliers and residual distributions in logistic regression , journal =
-
[10]
Rao, C. R. , year =. Linear statistical inference and its applications , edition =
-
[11]
Myers, R. H. , year =. Generalized linear models: With applications in engineering and the sciences , edition =
-
[12]
, year =
Agresti, A. , year =. Categorical data analysis , edition =
-
[13]
and Robinson, T
Shang, J. and Robinson, T. J. and Wulff, S. S. , year =. Goodness-of-fit tests in logistic regression with continuous covariates , journal =
-
[14]
Burnham, K. P. and Anderson, D. R. , year =. Multimodel inference: understanding AIC and BIC in model selection , journal =
-
[15]
Dunn, P. K. and Smyth, G. K. , year =. Generalized linear models with examples in R , publisher =
-
[16]
Menard, S. W. , year =. Applied logistic regression analysis , edition =
-
[17]
and Berger, R
Casella, G. and Berger, R. L. , year =. Statistical inference , volume =
-
[18]
Hosmer Jr, D. W. and Lemeshow, S. , year =. Goodness of fit tests for the multiple logistic regression model , journal =
-
[19]
Landwehr, J. M. and Pregibon, D. and Shoemaker, A. C. , year =. Graphical methods for assessing logistic regression models , journal =
-
[20]
and Sabolova, R
Marriott, P. and Sabolova, R. and Van Bever, G. and Critchley, F. , year =. Geometry of goodness-of-fit testing in high dimensional low sample size modelling , booktitle =
-
[21]
SAS/STAT® 15.2 , organization =
-
[22]
Slaughter, S. J. and Delwiche, L. D. , year =. The little
-
[23]
How Data Formats Affect Goodness-of-Fit in Binary Logistic Regression , howpublished =
-
[24]
Ailobhio, D. T. and Ikughur, J. A. , year =. A review of some goodness-of-fit tests for logistic regression model , journal =
-
[25]
and Ding, J
Zhang, J. and Ding, J. and Yang, Y. , year =. A Binary-Regression Adaptive Goodness-of-fit Test (BAGofT) , journal =
-
[26]
and Pennell, M
Nattino, G. and Pennell, M. L. and Lemeshow, S. , title =. Statistics in Medicine , volume =. 2020 , doi =
2020
-
[27]
and van Houwelingen, J
le Cessie, S. and van Houwelingen, J. C. , year =. A goodness-of-fit test for binary regression models, based on smoothing methods , journal =
-
[28]
and van Houwelingen, J
le Cessie, S. and van Houwelingen, J. C. , year =. Testing the fit of a regression model via score tests in random effects models , journal =
-
[29]
Likelihood ratio, Wald, and Lagrange multiplier (score) tests , howpublished =
-
[30]
Kumar, V. V. P. and Duffull, S. B. , year =. Evaluation of graphical diagnostics for assessing goodness of fit of logistic regression models , journal =
-
[31]
and Ohigashi, T
Gosho, M. and Ohigashi, T. and Nagashima, K. and Ito, Y. and Maruo, K. , year =. Bias in odds ratios from logistic regression methods with sparse data sets , journal =
-
[32]
and Lund, R
Bruce, P. and Lund, R. , year =. Fitting and evaluating logistic regression models , journal =
-
[33]
, year =
Pho, K.-H. , year =. Goodness of fit test for a zero-inflated Bernoulli regression model , journal =
-
[34]
and Tran, P.-L
Lee, S.-M. and Tran, P.-L. and Li, C.-S. , year =. Goodness-of-fit tests for a logistic regression model with missing covariates , journal =
-
[35]
, title =
Weesie, J. , title =. Proceedings of the 1st North American Stata Users Group Meeting , year =
-
[36]
and Wunsch, D
Xu, R. and Wunsch, D. , year =. Clustering algorithms in biomedical research: a review , journal =
-
[37]
Xie, X. J. and Pendergast, J. and Clarke, W. , year =. Increasing the power: A practical approach to goodness-of-fit test for logistic regression models with continuous predictors , journal =
-
[38]
and Robinson, T
Pulkstenis, E. and Robinson, T. J. , year =. Two goodness-of-fit tests for logistic regression models with continuous covariates , journal =
-
[39]
and Nelder, J
McCullagh, P. and Nelder, J. A. , year =. Generalized linear models , edition =
-
[40]
and Rojek, D
Osius, G. and Rojek, D. , year =. Normal goodness-of-fit tests for multinomial models with large degrees of freedom , journal =
-
[41]
Stukel, T. A. , year =. Generalized logistic models , journal =
-
[42]
Rady, E. A. and Abonazel, M. R. and Metawe'e, M. H. , year =. A comparison study of goodness of fit tests of logistic regression in R: Simulation and application to breast cancer data , journal =
-
[43]
and Li, X
Liu, H. and Li, X. and Chen, F. and Härdle, W. and Liang, H. , year =. A comprehensive comparison of goodness-of-fit tests for logistic regression models , journal =
-
[44]
, year =
Dardis, C. , year =. LogisticDx: Diagnostic tests and plots for logistic regression models , note =
-
[45]
LDdiag: Link Function and Distribution Diagnostic Test for Social Science Researchers , note =
Yongmei Ni , year =. LDdiag: Link Function and Distribution Diagnostic Test for Social Science Researchers , note =
-
[46]
American Journal of Epidemiology , volume =
A review of goodness of fit statistics for use in the development of logistic regression models , author =. American Journal of Epidemiology , volume =
-
[47]
Journal of the Royal Statistical Society: Series C (Applied Statistics) , volume =
Unweighted sum of squares test for proportions , author =. Journal of the Royal Statistical Society: Series C (Applied Statistics) , volume =
-
[48]
Biometrical Journal , volume =
An improved goodness of fit statistic for probability prediction models , author =. Biometrical Journal , volume =
-
[49]
Canary, J. D. and Blizzard, L. and Barry, R. P. and Hosmer, D. W. and Quinn, S. J. , journal =. A comparison of the
-
[50]
Philosophical Magazine , volume =
On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling , author =. Philosophical Magazine , volume =
-
[51]
Dunn, P. K. and Smyth, G. K. , title =. 2018 , publisher =
2018
-
[52]
Smyth, G. K. , title =. Science and Statistics: A Festschrift for Terry Speed , editor =. 2003 , publisher =
2003
-
[53]
, title =
Pregibon, D. , title =. The Annals of Statistics , volume =
-
[54]
, title =
Pregibon, D. , title =. GLIM82: Proceedings of the International Conference on Generalized Linear Models , editor =. 1982 , publisher =
1982
-
[55]
Journal of the American Medical Informatics Association , volume =
A tutorial on calibration measurements and calibration models for clinical prediction models , author =. Journal of the American Medical Informatics Association , volume =
-
[56]
Circulation: Cardiovascular Quality and Outcomes , volume =
Clinical prediction models for cardiovascular disease: tufts predictive analytics and comparative effectiveness clinical prediction model database , author =. Circulation: Cardiovascular Quality and Outcomes , volume =
-
[57]
and Kronenberger, J
Küppers, F. and Kronenberger, J. and Shantia, A. and Haselhoff, A. , title =. The
-
[58]
Statistics in Medicine , year =
Probabilistic prediction in patient management and clinical trials , author =. Statistics in Medicine , year =
-
[59]
, title =
Molinari, L. , title =. Biometrika , volume =
-
[60]
, title =
Farrington, C.P. , title =. Journal of the Royal Statistical Society, Series B , year =
-
[61]
, title =
Kuss, O. , title =. Statistics in Medicine , year =
-
[62]
and Lemeshow, S
Hosmer, D. and Lemeshow, S. , title =. 2013 , publisher =
2013
-
[63]
Windmeijer, F. A. G. , title =. Statistica Neerlandica , volume =
-
[64]
, title =
Weesie, J. , title =. Stata Technical Bulletin , volume =
-
[65]
and Beyer, J
Kueppers, S. and Beyer, J. and J. Calibration in Machine Learning Models , journal =
-
[66]
and Fenske, N
Mayr, A. and Fenske, N. and Hofner, B. and Kneib, T. and Schmid, M. , title =. Statistical Methods in Medical Research , volume =
-
[67]
Calibration (Statistics) , journal =
-
[68]
Calibration of Machine Learning Models , journal =
-
[69]
Hosmer-Lemeshow Test , journal =
-
[70]
Probability Calibration , journal =
-
[71]
Marius , title =
C.P. Marius , title =. GitHub repository , howpublished =. 2023 , publisher =
2023
-
[72]
, title =
Kuss, O. , title =. 2002 , howpublished =
2002
-
[73]
, title =
Kuss, O. , title =. Proceedings of the Twenty-Seventh Annual SAS Users Group International Conference (SUGI 27) , pages =. 2002 , address =
2002
-
[74]
and Puke, M
Henzi, A. and Puke, M. and Dimitriadis, T. and Ziegel, J. , title =. The New England Journal of Statistics in Data Science , volume =. 2024 , pages =
2024
-
[75]
Journal of the Royal Statistical Society: Series A (Statistics in Society) , volume =
Shafer, Glenn , title =. Journal of the Royal Statistical Society: Series A (Statistics in Society) , volume =
-
[76]
Wasserstein, R. L. and Lazar, N. A. , title =. The American Statistician , volume =
-
[77]
, title =
Harrell, Jr., Frank E. , title =. 2015 , publisher =
2015
-
[78]
Cox, D. R. , title =. Biometrika , volume =
-
[79]
Miller, M. E. and Hui, S. L. and Tierney, W. M. , title =. Statistics in Medicine , volume =
-
[80]
Barlow, R. E. and Bartholomew, D. J. and Bremner, J. M. and Brunk, H. D. , title =. Wiley Series in Probability and Mathematical Statistics , year =
-
[81]
and Caruana, R
Niculescu-Mizil, A. and Caruana, R. , title =. Proceedings of the 22nd International Conference on Machine Learning (ICML) , year =
-
[82]
and Petersen, I
Lindhiem, O. and Petersen, I. T. and Mentch, L. K. and Youngstrom, E. A. , title =. Assessment , volume =. 2018 , publisher =. doi:10.1177/1073191117750465 , url =
2018 doi
-
[83]
and Sun, H
Cabanillas Silva, P. and Sun, H. and Rezk, M. and Roccaro-Waldmeyer, D. and Fliegenschmidt, J. and Hulde, N. and von Dossow, V. and Meesseman, L. and Depraetere, K. and Stieg, J. and Szymanowsky, R. and Dahlweid, F. , title =. J Med Internet Res , volume =. 2024 , url =
2024
-
[84]
and Silva Filho, T
Kull, M. and Silva Filho, T. M. and Flach, P. , title =. Artificial Intelligence and Statistics , volume =. 2017 , url =
2017
-
[85]
and Finazzi, S
Nattino, G. and Finazzi, S. and Bertolini, G. , title =. Statistics in Medicine , volume =. 2014 , doi =
2014
-
[86]
and Finazzi, S
Nattino, G. and Finazzi, S. and Bertolini, G. , title =. Statistics in Medicine , volume =. 2016 , doi =
2016
-
[87]
, title =
Chesher, A. , title =. Economics Letters , volume =
-
[88]
, title =
White, H. , title =. Econometrica , volume =
-
[89]
, title =
Orme, C. , title =. The Manchester School , volume =
-
[90]
and Yang, K
Yang, A. and Yang, K. , title =. The New England Journal of Statistics in Data Science , volume =. 2025 , doi =
2025
-
[91]
Mitigating algorithmic bias through probability calibration: A case study on lead generation data , journal =
Nikoli. Mitigating algorithmic bias through probability calibration: A case study on lead generation data , journal =. 2025 , doi =
2025
-
[92]
and Zhang, H
Xu, J. and Zhang, H. , title =. Statistical Science , volume =. 2019 , doi =
2019
-
[93]
and Hart, J
Ma, Y. and Hart, J. D. and Janicki, R. and Carroll, R. J. , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =. 2011 , doi =
2011
-
[94]
Tsiatis, A. A. , title =. Biometrika , volume =
-
[95]
Newey, W. K. , title =. Econometrica , volume =
-
[96]
Copas, J. B. , title =. Applied Statistics , volume =
-
[97]
and Schemper, M
Mittlbock, M. and Schemper, M. , title =. Statistics in Medicine , volume =
-
[98]
and Zhang, B
Qin, J. and Zhang, B. , title =. Biometrika , volume =
-
[99]
and Bowman, A
Azzalini, A. and Bowman, A. and Hardle, A. W. , title =. Biometrika , volume =
-
[100]
, title =
McCullagh, P. , title =. International Statistical Review , year =
-
[101]
Staniswalis, J. G. and Severini, T. A. , title =. Annals of Statistics , volume =
-
[102]
and Glosup, J
Firth, D. and Glosup, J. and Hinkley, D. V. , title =. Biometrika , volume =
-
[103]
2009 , publisher =
The Elements of Statistical Learning: Data Mining, Inference, and Prediction , author =. 2009 , publisher =
2009
-
[104]
Wood, S. N. , title =. Chapman and Hall/CRC , year =
-
[105]
Cleveland, W. S. , title =. Journal of the American Statistical Association , volume =
-
[106]
Cleveland, W. S. and Grosse, E. and Shyu, W. M. , title =. Statistical Models in S , editor =
-
[107]
Landwehr, J. M. and Pregibon, D. and Shoemaker, A. C. , title =. Journal of the American Statistical Association , volume =
-
[108]
Dunn, P. K. and Smyth, G. K. , title =. Journal of Computational and Graphical Statistics , volume =
-
[109]
Journal of Statistical Computation and Simulation , volume =
Lai, Xin and Liu, Liu , title =. Journal of Statistical Computation and Simulation , volume =. 2018 , doi =
2018
-
[110]
Eubank, R. L. and Spiegelman, C. H. , title =. Biometrika , volume =
-
[111]
, title =
Efron, B. , title =. Annals of Statistics , volume =
-
[112]
and Tibshirani, R
Efron, B. and Tibshirani, R. J. , title =. 1993 , publisher =
1993
-
[113]
and Zhu, L.-X
Stute, W. and Zhu, L.-X. , title =. Scandinavian Journal of Statistics , volume =
-
[114]
and Yang, Y
Bohdal, O. and Yang, Y. and Hospedales, T. , title =. arXiv preprint arXiv:2308.01222 , year =
-
[115]
and Vergis, N
Denaxas, S. and Vergis, N. and Young, A. , title =. BMC Medicine , volume =
-
[116]
Cook, N. R. , title =. Circulation , volume =
-
[117]
Proceedings of the 34th International Conference on Machine Learning-Volume 70 , pages =
On calibration of modern neural networks , author =. Proceedings of the 34th International Conference on Machine Learning-Volume 70 , pages =
-
[118]
BMC medicine , volume =
Calibration: the Achilles heel of predictive analytics , author =. BMC medicine , volume =
-
[119]
Ganaie, M. A. and Hu, M. and Tanveer, M. and Suganthan, P. N. , title =. Engineering Applications of Artificial Intelligence , volume =. 2022 , doi =
2022
-
[120]
and Liu, L
Yang, Y. and Liu, L. and Wang, J. and Li, X. , title =. Artificial Intelligence Review , volume =. 2022 , doi =
2022
-
[121]
arXiv preprint arXiv:2501.10089 , year =
Classifier ensemble for efficient uncertainty calibration of deep neural networks for image classification , author =. arXiv preprint arXiv:2501.10089 , year =
-
[122]
Journal of Biomedical Informatics , volume =
Beyond discrimination: a comparison of calibration methods and clinical usefulness of predictive models of readmission risk , author =. Journal of Biomedical Informatics , volume =. 2017 , doi =
2017
-
[123]
and Bertolini, G
Finazzi, S. and Bertolini, G. , year =. Calibration belt for quality-of-care assessment based on dichotomous outcomes , journal =
-
[124]
and Finazzi, S
Nattino, G. and Finazzi, S. and Bertolini, G. , title =. Statistical Methods in Medical Research , volume =. 2018 , doi =
2018
-
[125]
Brier, G. W. , title =. Monthly Weather Review , volume =. 1950 , doi =
1950
-
[126]
Oza, N. C. , title =. 2004 , note =
2004
-
[127]
, title =
Wickham, H. , title =. Journal of Statistical Software , year =. doi:10.18637/jss.v059.i10 , url =
-
[128]
Moore, D. S. and Spruill, M. C. , title =. The Annals of Statistics , volume =. 1975 , doi =
1975
-
[129]
Archer, K. J. and Lemeshow, S. and Hosmer, D. W. , title =. Computational Statistics & Data Analysis , year =
-
[130]
, title =
Nygaard, E. , title =. 2019 , address =
2019
-
[131]
Hosmer, D. W. and Hjort, N. L. , title =. Statistics in Medicine , year =
-
[132]
Ebrahim, K. E. , title =. 2025 , howpublished =
2025
-
[133]
, title =
Wald, A. , title =. Transactions of the American Mathematical Society , year =
-
[134]
Dawid, A. P. , title =. Journal of the American Statistical Association , year =
-
[135]
Hosmer, D. W. and Lemeshow, S. , title =. Statistics in Medicine , year =
-
[136]
Steyerberg, E. W. and Vickers, A. J. and Cook, N. R. and Gerds, T. and Gonen, M. and Obuchowski, N. and Pencina, M. J. and Kattan, M. W. , title =. Epidemiology , year =
-
[137]
, title =
Serrano, N. , title =. Intensive Care Medicine , volume =. 2012 , doi =
2012
-
[138]
and Caputo, M
Cocomello, L. and Caputo, M. and Cornish, R. , title =. BMJ Open , year =
-
[139]
and Poole, D
Finazzi, S. and Poole, D. and Luciani, D. and Cogo, P. E. and Bertolini, G. , title =. PLoS ONE , year =
-
[140]
E. L. Lehmann and Joseph P. Romano , title =. 2005 , address =
2005
-
[141]
E. L. Lehmann , title =. 1999 , address =
1999
-
[142]
Rossi , title =
Richard J. Rossi , title =. 2018 , address =
2018
-
[143]
The Annals of Statistics , volume =
Vovk, Vladimir and Wang, Ruodu , title =. The Annals of Statistics , volume =. 2021 , doi =
2021
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.