Pith. sign in

REVIEW 3 major objections 5 minor 29 references

A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression

T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read This paper shows that adding a single signed correction term to the Hosmer–Lemeshow statistic—weighted by (1−2π̄_g) and referred to a χ²_{G−2} distribution—produces a partition goodness-of-fit test that holds its size, matches Hosmer–Lemesh

desk verdict A small, honest refinement of the Hosmer–Lemeshow test whose simulations back its claims, except for one superlative in the abstract that its own table does not support. read the letter →

arxiv 2607.15454 v1 pith:4M6PIQX3 submitted 2026-07-16 stat.ME stat.AP

classification stat.MEstat.AP MSC 62F0362J12
keywords goodnessoffitHosmer–Lemeshowtestlogisticregressionsparsedatalinkmisspecificationdirectionalcorrectionalignmentfunctionalcalibration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a small, closed-form alteration to the Hosmer–Lemeshow (HL) test can make the standard tool sensitive to a specific and common kind of model failure: a misspecified, asymmetric link function. The alteration subtracts a signed, risk-weighted correction term from HL's Pearson sum, and refers the result to the same χ²_{G−2} distribution. The paper shows in simulations that the new test keeps HL's size exactly, gains no power on symmetric departures or covariate-space structure, and clearly beats HL when the fitted link is the wrong asymmetric shape—most sharply for complementary log–log misfit. A single alignment functional, the inner product of the group-residual bias pattern with the signed correction weight, predicts both where the gain appears and where it reverses, so the test's behavior is explained rather than just observed. A reader should care because HL is the default goodness-of-fit check in applied logistic regression, and this is a drop-in refinement that costs nothing in size, requires no model refit, and adds a directional diagnostic.

What carries the argument

The statistic T_EF = Ĉ_G − C, where Ĉ_G is the Hosmer–Lemeshow Pearson sum over equal-size deciles of fitted risk and C = Σ_g (1−2π̄_g)(o_g−e_g)/V_g is a signed correction with per-group variance V_g = n_g π̄_g(1−π̄_g). The carrying mechanism is the alignment functional A(δ)=Σ_g (1−2π̄_g)δ_g/V_g: a linear functional of the group-residual bias pattern whose odd weight (1−2π̄_g) annihilates symmetric misfit and retains only directional misfit. This functional does the explanatory work, predicting the sign and magnitude of the power difference between EF and HL; the χ²_{G−2} reference is justified by the standard large-sample model-based-grouping argument, with the correction itself vanishing

What would settle it

Simulate the Aranda–Ordaz asymmetric-link setting (α near 0.1) at n=5000, but recompute decile boundaries under the true model and compare the empirical null distribution of T_EF using estimated versus oracle deciles; if the two differ by more than Monte-Carlo error, or if the EF-versus-HL power difference at moderate n fails to follow the sign of A(δ) when grouping is re-estimated, the fixed-grouping assumption breaks.

Watch

Extended reading notes

Core claim

The central claim is that the directional Hosmer–Lemeshow correction T_EF = Ĉ_G − C, with C = Σ_g (1−2π̄_g)(o_g−e_g)/V_g, turns the statistic's sensitivity on and off exactly according to the sign and size of the alignment functional A(δ)=Σ_g (1−2π̄_g)δ_g/V_g. Because the weight (1−2π̄_g) is odd about the midpoint of predicted risk, it cancels any symmetric pattern of group residual bias; only the odd, directional component survives. Under local alternatives, a negatively aligned odd pattern—the signature of an asymmetric link such as complementary log–log—raises the test's non-centrality and gives T_EF more power than HL, while a positively aligned pattern (an omitted quadratic in a skewed

Load-bearing premise

The load-bearing premise is that the decile grouping can be treated as fixed to first order under the null and local alternatives, so that the correction weights w_g/V_g are non-stochastic and the group residuals have the standard model-based covariance; if misspecification shifts the decile boundaries materially, both the χ²_{G−2} calibration and the predicted power gain could change (Section 3.1 and Appendix A).

Editorial extensions

If this is right

  • Practitioners can replace or accompany the HL p-value with T_EF at no extra computational cost: the test uses the same decile grouping, the same reference distribution, and a single closed-form correction, making it a drop-in refinement.
  • For asymmetric-link misfit—common in dose–response, discrete-time survival, and rare-event models—EF is the most sensitive well-calibrated partition test, with the gain visible at moderate n and fading as n grows; the effect is finite-sample, not asymptotic.
  • For symmetric departures such as omitted quadratic terms, EF is less powerful than HL because the alignment functional is positive; the test is not a universal improvement.
  • For covariate-space departures like omitted interactions, EF ties HL and both are beaten by covariate-space grouping tests; no probability-grouping test can detect such structure.
  • The per-group signed contributions {c_g} provide a directional read-out that can localize misfit on the risk scale and suggest trying a complementary-log–log-style link rather than an added polynomial.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the alignment functional depends only on the odd component of the residual-bias pattern and the odd weight, the same correction idea could be ported to other symmetric grouping schemes (equal-width bins, covariate-based partitions) as long as the weight remains odd about the center of the risk scale; the paper does not explore this.
  • The correction is O(n^{-1/2}) under the null, so its advantage is inherently finite-sample; a studentized version that gives the directional term its own rejection region could make the signal first-order and more useful at large n.
  • The calibration-slope appendix implies that in external validation the test is direction-dependent: it helps detect over-shrunk models but is less powerful than HL against overfitted transported models, so the sign of the correction can itself be read as a calibration diagnostic.
  • One testable extension: since the paper shows power tracks |A(δ)| monotonically, one could compute A(δ) under a menu of candidate asymmetric links for a given dataset and use it to decide which link to try after EF rejects, turning the test into a model-selection aid.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes and studies a modification of the Hosmer–Lemeshow (HL) test for sparse logistic regression: T_EF = C_hat_G − C, where C = Σ_g (1−2π̄_g)(o_g−e_g)/V_g is a signed directional correction referred to a χ²_{G−2} distribution. The correction is the grouped form of the Osius–Rojek/Farrington standardization. The authors claim that this correction is exactly inert for symmetric misfit and beneficial for asymmetric-link misfit, that a single 'alignment functional' A(δ) predicts when the correction helps or hurts, and that no well-calibrated partition test is more sensitive to asymmetric-link misspecification. Evidence includes a 24-cell size factorial, power sweeps over omitted quadratic, omitted interaction, and Aranda–Ordaz link families, a seven-test family comparison at n=1000, real-data concordance checks, and an R package with archived simulation code.

Significance. If the claims are supported, the paper would provide a useful, refit-free drop-in companion to HL with a directional diagnostic and an honest map of where the test wins and loses. Strengths include the deliberate reporting of Monte-Carlo error, the explicit 'no-free-lunch' rows, the reproducible R package and Zenodo archive, and the candid limitations section. The main contribution is not a new correction — the authors correctly credit Farrington and Osius–Rojek — but a characterization of the behavior of a grouped version of that correction. The central superlative claim, however, is not supported by the reported precision, and the theoretical 'prediction' is partly an identity; after revision of those claims the paper would be a solid empirical and interpretive contribution.

major comments (3)
  1. [Abstract; §5.2, Table 5; §7.4] The claim that 'no well-calibrated partition test is more sensitive to asymmetric-link misfit' is not supported by Table 5. At K=3000, EF gives 49.7% on cloglog versus HLeqw 48.3%, and 42.2% versus 40.9% on loglog. The Monte-Carlo standard error for the difference of two independent proportions near 0.5 is approximately sqrt(0.5·0.5/3000)·sqrt(2) ≈ 1.3 percentage points, so both margins (1.4 and 1.3) lie within the ±2 SE band of about ±2.6 points. HLeqw is well-calibrated (null size 4.7%), so the data support EF matching the best calibrated partition test, not strictly exceeding it. The text in §5.2 concedes the margin is 'within Monte-Carlo error', but the abstract and §7.4 do not carry this qualification. Please rephrase the claim to 'among the tests compared, EF exceeds the decile-based HL and ties the equal-width HL', or increase K to resolve the comparison. A universal statement ove
  2. [§3.1, Proposition 1; Appendix A, Eq. (4)] Proposition 1(i) is an identity: E(C) = A(δ) follows immediately from the definitions of C and A(δ). The statement that 'A(δ) predicts' where the test gains power is therefore a restatement of the mean shift induced by the correction, not an independent theoretical prediction. Moreover, A(δ) is computed from the true data-generating process (§4.4(c), Table 6), not from data, so it cannot be evaluated by a practitioner without knowing the misspecification. This does not invalidate the simulation results, but the paper should present A(δ) as an interpretive/explanatory device that organizes the simulation outcomes, rather than as a falsifiable prediction independent of the definition. The empirical monotonicity in Table 6 is then evidence of the mechanism, not a test of a separate hypothesis.
  3. [§3.1; Appendix A(iii)] The asymptotic arguments treat the decile grouping as fixed. The null calibration relies on the Moore–Spruill result (cited, not verified) together with the claim that C = b^T Z with b = O(sqrt(G/n)), so C = o_p(1) under fixed grouping. But under the local alternatives used for the power analysis, the fitted probabilities shift at O(n^{−1/2}), and the decile boundaries can shift at the same order as the residual biases δ_g. The first-order argument does not account for this boundary movement. The size simulations and KS checks are reassuring for the null, but the A(δ) computations and the power predictions implicitly assume fixed grouping. Please either supply an argument covering the boundary shift or state explicitly that the theory is first-order under fixed grouping and that the simulation evidence is the primary support.
minor comments (5)
  1. [Table 3; §4.4(b)] Table 3 labels the interaction row as 'continuous-by-continuous product term', but §4.4(b) defines the interaction as ψ x d with d ~ Bernoulli(0.5), i.e. continuous-by-binary. Please correct the table header or the text.
  2. [Table 5; §4.1] The equal-width HL variant HLeqw appears first in Table 5 but is not defined in Table 1 or in §4.1. Add it to the family list or define it in the text before first use.
  3. [Figure 7 caption] The caption says 'the power gain rises monotonically with A(δ) (n 2000)', but the figure plots n = 500, 1000, 2000, and 5000. Please correct the caption.
  4. [§5.3] The representative read-out dataset is selected from among datasets where EF rejects and HL does not. This conditioning should be stated in the text, because the displayed tilt may be stronger than in a typical dataset from the same design.
  5. [§7.3] The sentence 'the most sensitive of the well-calibrated partition tests (§5.2)' repeats the unsupported superlative from the abstract; it should be revised consistently with the response to the first major comment.

Circularity Check

1 steps flagged · score 2.0 of 10

The alignment functional A(δ) is by definition E(C); Proposition 1(i) is an acknowledged identity, but the central power claims rest on independent simulations, so circularity is minor.

  1. self definitional [Section 3.1, Eq. (2), Proposition 1(i)]
    "Define the alignment functional A(δ) = Σ_g (1−2π̄_g) δ_g / V_g ... The identity in part (i) is immediate—C is linear in the group residuals, so E(C) = A(δ)."

    A(δ) is defined to equal the expected value of the correction term C. Proposition 1(i) (E(T_EF) − E(Ĉ_G) = −A(δ)) is therefore true by construction. The paper then presents A(δ) as an explanatory quantity that 'predicts' where the test gains power, but that prediction is a restatement of the mean shift already encoded in the definition of C. The paper openly acknowledges the identity is immediate, and the empirical power results are independently simulated rather than fitted, so this is a presentational tautology rather than a fabricated prediction.

full rationale

The paper's derivation chain is largely self-contained and honest. The only definitional shortcut is the alignment functional: A(δ) is literally E(C) under the local-alternative setup, so Proposition 1(i) is an identity rather than a derived result. However, the paper explicitly says 'The identity in part (i) is immediate', so no hidden circularity is involved. The empirical content—size control, power comparisons, the monotone relationship between |A(δ)| and the EF advantage—comes from K=5000 and K=3000 simulations and is not manufactured by fitting parameters to make the test work. Self-citation is limited to the authors' own R package [3], which is not load-bearing for the statistical claims; the key references (Farrington, Osius–Rojek, Moore–Spruill, Hosmer–Lemeshow) are external. The superlative claim about 'no well-calibrated partition test is more sensitive to asymmetric-link misfit' is weakened by the Monte-Carlo error against the equal-width HL variant in Table 5, but that is a correctness/evidential concern, not circularity. Overall score 2: one acknowledged definitional identity presented as an explanatory mechanism, with independent simulation support for the central claims.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The method introduces no free parameters fitted to data beyond the conventional group count; it relies on standard asymptotic machinery for grouped Pearson statistics and on the local-alternative assumption that group biases are first-order with fixed grouping. No new physical entities are postulated.

free parameters (1)
  • Number of groups G = 10 (default)
    Chosen by hand following HL convention; size and power are group-count sensitive, and the paper recommends reporting G. Not fitted to data.
assumptions (4)
  • standard math Moore-Spruill CLT gives the grouped Pearson statistic a χ²_{G−2} limit under model-based grouping.
    Invoked in §3.1 and Appendix A(iii) to justify the χ²_{G−2} reference for T_EF.
  • domain assumption Under O(n^{−1/2}) local alternatives, group residuals have first-order means δ_g and variances V_g, with grouping fixed to first order.
    Section 3.1 and Appendix A; the entire alignment-functional argument and the op(1) claim for C depend on it.
  • standard math The Farrington/Osius-Rojek first-order standardization is valid and applicable after grouping.
    Cited in §2.2 and Appendix A rather than reproduced; the correction C is built from this line.
  • domain assumption For symmetric misfit, fitted risks are placed symmetrically about 1/2 so the odd/even decomposition holds.
    Proposition 1(ii); used to conclude A(δ)=0 for symmetric departures.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression." pith.science (2026). https://pith.science/paper/4M6PIQX3

@misc{pith2026260715454,
  author       = {Pith},
  title        = {Pith review of: A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4M6PIQX3}},
  note         = {Machine review of arXiv:2607.15454}
}
abstract

Goodness-of-fit assessment for the binary logistic regression model is difficult when covariates are continuous: the data are effectively sparse, the classical Pearson and deviance tests fail, and practitioners rely on partition-based tests, such as the Hosmer-Lemeshow test, that group observations before comparing observed and expected counts. We study a partition test that modifies the Hosmer-Lemeshow statistic with a single directional correction term, weighted by $(1-2\bar\pi_g)$ and referred to a $\chi^2_{G-2}$ distribution. The correction is the grouped form of the Osius-Rojek/Farrington standardization; grouping makes it well defined in the sparse regime, and it targets the asymmetric over- and under-prediction that a misspecified link induces. A single alignment functional captures its effect, predicting where the test gains power (asymmetric-link misspecification) and where it does not (symmetric departures, and covariate-space structure that no probability-grouping test can see). In simulations the test holds its size; no well-calibrated partition test is more sensitive to asymmetric-link misfit, and it clearly exceeds Hosmer-Lemeshow there, most so for the complementary log-log link -- a modest gain that fades as $n$ grows; it ties Hosmer-Lemeshow on an omitted interaction and is less powerful on an omitted quadratic (by about ten percentage points at $n=1000$). A real-data application illustrates its use, and the test is implemented in the R package ebrahim.gof.

Figures

Figures reproduced from arXiv: 2607.15454 by the authors.

Figure 1
Figure 1. Empirical type I error (%) of EF and HL under correct specification, across sample [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Power to detect a misspecified (Aranda–Ordaz) link at the 5% level, [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. The EF power advantage over HL, power(EF) − power(HL), mapped over link asymmetry (Aranda–Ordaz 1 − α: 0 = logit, 1 = complementary log–log) and sample size n (tests at the 5% level, K = 5000). The advantage is largest at strong asymmetry and moderate n, and vanishes toward the logit (left) and as n grows (top): a finite-sample, directional phenomenon. two columns are within a few points of each other in almost ever… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Empirical power (%) of five partition-based tests—EF, Hosmer–Lemeshow (HL), [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Power difference power(EF) − power(HL) by scenario (n = 1000, the 5% level). EF gains only on asymmetric links (green, top); it is within Monte-Carlo error of HL on most departures (gray) and trails on symmetric curvature and covariate-space interactions (orange, botto…
Figure 6
Figure 6. Figure 6: Power of the Ebrahim–Farrington (EF) test against the Hosmer–Lemeshow (HL) test [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: The alignment functional A(δ) predicts EF’s advantage. For the Aranda–Ordaz link sweep, the power difference power(EF) − power(HL) rises monotonically with −A(δ) (more directional misfit) for n ≤ 2000; at n = 5000 both tests saturate and the difference collapses. Symme…
Figure 8
Figure 8. Figure 8: The directional read-out on a representative complementary-log–log misfit (Aranda– [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

29 extracted references · 11 canonical work pages

  1. [1]

    Wiley Series in Probability and Statistics

    Alan Agresti.Categorical Data Analysis. Wiley Series in Probability and Statistics. John Wiley & Sons, Hoboken, NJ, 3rd edition, 2013

  2. [2]

    Aranda-Ordaz

    Francisco J. Aranda-Ordaz. On two families of transformations to additivity for binary response data.Biometrika, 68(2):357–363, 1981. doi: 10.1093/biomet/68.2.357

  3. [3]

    URL https://CRAN.R-project.org/package=ebrahim.gof

    Ebrahim Khaled Ebrahim.ebrahim.gof: Ebrahim–Farrington Goodness-of-Fit Test for Logistic Regression, 2026. URL https://CRAN.R-project.org/package=ebrahim.gof. R package version 2.1.0,https://doi.org/10.32614/CRAN.package.ebrahim.gof

  4. [4]

    C. P. Farrington. On assessing goodness of fit of generalized linear models to sparse data. Journal of the Royal Statistical Society: Series B (Methodological), 58(2):349–360, 1996. doi: 10.1111/j.2517-6161.1996.tb02084.x

  5. [5]

    Alexander Henzi, Marius Puke, Timo Dimitriadis, and Johanna F. Ziegel. A safe Hosmer– Lemeshow test.The New England Journal of Statistics in Data Science, 2(2):175–189, 2024. doi: 10.51387/23-NEJSDS56

  6. [6]

    Hosmer and Nils Lid Hjort

    David W. Hosmer and Nils Lid Hjort. Goodness-of-fit processes for logistic regression: Simulation results.Statistics in Medicine, 21(18):2723–2738, 2002. doi: 10.1002/sim.1200

  7. [7]

    Hosmer and Stanley Lemeshow

    David W. Hosmer and Stanley Lemeshow. Goodness of fit tests for the multiple logistic regression model.Communications in Statistics – Theory and Methods, 9(10):1043–1069,

  8. [8]

    Hosmer, Trina Hosmer, Saskia Le Cessie, and Stanley Lemeshow

    David W. Hosmer, Trina Hosmer, Saskia Le Cessie, and Stanley Lemeshow. A comparison of goodness-of-fit tests for the logistic regression model.Statistics in Medicine, 16(9):965–980,

Show all 29 references
  1. [9]

    Hosmer, Stanley Lemeshow, and Rodney X

    David W. Hosmer, Stanley Lemeshow, and Rodney X. Sturdivant.Applied Logistic Regres- sion. Wiley Series in Probability and Statistics. John Wiley & Sons, Hoboken, NJ, 3rd edition, 2013. doi: 10.1002/9781118548387

  2. [10]

    Global goodness-of-fit tests in logistic regression with sparse data.Statistics in Medicine, 21(24):3789–3801, 2002

    Oliver Kuss. Global goodness-of-fit tests in logistic regression with sparse data.Statistics in Medicine, 21(24):3789–3801, 2002. doi: 10.1002/sim.1421

  3. [11]

    van Houwelingen

    Saskia le Cessie and Johannes C. van Houwelingen. A goodness-of-fit test for binary regression models, based on smoothing methods.Biometrics, 47(4):1267–1282, 1991. doi: 10.2307/2532385

  4. [12]

    A comprehensive comparison of goodness-of-fit tests for logistic regression models.Journal of Statistical Computation and Simulation, 94(9): 1877–1906, 2024

    Yiwen Liu, Yisha Li, and Jiaxin Xie. A comprehensive comparison of goodness-of-fit tests for logistic regression models.Journal of Statistical Computation and Simulation, 94(9): 1877–1906, 2024. doi: 10.1080/00949655.2023.2301037

  5. [13]

    On the asymptotic distribution of Pearson’s statistic in linear exponential- family models.International Statistical Review, 53(1):61–67, 1985

    Peter McCullagh. On the asymptotic distribution of Pearson’s statistic in linear exponential- family models.International Statistical Review, 53(1):61–67, 1985. doi: 10.2307/1402880

  6. [14]

    Moore and M

    David S. Moore and M. C. Spruill. Unified large-sample theory of general chi-squared statistics for tests of fit.The Annals of Statistics, 3(3):599–616, 1975. doi: 10.1214/aos/ 1176343125

  7. [15]

    Pennell, and Stanley Lemeshow

    Giovanni Nattino, Michael L. Pennell, and Stanley Lemeshow. Assessing the goodness of fit of logistic regression models in large samples: A modification of the Hosmer–Lemeshow test.Biometrics, 76(2):549–560, 2020. doi: 10.1111/biom.13249

  8. [16]

    Normal goodness-of-fit tests for multinomial models with large degrees of freedom.Journal of the American Statistical Association, 87(420): 1145–1152, 1992

    Gerhard Osius and Dieter Rojek. Normal goodness-of-fit tests for multinomial models with large degrees of freedom.Journal of the American Statistical Association, 87(420): 1145–1152, 1992. doi: 10.1080/01621459.1992.10476271

  9. [17]

    Pigeon and Joseph F

    Joseph G. Pigeon and Joseph F. Heyse. An improved goodness of fit statistic for prob- ability prediction models.Biometrical Journal, 41(1):71–82, 1999. doi: 10.1002/(SICI) 1521-4036(199903)41:1⟨71::AID-BIMJ71⟩3.0.CO;2-O

  10. [18]

    Robinson

    Erik Pulkstenis and Timothy J. Robinson. Two goodness-of-fit tests for logistic regression models with continuous covariates.Statistics in Medicine, 21(1):79–93, 2002. doi: 10.1002/ sim.943

  11. [19]

    Steyerberg.Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating

    Ewout W. Steyerberg.Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating. Statistics for Biology and Health. Springer, Cham, Switzerland, 2nd edition, 2019. doi: 10.1007/978-3-030-16399-0

  12. [20]

    Therese A. Stukel. Generalized logistic models.Journal of the American Statistical Association, 83(402):426–431, 1988. doi: 10.1080/01621459.1988.10478613

  13. [21]

    Nikola Surjanovic and Thomas M. Loughin. Improving the Hosmer–Lemeshow goodness-of- fit test in large models with replicated Bernoulli trials.Journal of Applied Statistics, 51(7): 1399–1411, 2024. doi: 10.1080/02664763.2023.2272223

  14. [22]

    Anastasios A. Tsiatis. A note on a goodness-of-fit test for the logistic regression model. Biometrika, 67(1):250–251, 1980. doi: 10.1093/biomet/67.1.250. 28

  15. [23]

    McLernon, Maarten van Smeden, Laure Wynants, and Ewout W

    Ben Van Calster, David J. McLernon, Maarten van Smeden, Laure Wynants, and Ewout W. Steyerberg. Calibration: the Achilles heel of predictive analytics.BMC Medicine, 17(1): 230, 2019. doi: 10.1186/s12916-019-1466-7

  16. [24]

    Venables and Brian D

    William N. Venables and Brian D. Ripley.Modern Applied Statistics with S. Springer, New York, 4th edition, 2002. doi: 10.1007/978-0-387-21706-2. R packageMASS, including the birthwtdata set

  17. [25]

    Increasing the power: A practical approach to goodness-of-fit test for logistic regression models with continuous predictors

    Xian-Jin Xie, Jane Pendergast, and William Clarke. Increasing the power: A practical approach to goodness-of-fit test for logistic regression models with continuous predictors. Computational Statistics & Data Analysis, 52(5):2703–2713, 2008. doi: 10.1016/j.csda.2007. 10.004

  18. [26]

    BAGofT: A binary adap- tive goodness-of-fit test for the logistic regression model.arXiv preprint arXiv:1911.03068,

    Jiawei Zhang, Zhigang Zhang, Kani Chen, Ai Ni, and Zhiliang Lin. BAGofT: A binary adap- tive goodness-of-fit test for the logistic regression model.arXiv preprint arXiv:1911.03068,

  19. [1980]

    doi: 10.1080/03610928008827941. 27

  20. [1997]

    doi: 10.1002/(SICI)1097-0258(19970515)16:9⟨965::AID-SIM509⟩3.0.CO;2-O

  21. [2019]

    Zhang et al

    See also J. Zhang et al. (2023),Statistica Sinica. 29

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.