REVIEW 3 major objections 5 minor 29 references
A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression
T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read This paper shows that adding a single signed correction term to the Hosmer–Lemeshow statistic—weighted by (1−2π̄_g) and referred to a χ²_{G−2} distribution—produces a partition goodness-of-fit test that holds its size, matches Hosmer–Lemesh
desk verdict A small, honest refinement of the Hosmer–Lemeshow test whose simulations back its claims, except for one superlative in the abstract that its own table does not support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The statistic T_EF = Ĉ_G − C, where Ĉ_G is the Hosmer–Lemeshow Pearson sum over equal-size deciles of fitted risk and C = Σ_g (1−2π̄_g)(o_g−e_g)/V_g is a signed correction with per-group variance V_g = n_g π̄_g(1−π̄_g). The carrying mechanism is the alignment functional A(δ)=Σ_g (1−2π̄_g)δ_g/V_g: a linear functional of the group-residual bias pattern whose odd weight (1−2π̄_g) annihilates symmetric misfit and retains only directional misfit. This functional does the explanatory work, predicting the sign and magnitude of the power difference between EF and HL; the χ²_{G−2} reference is justified by the standard large-sample model-based-grouping argument, with the correction itself vanishing
What would settle it
Simulate the Aranda–Ordaz asymmetric-link setting (α near 0.1) at n=5000, but recompute decile boundaries under the true model and compare the empirical null distribution of T_EF using estimated versus oracle deciles; if the two differ by more than Monte-Carlo error, or if the EF-versus-HL power difference at moderate n fails to follow the sign of A(δ) when grouping is re-estimated, the fixed-grouping assumption breaks.
Extended reading notes
Core claim
The central claim is that the directional Hosmer–Lemeshow correction T_EF = Ĉ_G − C, with C = Σ_g (1−2π̄_g)(o_g−e_g)/V_g, turns the statistic's sensitivity on and off exactly according to the sign and size of the alignment functional A(δ)=Σ_g (1−2π̄_g)δ_g/V_g. Because the weight (1−2π̄_g) is odd about the midpoint of predicted risk, it cancels any symmetric pattern of group residual bias; only the odd, directional component survives. Under local alternatives, a negatively aligned odd pattern—the signature of an asymmetric link such as complementary log–log—raises the test's non-centrality and gives T_EF more power than HL, while a positively aligned pattern (an omitted quadratic in a skewed
Load-bearing premise
The load-bearing premise is that the decile grouping can be treated as fixed to first order under the null and local alternatives, so that the correction weights w_g/V_g are non-stochastic and the group residuals have the standard model-based covariance; if misspecification shifts the decile boundaries materially, both the χ²_{G−2} calibration and the predicted power gain could change (Section 3.1 and Appendix A).
Editorial extensions
If this is right
- Practitioners can replace or accompany the HL p-value with T_EF at no extra computational cost: the test uses the same decile grouping, the same reference distribution, and a single closed-form correction, making it a drop-in refinement.
- For asymmetric-link misfit—common in dose–response, discrete-time survival, and rare-event models—EF is the most sensitive well-calibrated partition test, with the gain visible at moderate n and fading as n grows; the effect is finite-sample, not asymptotic.
- For symmetric departures such as omitted quadratic terms, EF is less powerful than HL because the alignment functional is positive; the test is not a universal improvement.
- For covariate-space departures like omitted interactions, EF ties HL and both are beaten by covariate-space grouping tests; no probability-grouping test can detect such structure.
- The per-group signed contributions {c_g} provide a directional read-out that can localize misfit on the risk scale and suggest trying a complementary-log–log-style link rather than an added polynomial.
Reading between the lines
- Because the alignment functional depends only on the odd component of the residual-bias pattern and the odd weight, the same correction idea could be ported to other symmetric grouping schemes (equal-width bins, covariate-based partitions) as long as the weight remains odd about the center of the risk scale; the paper does not explore this.
- The correction is O(n^{-1/2}) under the null, so its advantage is inherently finite-sample; a studentized version that gives the directional term its own rejection region could make the signal first-order and more useful at large n.
- The calibration-slope appendix implies that in external validation the test is direction-dependent: it helps detect over-shrunk models but is less powerful than HL against overfitted transported models, so the sign of the correction can itself be read as a calibration diagnostic.
- One testable extension: since the paper shows power tracks |A(δ)| monotonically, one could compute A(δ) under a menu of candidate asymmetric links for a given dataset and use it to decide which link to try after EF rejects, turning the test into a model-selection aid.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes and studies a modification of the Hosmer–Lemeshow (HL) test for sparse logistic regression: T_EF = C_hat_G − C, where C = Σ_g (1−2π̄_g)(o_g−e_g)/V_g is a signed directional correction referred to a χ²_{G−2} distribution. The correction is the grouped form of the Osius–Rojek/Farrington standardization. The authors claim that this correction is exactly inert for symmetric misfit and beneficial for asymmetric-link misfit, that a single 'alignment functional' A(δ) predicts when the correction helps or hurts, and that no well-calibrated partition test is more sensitive to asymmetric-link misspecification. Evidence includes a 24-cell size factorial, power sweeps over omitted quadratic, omitted interaction, and Aranda–Ordaz link families, a seven-test family comparison at n=1000, real-data concordance checks, and an R package with archived simulation code.
Significance. If the claims are supported, the paper would provide a useful, refit-free drop-in companion to HL with a directional diagnostic and an honest map of where the test wins and loses. Strengths include the deliberate reporting of Monte-Carlo error, the explicit 'no-free-lunch' rows, the reproducible R package and Zenodo archive, and the candid limitations section. The main contribution is not a new correction — the authors correctly credit Farrington and Osius–Rojek — but a characterization of the behavior of a grouped version of that correction. The central superlative claim, however, is not supported by the reported precision, and the theoretical 'prediction' is partly an identity; after revision of those claims the paper would be a solid empirical and interpretive contribution.
major comments (3)
- [Abstract; §5.2, Table 5; §7.4] The claim that 'no well-calibrated partition test is more sensitive to asymmetric-link misfit' is not supported by Table 5. At K=3000, EF gives 49.7% on cloglog versus HLeqw 48.3%, and 42.2% versus 40.9% on loglog. The Monte-Carlo standard error for the difference of two independent proportions near 0.5 is approximately sqrt(0.5·0.5/3000)·sqrt(2) ≈ 1.3 percentage points, so both margins (1.4 and 1.3) lie within the ±2 SE band of about ±2.6 points. HLeqw is well-calibrated (null size 4.7%), so the data support EF matching the best calibrated partition test, not strictly exceeding it. The text in §5.2 concedes the margin is 'within Monte-Carlo error', but the abstract and §7.4 do not carry this qualification. Please rephrase the claim to 'among the tests compared, EF exceeds the decile-based HL and ties the equal-width HL', or increase K to resolve the comparison. A universal statement ove
- [§3.1, Proposition 1; Appendix A, Eq. (4)] Proposition 1(i) is an identity: E(C) = A(δ) follows immediately from the definitions of C and A(δ). The statement that 'A(δ) predicts' where the test gains power is therefore a restatement of the mean shift induced by the correction, not an independent theoretical prediction. Moreover, A(δ) is computed from the true data-generating process (§4.4(c), Table 6), not from data, so it cannot be evaluated by a practitioner without knowing the misspecification. This does not invalidate the simulation results, but the paper should present A(δ) as an interpretive/explanatory device that organizes the simulation outcomes, rather than as a falsifiable prediction independent of the definition. The empirical monotonicity in Table 6 is then evidence of the mechanism, not a test of a separate hypothesis.
- [§3.1; Appendix A(iii)] The asymptotic arguments treat the decile grouping as fixed. The null calibration relies on the Moore–Spruill result (cited, not verified) together with the claim that C = b^T Z with b = O(sqrt(G/n)), so C = o_p(1) under fixed grouping. But under the local alternatives used for the power analysis, the fitted probabilities shift at O(n^{−1/2}), and the decile boundaries can shift at the same order as the residual biases δ_g. The first-order argument does not account for this boundary movement. The size simulations and KS checks are reassuring for the null, but the A(δ) computations and the power predictions implicitly assume fixed grouping. Please either supply an argument covering the boundary shift or state explicitly that the theory is first-order under fixed grouping and that the simulation evidence is the primary support.
minor comments (5)
- [Table 3; §4.4(b)] Table 3 labels the interaction row as 'continuous-by-continuous product term', but §4.4(b) defines the interaction as ψ x d with d ~ Bernoulli(0.5), i.e. continuous-by-binary. Please correct the table header or the text.
- [Table 5; §4.1] The equal-width HL variant HLeqw appears first in Table 5 but is not defined in Table 1 or in §4.1. Add it to the family list or define it in the text before first use.
- [Figure 7 caption] The caption says 'the power gain rises monotonically with A(δ) (n 2000)', but the figure plots n = 500, 1000, 2000, and 5000. Please correct the caption.
- [§5.3] The representative read-out dataset is selected from among datasets where EF rejects and HL does not. This conditioning should be stated in the text, because the displayed tilt may be stronger than in a typical dataset from the same design.
- [§7.3] The sentence 'the most sensitive of the well-calibrated partition tests (§5.2)' repeats the unsupported superlative from the abstract; it should be revised consistently with the response to the first major comment.
Circularity Check
The alignment functional A(δ) is by definition E(C); Proposition 1(i) is an acknowledged identity, but the central power claims rest on independent simulations, so circularity is minor.
-
self definitional
[Section 3.1, Eq. (2), Proposition 1(i)]
"Define the alignment functional A(δ) = Σ_g (1−2π̄_g) δ_g / V_g ... The identity in part (i) is immediate—C is linear in the group residuals, so E(C) = A(δ)."
A(δ) is defined to equal the expected value of the correction term C. Proposition 1(i) (E(T_EF) − E(Ĉ_G) = −A(δ)) is therefore true by construction. The paper then presents A(δ) as an explanatory quantity that 'predicts' where the test gains power, but that prediction is a restatement of the mean shift already encoded in the definition of C. The paper openly acknowledges the identity is immediate, and the empirical power results are independently simulated rather than fitted, so this is a presentational tautology rather than a fabricated prediction.
full rationale
The paper's derivation chain is largely self-contained and honest. The only definitional shortcut is the alignment functional: A(δ) is literally E(C) under the local-alternative setup, so Proposition 1(i) is an identity rather than a derived result. However, the paper explicitly says 'The identity in part (i) is immediate', so no hidden circularity is involved. The empirical content—size control, power comparisons, the monotone relationship between |A(δ)| and the EF advantage—comes from K=5000 and K=3000 simulations and is not manufactured by fitting parameters to make the test work. Self-citation is limited to the authors' own R package [3], which is not load-bearing for the statistical claims; the key references (Farrington, Osius–Rojek, Moore–Spruill, Hosmer–Lemeshow) are external. The superlative claim about 'no well-calibrated partition test is more sensitive to asymmetric-link misfit' is weakened by the Monte-Carlo error against the equal-width HL variant in Table 5, but that is a correctness/evidential concern, not circularity. Overall score 2: one acknowledged definitional identity presented as an explanatory mechanism, with independent simulation support for the central claims.
Assumptions & free parameters
free parameters (1)
- Number of groups G =
10 (default)
assumptions (4)
- standard math Moore-Spruill CLT gives the grouped Pearson statistic a χ²_{G−2} limit under model-based grouping.
- domain assumption Under O(n^{−1/2}) local alternatives, group residuals have first-order means δ_g and variances V_g, with grouping fixed to first order.
- standard math The Farrington/Osius-Rojek first-order standardization is valid and applicable after grouping.
- domain assumption For symmetric misfit, fitted risks are placed symmetrically about 1/2 so the odd/even decomposition holds.
Cite this review
Pith. "Pith review of A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression." pith.science (2026). https://pith.science/paper/4M6PIQX3
@misc{pith2026260715454,
author = {Pith},
title = {Pith review of: A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression},
year = {2026},
howpublished = {\url{https://pith.science/paper/4M6PIQX3}},
note = {Machine review of arXiv:2607.15454}
}
abstract
Goodness-of-fit assessment for the binary logistic regression model is difficult when covariates are continuous: the data are effectively sparse, the classical Pearson and deviance tests fail, and practitioners rely on partition-based tests, such as the Hosmer-Lemeshow test, that group observations before comparing observed and expected counts. We study a partition test that modifies the Hosmer-Lemeshow statistic with a single directional correction term, weighted by $(1-2\bar\pi_g)$ and referred to a $\chi^2_{G-2}$ distribution. The correction is the grouped form of the Osius-Rojek/Farrington standardization; grouping makes it well defined in the sparse regime, and it targets the asymmetric over- and under-prediction that a misspecified link induces. A single alignment functional captures its effect, predicting where the test gains power (asymmetric-link misspecification) and where it does not (symmetric departures, and covariate-space structure that no probability-grouping test can see). In simulations the test holds its size; no well-calibrated partition test is more sensitive to asymmetric-link misfit, and it clearly exceeds Hosmer-Lemeshow there, most so for the complementary log-log link -- a modest gain that fades as $n$ grows; it ties Hosmer-Lemeshow on an omitted interaction and is less powerful on an omitted quadratic (by about ten percentage points at $n=1000$). A real-data application illustrates its use, and the test is implemented in the R package ebrahim.gof.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Wiley Series in Probability and Statistics
Alan Agresti.Categorical Data Analysis. Wiley Series in Probability and Statistics. John Wiley & Sons, Hoboken, NJ, 3rd edition, 2013
2013
-
[2]
Francisco J. Aranda-Ordaz. On two families of transformations to additivity for binary response data.Biometrika, 68(2):357–363, 1981. doi: 10.1093/biomet/68.2.357
-
[3]
URL https://CRAN.R-project.org/package=ebrahim.gof
Ebrahim Khaled Ebrahim.ebrahim.gof: Ebrahim–Farrington Goodness-of-Fit Test for Logistic Regression, 2026. URL https://CRAN.R-project.org/package=ebrahim.gof. R package version 2.1.0,https://doi.org/10.32614/CRAN.package.ebrahim.gof
-
[4]
C. P. Farrington. On assessing goodness of fit of generalized linear models to sparse data. Journal of the Royal Statistical Society: Series B (Methodological), 58(2):349–360, 1996. doi: 10.1111/j.2517-6161.1996.tb02084.x
arXiv 1996
-
[5]
Alexander Henzi, Marius Puke, Timo Dimitriadis, and Johanna F. Ziegel. A safe Hosmer– Lemeshow test.The New England Journal of Statistics in Data Science, 2(2):175–189, 2024. doi: 10.51387/23-NEJSDS56
-
[6]
David W. Hosmer and Nils Lid Hjort. Goodness-of-fit processes for logistic regression: Simulation results.Statistics in Medicine, 21(18):2723–2738, 2002. doi: 10.1002/sim.1200
-
[7]
Hosmer and Stanley Lemeshow
David W. Hosmer and Stanley Lemeshow. Goodness of fit tests for the multiple logistic regression model.Communications in Statistics – Theory and Methods, 9(10):1043–1069,
-
[8]
Hosmer, Trina Hosmer, Saskia Le Cessie, and Stanley Lemeshow
David W. Hosmer, Trina Hosmer, Saskia Le Cessie, and Stanley Lemeshow. A comparison of goodness-of-fit tests for the logistic regression model.Statistics in Medicine, 16(9):965–980,
Show all 29 references
-
[9]
Hosmer, Stanley Lemeshow, and Rodney X
David W. Hosmer, Stanley Lemeshow, and Rodney X. Sturdivant.Applied Logistic Regres- sion. Wiley Series in Probability and Statistics. John Wiley & Sons, Hoboken, NJ, 3rd edition, 2013. doi: 10.1002/9781118548387
2013 doi
-
[10]
Global goodness-of-fit tests in logistic regression with sparse data.Statistics in Medicine, 21(24):3789–3801, 2002
Oliver Kuss. Global goodness-of-fit tests in logistic regression with sparse data.Statistics in Medicine, 21(24):3789–3801, 2002. doi: 10.1002/sim.1421
2002 doi
-
[11]
van Houwelingen
Saskia le Cessie and Johannes C. van Houwelingen. A goodness-of-fit test for binary regression models, based on smoothing methods.Biometrics, 47(4):1267–1282, 1991. doi: 10.2307/2532385
1991 doi
-
[12]
A comprehensive comparison of goodness-of-fit tests for logistic regression models.Journal of Statistical Computation and Simulation, 94(9): 1877–1906, 2024
Yiwen Liu, Yisha Li, and Jiaxin Xie. A comprehensive comparison of goodness-of-fit tests for logistic regression models.Journal of Statistical Computation and Simulation, 94(9): 1877–1906, 2024. doi: 10.1080/00949655.2023.2301037
1906
-
[13]
On the asymptotic distribution of Pearson’s statistic in linear exponential- family models.International Statistical Review, 53(1):61–67, 1985
Peter McCullagh. On the asymptotic distribution of Pearson’s statistic in linear exponential- family models.International Statistical Review, 53(1):61–67, 1985. doi: 10.2307/1402880
1985 doi
-
[14]
Moore and M
David S. Moore and M. C. Spruill. Unified large-sample theory of general chi-squared statistics for tests of fit.The Annals of Statistics, 3(3):599–616, 1975. doi: 10.1214/aos/ 1176343125
1975 doi
-
[15]
Pennell, and Stanley Lemeshow
Giovanni Nattino, Michael L. Pennell, and Stanley Lemeshow. Assessing the goodness of fit of logistic regression models in large samples: A modification of the Hosmer–Lemeshow test.Biometrics, 76(2):549–560, 2020. doi: 10.1111/biom.13249
2020 doi
-
[16]
Normal goodness-of-fit tests for multinomial models with large degrees of freedom.Journal of the American Statistical Association, 87(420): 1145–1152, 1992
Gerhard Osius and Dieter Rojek. Normal goodness-of-fit tests for multinomial models with large degrees of freedom.Journal of the American Statistical Association, 87(420): 1145–1152, 1992. doi: 10.1080/01621459.1992.10476271
1992
-
[17]
Pigeon and Joseph F
Joseph G. Pigeon and Joseph F. Heyse. An improved goodness of fit statistic for prob- ability prediction models.Biometrical Journal, 41(1):71–82, 1999. doi: 10.1002/(SICI) 1521-4036(199903)41:1⟨71::AID-BIMJ71⟩3.0.CO;2-O
1999 doi
-
[18]
Robinson
Erik Pulkstenis and Timothy J. Robinson. Two goodness-of-fit tests for logistic regression models with continuous covariates.Statistics in Medicine, 21(1):79–93, 2002. doi: 10.1002/ sim.943
2002
-
[19]
Steyerberg.Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating
Ewout W. Steyerberg.Clinical Prediction Models: A Practical Approach to Development, Validation, and Updating. Statistics for Biology and Health. Springer, Cham, Switzerland, 2nd edition, 2019. doi: 10.1007/978-3-030-16399-0
2019 doi
-
[20]
Therese A. Stukel. Generalized logistic models.Journal of the American Statistical Association, 83(402):426–431, 1988. doi: 10.1080/01621459.1988.10478613
1988
-
[21]
Nikola Surjanovic and Thomas M. Loughin. Improving the Hosmer–Lemeshow goodness-of- fit test in large models with replicated Bernoulli trials.Journal of Applied Statistics, 51(7): 1399–1411, 2024. doi: 10.1080/02664763.2023.2272223
2024
-
[22]
Anastasios A. Tsiatis. A note on a goodness-of-fit test for the logistic regression model. Biometrika, 67(1):250–251, 1980. doi: 10.1093/biomet/67.1.250. 28
1980 doi
-
[23]
McLernon, Maarten van Smeden, Laure Wynants, and Ewout W
Ben Van Calster, David J. McLernon, Maarten van Smeden, Laure Wynants, and Ewout W. Steyerberg. Calibration: the Achilles heel of predictive analytics.BMC Medicine, 17(1): 230, 2019. doi: 10.1186/s12916-019-1466-7
2019 doi
-
[24]
Venables and Brian D
William N. Venables and Brian D. Ripley.Modern Applied Statistics with S. Springer, New York, 4th edition, 2002. doi: 10.1007/978-0-387-21706-2. R packageMASS, including the birthwtdata set
2002 doi
-
[25]
Increasing the power: A practical approach to goodness-of-fit test for logistic regression models with continuous predictors
Xian-Jin Xie, Jane Pendergast, and William Clarke. Increasing the power: A practical approach to goodness-of-fit test for logistic regression models with continuous predictors. Computational Statistics & Data Analysis, 52(5):2703–2713, 2008. doi: 10.1016/j.csda.2007. 10.004
2008 doi
-
[26]
BAGofT: A binary adap- tive goodness-of-fit test for the logistic regression model.arXiv preprint arXiv:1911.03068,
Jiawei Zhang, Zhigang Zhang, Kani Chen, Ai Ni, and Zhiliang Lin. BAGofT: A binary adap- tive goodness-of-fit test for the logistic regression model.arXiv preprint arXiv:1911.03068,
1911 arXiv
-
[1980]
doi: 10.1080/03610928008827941. 27
-
[1997]
doi: 10.1002/(SICI)1097-0258(19970515)16:9⟨965::AID-SIM509⟩3.0.CO;2-O
-
[2019]
Zhang et al
See also J. Zhang et al. (2023),Statistica Sinica. 29
2023
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.