REVIEW 2 major objections 6 minor 1 cited by
Receiver operating characteristic curve analysis with non-ignorable missing disease status
T0 review · 2 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper shows that the ROC curve and AUC of a biomarker can be estimated and tested with valid confidence intervals when disease status is missing non-ignorably, no instrumental variable required.
desk verdict This paper is a solid, mostly-correct step forward for ROC/AUC with non-ignorable missing disease status; it deserves review, but Theorem 2(b) needs an added positivity condition on f0 at the quantile. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing identity is model (7), which re-expresses the verification propensity as $P(R=1|x,v) = 1/(1+\exp\{\psi_1+\psi_2 x+\psi_3^T v + c(x,v;\mu,\beta)\})$, where the offset $c(x,v;\mu,\beta) = \log E(e^{\beta Y}|x,v,R=1)$ is computable from the disease model. This identity converts an unobserved-$Y$ problem into a logistic regression on observed data and is what makes identifiability without an instrumental variable plausible. The second piece is Proposition 1: for weights $g(X,V,R;\mu,\beta) = P(Y=1|X,V,R)$ and $1-g$, the identities $E\{g\,I(X\le x)\} = E\{g\}\,F_1(x)$ and $E\{(1-g)I(X\le x)\} = E\{1-g\}\,F_0(x)$ let the biomarker distributions in healthy and diseased groups be estimated by weighted empirical distribution functions without modelling the biomarker distribution. These two pieces together carry the estimation of $\mathrm{ROC}(s)$ and AUC.
What would settle it
Generate data under models (3)-(4) with $\mu_2=0$ and maximize the observed-data likelihood from many starting values; if the maximum is not unique (a flat ridge in parameter space), Theorem 1's identifiability claim fails outside Condition C1. Alternatively, obtain gold-standard $Y$ for a random subsample of patients with $R=0$ and compare the observed $R=1$ rates in $(Y,X,V)$ cells with the fitted probabilities from model (4); systematic mismatch would falsify the verification model.
Extended reading notes
Core claim
On its own terms, the paper claims that the two logistic models (3) and (4) — disease status among verified patients, and verification status given disease, biomarker, and covariates — are identifiable from observed data alone, with no instrumental variable, as long as the biomarker is continuous, the regressors are stochastically linearly independent, and the biomarker coefficient $\mu_2$ in the disease model is nonzero (Theorem 1). The observable verification propensity $P(R=1|x,v)$ is shown to be logistic with an offset $c(x,v;\mu,\beta) = \log E(e^{\beta Y}|x,v,R=1)$, and identifiability is established through that representation. The full-likelihood estimator is asymptotically normal, and the plug-in estimators of $\mathrm{ROC}(s)$ and AUC inherit that property, so Wald intervals with plug-in variances have correct coverage under correct specification (Theorem 2). Proposition 1 supplies the weighted empirical distribution identities used to turn the fitted models into ROC and AUC estimates.
Load-bearing premise
The load-bearing premise is that the verification model (4) is correctly specified as a logistic regression on the biomarker, covariates, and the unobserved disease status $Y$, since $Y$ is missing whenever verification is skipped and the data cannot directly check this assumption; identifiability also requires the biomarker coefficient in the disease model to be nonzero.
Editorial extensions
If this is right
- Researchers no longer need to search for an instrumental variable before applying the method; identifiability holds under Condition C1 alone.
- Using all observed patients, including those with unverified disease status, the full-likelihood AUC estimator has smaller mean squared error than the IPW estimator in the paper's simulations.
- Wald confidence intervals for AUC and $\mathrm{ROC}(s)$ achieve coverage close to the nominal 95% level when models (3) and (4) are correctly specified.
- The two-step goodness-of-fit procedure provides a practical check of both models despite the missingness of $Y$ in unverified patients.
- On the NACC Alzheimer's data, the method yields MMSE AUC 0.791 with 95% CI (0.779, 0.803), and the estimated non-ignorability parameter $\beta$ is significant, whereas the IPW confidence interval is extremely wide.
Reading between the lines
- Editorial: the offset-logistic identification argument should extend to other binary links, such as probit, and to ordinal disease categories, since only the induced form of the verification propensity and a nonzero biomarker slope are needed.
- Editorial: the paper's two-step goodness-of-fit check validates the implied model (7), not model (4) itself, so a passing p-value does not rule out misspecification of how $Y$ enters the verification model.
- Editorial: in applications, one should inspect how the AUC estimate changes with the assumed value of $\beta$, since $\beta$ is identified through the nonlinear shape of the offset rather than through direct observation of $Y$ in unverified patients.
- Editorial: the NACC comparison suggests that treating autopsy-based verification as ignorable can shift the MMSE AUC from 0.584 (IG) to 0.791, a difference large enough to change clinical interpretation if it replicates in other cohorts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper considers estimation of the ROC curve and AUC when disease status is missing not at random. The authors assume logistic models for the disease status conditional on verification and for the verification indicator, establish identifiability without an instrumental variable (Theorem 1), propose a maximum likelihood estimator for the model parameters and weighted empirical distribution estimators for F0 and F1, and derive asymptotic normality of the resulting AUC and ROC estimators (Theorem 2). They also propose a two-step goodness-of-fit procedure for model checking. The methodology is evaluated in simulations and applied to the NACC Alzheimer's disease data.
Significance. If the results are correct, the paper makes a useful contribution to diagnostic accuracy studies with non-ignorable verification bias. Its estimands are defined through interpretable models, and the likelihood-based approach uses all observed data, leading to substantially smaller MSEs than the IPW competitor in the reported simulations. The paper is transparent in reporting undercoverage under model misspecification, and the main derivation structure (Bayes identity, iterated expectation, likelihood factorization) is internally consistent. The identifiability result without an instrumental variable is an important theoretical clarification. However, the universal statement of Theorem 2(b) needs correction, and the model verification procedure requires further justification.
major comments (2)
- [Theorem 2(b) and Appendix, Conditions C1–C6] Theorem 2(b) claims that sqrt(n)(ROC_hat(s) - ROC(s)) is asymptotically normal for each s in (0,1) under Conditions C1–C6, but these conditions do not ensure that the quantile ξ_{1-s} = F0^{-1}(1-s) is a point at which f0 is positive or at which F1 is differentiable. The influence-function expression for σ_s^2 in Theorem 2(b) contains the ratio f1(ξ_{1-s})/f0(ξ_{1-s}); if f0(ξ_{1-s}) = 0, as can happen when X has a continuous distribution with a gap in its support (e.g., X ~ 0.5·U(0,1) + 0.5·U(2,3) and 1-s is chosen so that the quantile falls in the gap), the quantile map F0^{-1} is not Hadamard-differentiable at 1-s and the asymptotic normality result need not hold. This is an internal gap between the theorem's universal quantifier over s and the listed conditions. I recommend either adding a condition such as f0(x) > 0 and f1(x) < ∞ on a neighborhood of ξ_{1-s} for each s considered, or restricting the statement of Theorem 2(b) to s for which such a condition holds.
- [Section 2.5, model verification] The proposed goodness-of-fit test for the verification model (4) uses T2/se(T2) as the test statistic and assumes a standard normal reference distribution, citing Hosmer et al. (1997). However, Hosmer et al. (1997) considered a fully observed logistic regression, whereas here the fitted probabilities π_i = π(X_i,V_i; μ_hat, φ_hat) come from the full likelihood with estimated parameters from both models, and the residuals are not those of a standard logistic regression. The asymptotic null distribution of T2/se(T2) is asserted without derivation or simulation evidence; the bootstrap is used only to estimate the standard error, not to calibrate the p-value. If the reference distribution is incorrect, the p-values in the real data application (Section 5) would be invalid. I recommend deriving the asymptotic distribution of a suitable test statistic or using a bootstrap calibration of the p-value.
minor comments (6)
- [Section 2.2, equation (5)] The notation P(X=x, V=v, Y=y, R=0) in equation (5) conflates probability mass functions with densities for the continuous variables X and V; this should be clarified, for instance by using f_{X,V|...} or a generic density notation.
- [Section 2.5, definition of T2] There is a typographical error in the definition of T2: the expression should be Σ_i ((R_i − π_i)^2 − π_i(1 − π_i)) with a matching closing parenthesis, but the text has a mismatched brace.
- [Tables 6 and 7] The label "Martial status" should be "Marital status" in Tables 6 and 7.
- [Table 8 and Section 5] The IPW 95% confidence interval for AUC is reported as (−1.289, 2.778), which lies outside the [0,1] support of AUC; this suggests numerical instability of the IPW variance estimator, and the paper should either discuss this or present the interval on a truncated scale.
- [Supplementary material] The proofs of Proposition 1, Theorem 1, and Theorem 2 are deferred to a supplementary file that is not included in the arXiv posting; for a journal submission, the supplementary material should be provided for review so that the derivations can be verified.
- [Figure 1 caption] The text says the confidence band is for s ∈ (0.05, 0.3), but the figure caption states s ∈ [0.05, 0.3]; the endpoints should be made consistent.
Circularity Check
No significant circularity: ROC/AUC estimators are plug-in functions of likelihood-based parameter estimates, not refitted to ROC/AUC values.
full rationale
The derivation chain is self-contained. The likelihood (Section 2.3) targets only (mu, phi) through models (3) and (4); it never uses ROC/AUC values as fitting targets. The ROC and AUC estimates in Section 2.4 are obtained by applying Proposition 1's identities E{g(X,V,R;mu,beta) I(X<=x)} = E{g(...)} F1(x) and E{(1-g(...)) I(X<=x)} = E{1-g(...)} F0(x) to the MLE, then plugging the resulting weighted empirical CDFs into the definitions ROC(s)=1-F1(F0^{-1}(1-s)) and AUC=int F0 dF1. No parameter is fitted to a subset of ROC/AUC data and then 'predicted'. The self-citations (Hu et al. 2023; Liu et al. 2022) are contextual and do not carry the argument: the identifiability proof is the authors' own Theorem 1 (deferred to supplementary), and the asymptotic proof is Theorem 2 with regularity conditions C1-C6. The two-step verification of model (4) explicitly acknowledges that (4) cannot be validated directly and only checks the implied model (7); that is an honest limitation, not a circular step. There is no equation in which the claimed output equals its input by construction. Concerns about the missing f0>0/differentiability condition in Theorem 2(b) are a correctness or regularity gap, not circularity.
Assumptions & free parameters
free parameters (1)
- Kernel bandwidths h0 and h1 for density estimates of f0 and f1 =
1.06 n^{-1/5} min(sigma_hat, IQR/1.34)
assumptions (4)
- domain assumption Disease model (3): P(Y=1|X,V,R=1) is logistic with linear index mu1+mu2*x+mu3^T v.
- domain assumption Verification model (4): P(R=1|Y,X,V) is logistic with linear index psi1+psi2*x+psi3^T v+beta*y.
- domain assumption Condition C1: X continuous, (1,X,V) stochastically linearly independent, mu2 != 0.
- standard math Regularity conditions C2-C5 for MLE asymptotic normality.
Cite this review
Pith. "Pith review of Receiver operating characteristic curve analysis with non-ignorable missing disease status." pith.science (2026). https://pith.science/paper/W6GGSOP2
@misc{pith2026241117402,
author = {Pith},
title = {Pith review of: Receiver operating characteristic curve analysis with non-ignorable missing disease status},
year = {2026},
howpublished = {\url{https://pith.science/paper/W6GGSOP2}},
note = {Machine review of arXiv:2411.17402}
}
read the original abstract
This article considers the receiver operating characteristic (ROC) curve analysis for medical data with non-ignorable missingness in the disease status. In the framework of the logistic regression models for both the disease status and the verification status, we first establish the identifiability of model parameters, and then propose a likelihood method to estimate the model parameters, the ROC curve, and the area under the ROC curve (AUC) for the biomarker. The asymptotic distributions of these estimators are established. Via extensive simulation studies, we compare our method with competing methods in the point estimation and assess the accuracy of confidence interval estimation under various scenarios. To illustrate the application of the proposed method in practical data, we apply our method to the National Alzheimer's Coordinating Center data set.
Figures
Forward citations
Cited by 1 Pith paper
-
Estimation with missing not at random binary outcomes via exponential tilts
Binary-outcome MNAR estimation is made identifiable through an exponential tilt model with a sufficient identifiability condition, estimated by KL matching without instruments or shadow variables.
Reference graph
Works this paper leans on
-
[1]
Begg, C. B. and Greenes, R. A. (1983). Assessment of diagnostic tests when disease verification is subject to selection bias. Biometrics , 39:207--215
work page 1983
-
[2]
Fluss, R., Reiser, B., Faraggi, D., and Rotnitzky, A. (2009). Estimation of the ROC curve under verification bias. Biometrical Journal , 51:475--490
work page 2009
-
[3]
Folstein, M. F., Folstein, S. E., and McHugh, P. R. (1975). “ M ini-mental state”: a practical method for grading the cognitive state of patients for the clinician. Journal of Psychiatric Research , 12:189--198
work page 1975
-
[4]
W., Hosmer, T., Le Cessie, S., and Lemeshow, S
Hosmer, D. W., Hosmer, T., Le Cessie, S., and Lemeshow, S. (1997). A comparison of goodness-of-fit tests for the logistic regression model. Statistics in Medicine , 16:965--980
work page 1997
-
[5]
Hu, D., Yuan, M., Yu, T., and Li, P. (2023). Statistical inference for the two-sample problem under likelihood ratio ordering, with application to the ROC curve estimation. Statistics in Medicine , 42:3649--3664
work page 2023
-
[6]
Liu, D. and Zhou, X. (2010). A model for adjusting for nonignorable verification bias in estimation of ROC curve and its area with likelihood-based approach. Biometrics , 66:1119--1128
work page 2010
-
[7]
Liu, Y., Li, P., and Qin, J. (2022). Full-semiparametric-likelihood-based inference for non-ignorable missing data. Statistica Sinica , 32:271--292
work page 2022
-
[8]
Tombaugh, T. N. and McIntyre, N. J. (1992). The mini-mental state examination: a comprehensive review. Journal of the American Geriatrics Society , 40:922--935
work page 1992
Show all 14 references
-
[9]
and Tian, L
Yin, J. and Tian, L. (2014). Joint inference about sensitivity and specificity at the optimal cut-off point associated with youden index. Computational Statistics & Data Analysis , 77:1--13
2014
-
[10]
K., and Park, T
Yu, W., Kim, J. K., and Park, T. (2018). Estimation of area under the ROC curve under nonignorable verification bias. Statistica Sinica , 28:2149--2166
2018
-
[11]
Yuan, M., Li, P., and Wu, C. (2021). Semiparametric inference of the youden index and the optimal cut-off point under density ratio models. The Canadian Journal of Statistics , 49:965--986
2021
-
[12]
Zhou, X. (1998). Comparing correlated areas under the ROC curves of two diagnostic tests in the presence of verification bias. Biometrics , 54:453--470
1998
-
[13]
and Castelluccio, P
Zhou, X. and Castelluccio, P. (2004). Adjusting for non-ignorable verification bias in clinical studies for alzheimer's disease. Statistics in Medicine , 23:221--230
2004
-
[14]
Zhou, X., McClish, D., and Obuchowski, N. (2011). Statistical Methods in Diagnostic Medicine. New York: Wiley
2011
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.