{"id":"b5d48464-e463-4c3e-bdf4-1b864b148029","arxiv_id":"2411.17402","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A maximum-likelihood approach estimates ROC curves and AUC under non-ignorable missing disease status, with identifiability achieved without an instrumental variable.","lead":"This paper proposes a likelihood-based method for estimating the diagnostic accuracy of a biomarker (ROC curve and AUC) when many patients' true disease status is missing, and the missingness depends on the disease itself. The method is shown to be more efficient than an existing inverse-probability-weighting approach and is applied to Alzheimer's disease data, yielding a tight confidence interval for the MMSE test's accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2(b) claims root-n normality of ROC(s) for all s in (0,1), but Conditions C1-C6 never require f0 > 0 at the quantile or differentiability of F1 there; if the quantile falls in a zero-density stretch, Wald intervals (11) can fail even under correct specification.","rationale":"The reader's weakest assumption was that the verification model (4) is untestable and that the two-step goodness-of-fit check only validates the implied model (7). That is a legitimate concern about practical applicability and about the strength of the model-checking contribution. My concern is different but related: even under exact correct specification, the stated regularity conditions do not appear to support the full scope of Theorem 2(b). The proof is in an unavailable supplementary file, and the Appendix lists no density-positivity or differentiability condition for F0 and F1 at the quantile of interest. This matters because the ROC curve estimator at a fixed s is a functional of the empirical quantile F0^{-1}(1-s) and of F1; the delta method / Hadamard differentiability argument needs f0(xi_{1-s}) > 0. If the quantile falls in a flat stretch of F0 or at the boundary of the support, root-n normality can fail or the plug-in variance estimator can be unstable. The proposed concrete simulation would settle whether this gap is real. Since the reader already assigned CONDITIONAL, my finding does not move the verdict; it sharpens the reason for the conditionality: the authors should either add the missing regularity conditions to Theorem 2(b) or restrict the claim to s where f0(xi_{1-s}) > 0, and they should make the supplementary proof available for verification.","tokens_in":13137,"tokens_out":25854,"duration_ms":257519,"concrete_test":"Simulate a correct-specification setting in which X has density vanishing on an interior interval, e.g., X ~ 0.5*U(0,1) + 0.5*U(2,3), with V absent, Y generated from model (3), R from model (4) with beta != 0, n = 10,000, and B = 1,000 replications. Compute ROC_hat(s) and the Wald interval (11) for s = 0.5, so that xi_{1-s} lies in the zero-density gap. If the empirical coverage is not near 0.95 or the Q-Q plot of sqrt(n) ROC_hat(s) shows clear non-normality, Theorem 2(b) needs an additional condition such as f0(xi_{1-s}) > 0 and differentiability of F1 at xi_{1-s}. Alternatively, inspect the supplementary proof: if it invokes such a condition, the main text should state it explicitly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of Theorem 2(b) is that sqrt(n)(ROC_hat(s) - ROC(s)) is asymptotically normal for each s in (0,1) under Conditions C1-C6. The proof is deferred to a supplementary file not included in the arXiv posting, so the only verifiable regularity conditions are those in the Appendix. Conditions C2-C6 require compactness, finite moments, positive definite Fisher information, and third-order smoothness of the likelihood; Condition C1 requires X continuous, stochastic linear independence of (1,X,V), and mu2 nonzero. None of these conditions states that the quantile xi_{1-s} = F0^{-1}(1-s) is a point where f0 is positive or where F1 is differentiable. Yet the influence-function formula for sigma_s^2 in Theorem 2(b) explicitly uses the ratio f1(xi_{1-s})/f0(xi_{1-s}). If f0(xi_{1-s}) = 0, which can happen for a continuous X whose density vanishes on an interior interval (e.g., X ~ 0.5*U(0,1) + 0.5*U(2,3) with s chosen so that xi_{1-s} lies in the gap), the quantile map F0^{-1} is not Hadamard-differentiable at 1-s, the influence-function expansion breaks down, and the Wald interval (11) is not guaranteed to attain its claimed asymptotic coverage even when models (3)-(4) are correctly specified. This is an internal gap between the theorem's universal quantifier over s and the listed conditions, not a disagreement with external consensus. The AUC result in Theorem 2(a) is not affected because it involves no density ratio.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper considers estimation of the ROC curve and AUC when disease status is missing not at random. The authors assume logistic models for the disease status conditional on verification and for the verification indicator, establish identifiability without an instrumental variable (Theorem 1), propose a maximum likelihood estimator for the model parameters and weighted empirical distribution estimators for F0 and F1, and derive asymptotic normality of the resulting AUC and ROC estimators (Theorem 2). They also propose a two-step goodness-of-fit procedure for model checking. The methodology is evaluated in simulations and applied to the NACC Alzheimer's disease data.","tokens_in":13477,"tokens_out":7084,"duration_ms":72177,"significance":"If the results are correct, the paper makes a useful contribution to diagnostic accuracy studies with non-ignorable verification bias. Its estimands are defined through interpretable models, and the likelihood-based approach uses all observed data, leading to substantially smaller MSEs than the IPW competitor in the reported simulations. The paper is transparent in reporting undercoverage under model misspecification, and the main derivation structure (Bayes identity, iterated expectation, likelihood factorization) is internally consistent. The identifiability result without an instrumental variable is an important theoretical clarification. However, the universal statement of Theorem 2(b) needs correction, and the model verification procedure requires further justification.","major_comments":[{"comment":"Theorem 2(b) claims that sqrt(n)(ROC_hat(s) - ROC(s)) is asymptotically normal for each s in (0,1) under Conditions C1–C6, but these conditions do not ensure that the quantile ξ_{1-s} = F0^{-1}(1-s) is a point at which f0 is positive or at which F1 is differentiable. The influence-function expression for σ_s^2 in Theorem 2(b) contains the ratio f1(ξ_{1-s})/f0(ξ_{1-s}); if f0(ξ_{1-s}) = 0, as can happen when X has a continuous distribution with a gap in its support (e.g., X ~ 0.5·U(0,1) + 0.5·U(2,3) and 1-s is chosen so that the quantile falls in the gap), the quantile map F0^{-1} is not Hadamard-differentiable at 1-s and the asymptotic normality result need not hold. This is an internal gap between the theorem's universal quantifier over s and the listed conditions. I recommend either adding a condition such as f0(x) > 0 and f1(x) < ∞ on a neighborhood of ξ_{1-s} for each s considered, or restricting the statement of Theorem 2(b) to s for which such a condition holds.","section":"Theorem 2(b) and Appendix, Conditions C1–C6"},{"comment":"The proposed goodness-of-fit test for the verification model (4) uses T2/se(T2) as the test statistic and assumes a standard normal reference distribution, citing Hosmer et al. (1997). However, Hosmer et al. (1997) considered a fully observed logistic regression, whereas here the fitted probabilities π_i = π(X_i,V_i; μ_hat, φ_hat) come from the full likelihood with estimated parameters from both models, and the residuals are not those of a standard logistic regression. The asymptotic null distribution of T2/se(T2) is asserted without derivation or simulation evidence; the bootstrap is used only to estimate the standard error, not to calibrate the p-value. If the reference distribution is incorrect, the p-values in the real data application (Section 5) would be invalid. I recommend deriving the asymptotic distribution of a suitable test statistic or using a bootstrap calibration of the p-value.","section":"Section 2.5, model verification"}],"minor_comments":[{"comment":"The notation P(X=x, V=v, Y=y, R=0) in equation (5) conflates probability mass functions with densities for the continuous variables X and V; this should be clarified, for instance by using f_{X,V|...} or a generic density notation.","section":"Section 2.2, equation (5)"},{"comment":"There is a typographical error in the definition of T2: the expression should be Σ_i ((R_i − π_i)^2 − π_i(1 − π_i)) with a matching closing parenthesis, but the text has a mismatched brace.","section":"Section 2.5, definition of T2"},{"comment":"The label \"Martial status\" should be \"Marital status\" in Tables 6 and 7.","section":"Tables 6 and 7"},{"comment":"The IPW 95% confidence interval for AUC is reported as (−1.289, 2.778), which lies outside the [0,1] support of AUC; this suggests numerical instability of the IPW variance estimator, and the paper should either discuss this or present the interval on a truncated scale.","section":"Table 8 and Section 5"},{"comment":"The proofs of Proposition 1, Theorem 1, and Theorem 2 are deferred to a supplementary file that is not included in the arXiv posting; for a journal submission, the supplementary material should be provided for review so that the derivations can be verified.","section":"Supplementary material"},{"comment":"The text says the confidence band is for s ∈ (0.05, 0.3), but the figure caption states s ∈ [0.05, 0.3]; the endpoints should be made consistent.","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The central methodological idea is sound, and the simulation study is honest about behavior under misspecification. The main concern is the gap in the statement of Theorem 2(b) regarding density positivity; this is fixable by adding a regularity condition or restricting the range of s. The goodness-of-fit test also needs stronger justification. The authors should be asked to provide the supplementary material with the revision so that the proofs can be checked. I do not see evidence of circularity or deliberate misrepresentation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Briefly: this paper is worth engaging. It moves past Yu et al.'s IPW estimator by establishing identifiability of the two logistic models without an instrumental variable and by fitting both models through maximum likelihood, using all observed data. Simulations show large MSE reductions relative to IPW, and the authors honestly report undercoverage under misspecification. The model setup is the same as Yu et al. (2018), which is not a flaw since the contribution is identification and estimation, but the paper should say more clearly what is genuinely new relative to Liu et al. (2022).\n\nThe likelihood factorization and the Bayes identity leading to equation (9) check out, and Proposition 1 is a straightforward iterative-expectation identity. The ROC/AUC estimators are built from fitted conditional probabilities plus weighted empirical CDFs; there is no circularity in the usual sense.\n\nSoft spots, in order of size. First, the stress-test concern about Theorem 2(b) holds up. The stated conditions C1--C6 never require f0 to be positive at the quantile xi_{1-s} = F0^{-1}(1-s), yet the variance formula is the ratio f1(xi_{1-s})/f0(xi_{1-s}). For a biomarker density that vanishes on an interior interval, the quantile map is not Hadamard differentiable and the Wald intervals (11) are not guaranteed to attain nominal coverage even under correct specification. This is fixable: add a condition like f0(F0^{-1}(1-s)) > 0 and continuity of f1 at that point, or restrict the theorem to values of s where this holds. It does not affect Theorem 2(a) for AUC.\n\nSecond, all proofs are in a supplementary file that is not included in the arXiv posting. As a reviewer I would need that file before signing off. Third, the goodness-of-fit test for model (4) is asserted to follow the Hosmer et al. construction, with N(0,1) as reference distribution, but no derivation of the test statistic's null asymptotics is given. The bootstrap standard error is plausible but needs a statement about why the null distribution is standard normal. Fourth, no code or data are shipped; the simulation tables are detailed enough that this is a minor issue, but a public implementation would help.\n\nOverall the central identification and estimation strategy looks sound, the simulations are honest, and the real-data example is appropriate. Citation pattern is fine; self-citations are to the methods being extended. This is a serious paper for a serious referee. I would ask for the supplementary proofs, the added density condition, and a clearer demarcation from Liu et al. (2022), then expect it to go through.","headline":"This paper is a solid, mostly-correct step forward for ROC/AUC with non-ignorable missing disease status; it deserves review, but Theorem 2(b) needs an added positivity condition on f0 at the quantile.","tokens_in":14025,"tokens_out":3686,"would_cite":true,"duration_ms":36484,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that the ROC curve and AUC of a biomarker can be estimated and tested with valid confidence intervals when disease status is missing non-ignorably, no instrumental variable required.","keywords":["ROC curve","AUC","non-ignorable missing data","verification bias","identifiability","maximum likelihood","asymptotic normality","Alzheimer's disease"],"falsifier":"Generate data under models (3)-(4) with $\\mu_2=0$ and maximize the observed-data likelihood from many starting values; if the maximum is not unique (a flat ridge in parameter space), Theorem 1's identifiability claim fails outside Condition C1. Alternatively, obtain gold-standard $Y$ for a random subsample of patients with $R=0$ and compare the observed $R=1$ rates in $(Y,X,V)$ cells with the fitted probabilities from model (4); systematic mismatch would falsify the verification model.","tokens_in":12902,"feed_emoji":"📊","tokens_out":11617,"duration_ms":95885,"temperature":0.7,"pith_summary":"The paper addresses a common clinical setting: disease status is verified only for some patients, and the decision to verify depends on the unobserved disease status itself, so standard ROC and AUC estimates are biased. It works with two logistic models, one for disease given biomarker and covariates among verified patients and one for verification given disease, biomarker, and covariates. The central result is that these models are identifiable without an instrumental variable, provided the biomarker coefficient in the disease model is nonzero and the regressors are not collinear. The authors then build a full maximum-likelihood estimator for the parameters, the ROC curve, and the AUC, prove asymptotic normality, and demonstrate in simulations that the estimator is more precise than inverse-probability weighting. An application to Alzheimer's disease data shows an MMSE AUC of 0.791 with a narrow confidence interval, and a significant non-ignorability coefficient.","feed_headline":"ROC estimation works without an instrument for missing disease status","feed_subtitle":"Full likelihood identifies the non-ignorable missingness model and gives valid confidence intervals for AUC.","key_machinery":"The load-bearing identity is model (7), which re-expresses the verification propensity as $P(R=1|x,v) = 1/(1+\\exp\\{\\psi_1+\\psi_2 x+\\psi_3^T v + c(x,v;\\mu,\\beta)\\})$, where the offset $c(x,v;\\mu,\\beta) = \\log E(e^{\\beta Y}|x,v,R=1)$ is computable from the disease model. This identity converts an unobserved-$Y$ problem into a logistic regression on observed data and is what makes identifiability without an instrumental variable plausible. The second piece is Proposition 1: for weights $g(X,V,R;\\mu,\\beta) = P(Y=1|X,V,R)$ and $1-g$, the identities $E\\{g\\,I(X\\le x)\\} = E\\{g\\}\\,F_1(x)$ and $E\\{(1-g)I(X\\le x)\\} = E\\{1-g\\}\\,F_0(x)$ let the biomarker distributions in healthy and diseased groups be estimated by weighted empirical distribution functions without modelling the biomarker distribution. These two pieces together carry the estimation of $\\mathrm{ROC}(s)$ and AUC.","core_discovery":"On its own terms, the paper claims that the two logistic models (3) and (4) — disease status among verified patients, and verification status given disease, biomarker, and covariates — are identifiable from observed data alone, with no instrumental variable, as long as the biomarker is continuous, the regressors are stochastically linearly independent, and the biomarker coefficient $\\mu_2$ in the disease model is nonzero (Theorem 1). The observable verification propensity $P(R=1|x,v)$ is shown to be logistic with an offset $c(x,v;\\mu,\\beta) = \\log E(e^{\\beta Y}|x,v,R=1)$, and identifiability is established through that representation. The full-likelihood estimator is asymptotically normal, and the plug-in estimators of $\\mathrm{ROC}(s)$ and AUC inherit that property, so Wald intervals with plug-in variances have correct coverage under correct specification (Theorem 2). Proposition 1 supplies the weighted empirical distribution identities used to turn the fitted models into ROC and AUC estimates.","pith_inferences":["Editorial: the offset-logistic identification argument should extend to other binary links, such as probit, and to ordinal disease categories, since only the induced form of the verification propensity and a nonzero biomarker slope are needed.","Editorial: the paper's two-step goodness-of-fit check validates the implied model (7), not model (4) itself, so a passing p-value does not rule out misspecification of how $Y$ enters the verification model.","Editorial: in applications, one should inspect how the AUC estimate changes with the assumed value of $\\beta$, since $\\beta$ is identified through the nonlinear shape of the offset rather than through direct observation of $Y$ in unverified patients.","Editorial: the NACC comparison suggests that treating autopsy-based verification as ignorable can shift the MMSE AUC from 0.584 (IG) to 0.791, a difference large enough to change clinical interpretation if it replicates in other cohorts."],"forward_implications":["Researchers no longer need to search for an instrumental variable before applying the method; identifiability holds under Condition C1 alone.","Using all observed patients, including those with unverified disease status, the full-likelihood AUC estimator has smaller mean squared error than the IPW estimator in the paper's simulations.","Wald confidence intervals for AUC and $\\mathrm{ROC}(s)$ achieve coverage close to the nominal 95% level when models (3) and (4) are correctly specified.","The two-step goodness-of-fit procedure provides a practical check of both models despite the missingness of $Y$ in unverified patients.","On the NACC Alzheimer's data, the method yields MMSE AUC 0.791 with 95% CI (0.779, 0.803), and the estimated non-ignorability parameter $\\beta$ is significant, whereas the IPW confidence interval is extremely wide."],"supporting_citations":[{"why":"Supplies the two logistic models for disease and verification status and the inverse-probability-weighting estimator that the proposed full-likelihood method extends and outperforms.","marker":"Yu et al. (2018)"},{"why":"Establishes the likelihood-based approach for non-ignorable verification bias in ROC/AUC estimation and provides the covariate set used in the NACC analysis.","marker":"Liu and Zhou (2010)"},{"why":"Provides the unweighted sum-of-squares goodness-of-fit test statistic that the two-step model verification procedure adapts.","marker":"Hosmer et al. (1997)"},{"why":"Gives the definitions of the ROC curve and AUC and the classical framework the paper builds on.","marker":"Zhou et al. (2011)"},{"why":"Introduces verification-bias adjustment under missing at random, the assumption the paper argues is often violated.","marker":"Begg and Greenes (1983)"},{"why":"Documents the two-stage Alzheimer's study design that motivates non-ignorable missing disease status.","marker":"Zhou and Castelluccio (2004)"}],"fun_headline_variants":["No instrument needed for ROC with missing disease","Full likelihood identifies ROC despite missing disease","ROC identifiability proven without instrumental variable","Likelihood-based ROC handles non-ignorable missingness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the verification model (4) is correctly specified as a logistic regression on the biomarker, covariates, and the unobserved disease status $Y$, since $Y$ is missing whenever verification is skipped and the data cannot directly check this assumption; identifiability also requires the biomarker coefficient in the disease model to be nonzero.","fun_headline_variants_meta":{"raw":{"variants":["No instrument needed for ROC with missing disease","Full likelihood identifies ROC despite missing disease","ROC identifiability proven without instrumental variable","Likelihood-based ROC handles non-ignorable missingness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001145,"raw_usage":{"total_tokens":4710,"prompt_tokens":868,"completion_tokens":3842,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":3785}},"tokens_in":484,"tokens_out":3842,"duration_ms":26609,"temperature":1.0,"reasoning_tokens":3785,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:09:10.443212+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate data under models (3)-(4) with $\\mu_2=0$ and maximize the observed-data likelihood from many starting values; if the maximum is not unique (a flat ridge in parameter space), Theorem 1's identifiability claim fails outside Condition C1. Alternatively, obtain gold-standard $Y$ for a random subsample of patients with $R=0$ and compare the observed $R=1$ rates in $(Y,X,V)$ cells with the fitted probabilities from model (4); systematic mismatch would falsify the verification model.","supporting_citations":[{"cited_title":"and Zhou, X","cited_arxiv_id":null,"evidence_quote":"Establishes the likelihood-based approach for non-ignorable verification bias in ROC/AUC estimation and provides the covariate set used in the NACC analysis."},{"cited_title":"W., Hosmer, T., Le Cessie, S., and Lemeshow, S","cited_arxiv_id":null,"evidence_quote":"Provides the unweighted sum-of-squares goodness-of-fit test statistic that the two-step model verification procedure adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the definitions of the ROC curve and AUC and the classical framework the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces verification-bias adjustment under missing at random, the assumption the paper argues is often violated."},{"cited_title":"and Castelluccio, P","cited_arxiv_id":null,"evidence_quote":"Documents the two-stage Alzheimer's study design that motivates non-ignorable missing disease status."}],"review_version":1}