REVIEW 3 major objections 5 minor 2 cited by
Efficient Estimation of Causal Effects Under Two-Phase Sampling with Error-Prone Outcome and Treatment Measurements
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper establishes that two one-step doubly robust estimators of the average treatment effect under biased two-phase validation sampling are both asymptotically normal and semiparametrically efficient, and that an…
desk verdict Worth reviewing: the EIC contribution and finite-sample fixes are real, but the appendix proofs need a serious rewrite before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are two efficient influence curves for the treatment-arm mean $\psi_a = E[Y(a)]$: the Approach 1 curve $\phi_a(O,P)$ from Theorem 2, built from the imputation models $\lambda_a(Z)$ and $\mu_a(Z)$, the marginalized outcome $\eta_a(X)$, the imputed propensity $\pi_a(X)$, and the sampling probability $\kappa(Z)$; and the Approach 2 curve $\phi^{\mathrm{ALT}}_a(O;P)$ of Proposition 1, obtained by mapping the complete-data curve $\chi_a(O;P_C)$ through the selection mechanism, with the pseudo-regression $\varphi_a(Z) = E[\chi_a(O;P_C)\mid Z,R=1]$ absorbing the complicated composite nuisance function. These curves are doing the work: they supply the one-step bias corrections, define the efficiency bound $E[\phi_a(O,P)^2]$, and make visible the two distinct finite-sample sources of instability, namely multiplicative weighting terms in Approach 1 and difficult estimation of $\varphi_a(Z)$ in Approach 2.
What would settle it
Run a simulation with a deliberately misspecified full-data outcome model and a correctly specified propensity model, so that $\|\hat m_a - m_a\|$ is bounded away from zero while $\|\hat g_a - g_a\|$ shrinks; if $\hat\psi^{\mathrm{OS},2}_a$ still achieves $\sqrt{n}$-normal coverage with variance at the efficiency bound, the product-rate condition of Theorem 4 is not necessary. Alternatively, in the HIV data application, hold the phase-two sample at a fixed small size while growing the phase-one sample and check whether confidence-interval coverage degrades when the estimated product $\|\hat m_a - m_a\|\,\|\hat g_a - g_a\|$ is not small.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the efficient influence curve for an average treatment effect under two-phase sampling with error-prone outcome and treatment can be written in two ways: either directly from the observed-data functional, or, through Proposition 1, as the complete-data influence curve $\chi_a(O;P_C)$ weighted by sampling probabilities $\kappa(Z)$ minus a projection $\varphi_a(Z)$. Both representations yield one-step estimators, $\hat\psi^{\mathrm{OS},1}_a$ and $\hat\psi^{\mathrm{OS},2}_a$, that are asymptotically equivalent, $\sqrt{n}$-consistent, and variance-bound attaining under the stated rate conditions, but the second requires estimating fewer nuisance functions and has less stringent consistency conditions. The paper further claims that the complex pseudo-outcome regression $\varphi_a(Z)$ can be estimated by empirical efficiency maximization to restore finite-sample efficiency, and that a weighted ensemble of the two estimators converges to the same influence curve with variance no larger than either component.
Load-bearing premise
The efficiency guarantee depends on the product of the full-data outcome regression error and the full-data propensity score error being $o_P(1/\sqrt{n})$ in Theorem 4, with analogous rates in Theorem 3, a condition the paper does not verify for phase-two samples as small as roughly one hundred patients in the HIV application.
Editorial extensions
If this is right
- With known sampling probabilities $\kappa(Z)$, the Approach 2 estimator requires only the usual full-data product-rate condition $\|\hat m_a - m_a\|\,\|\hat g_a - g_a\|=o_P(n^{-1/2})$ to be $\sqrt{n}$-consistent and efficient, so two-phase estimation is no harder than complete-data estimation under that condition.
- Estimators built from the Approach 2 influence curve take the same form across two-phase problems with different partially missing variables, so the empirical-efficiency-maximization modification transfers to other missing-data settings.
- The one-step Approach 1 estimator requires a stronger set of rate conditions, including $o_P(n^{-1/4})$ estimation of the imputed propensity score $\pi_a(X)$, and in practice it can suffer instability from high-degree weight terms.
- The proposed ensemble $\hat\psi^{\mathrm{OS},W}_a$, with weights chosen by empirical variances and covariance, retains the shared limiting normal distribution with variance $E[\phi_a(O,P)^2]$, and in simulations its RMSE is among the lowest across every setting considered.
Reading between the lines
- A testable extension beyond the paper is to apply the empirical-efficiency-maximization-modified Approach 2 estimator to other two-phase designs, such as missing confounders or missing exposure, because the same $\varphi_a(Z)$ pseudo-regression structure appears whenever the complete-data influence curve is mapped through the selection mechanism.
- Because the ensemble weight converges to $1/2$ when both estimators are efficient, the practical value of the ensemble lies almost entirely in small samples, where post-hoc model selection would otherwise invalidate inference; this suggests that guidance on when the finite-sample gains persist could matter more than the asymptotic statement.
- If the product-rate condition of Theorem 4 fails but $\hat\psi^{\mathrm{OS},2}_a$ still appears approximately unbiased in applications with phase-two samples of a few hundred patients, that would indicate the rate condition is sufficient rather than necessary, broadening the method's range.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper considers two-phase sampling designs in which the outcome and treatment are measured with error, gold-standard values are obtained only for a validation subsample, and selection into validation may depend on error-prone measurements. It derives two one-step, doubly robust estimators of the average treatment effect: one based on the efficient influence curve of the observed-data functional (Approach 1), and one based on a complete-data influence curve representation (Approach 2). The authors show that the two estimators are asymptotically equivalent and, under rate conditions on nuisance estimates, √n-consistent and semiparametric efficient. They propose an empirical efficiency maximization (EEM) modification for estimating the pseudo-outcome regression φ_a, and a weighted ensemble of the two one-step estimators. Finite-sample behavior is studied in simulations and in a plasmode analysis of HIV patients from the Vanderbilt Comprehensive Care Clinic, with the ensemble estimator performing well and the EEM modification improving efficiency of the Approach 2 estimator.
Significance. If the theoretical claims are correct, the paper makes a useful practical contribution by unifying two previously disconnected derivations of efficient estimators in two-phase sampling, and by offering finite-sample modifications that appear to help in small validation samples. The open-source R package and reproducible simulation code are strengths, and the VCCC application demonstrates the methods in a realistic EHR setting. However, the central efficiency theorems rest on appendix proofs that are not currently self-contained and contain internal errors; the paper's asymptotic claims therefore need a complete, corrected derivation before the contribution can be fully evaluated. The simulation evidence is supportive, but it does not by itself establish the asymptotic efficiency claims.
major comments (3)
- [Appendix A.3, proof of Theorem 3] The proof of Theorem 3 is not self-contained and cannot be verified as written. In the decomposition (A1), the second-order term T2 is dismissed with 'analogous arguments made in the proof of Theorem 3,' which is circular because this is the proof of Theorem 3. The key bias decomposition (A2) is imported from Levis et al. (2024a) without being stated as a lemma or derived, and the subsequent 'full data bias' assertion is used to conclude that the displayed product rates are the only remainder terms to control. Because Theorem 3 is the basis for the Approach 1 efficiency claim, this gap is load-bearing; the authors should provide a complete derivation of (A2) and of T2, or state (A2) as a lemma with proof.
- [Appendix A.4, proof of Theorem 4] The proof of Theorem 4 contains errors that prevent verification of the central efficiency claim. In (A3), the first term should be (Pn−P)φ_a, not (Pn−P)ψ_a, and the middle term uses ψ̂_a where φ̂_a is meant. More seriously, the final displayed equation sets the scalar bias term ψPI,1_a(P̂C)+E[χ_a(O,P̂C)]−Ψ_a(PC) equal to the random variable ma(W)+I(A=a)/πa(W)(Y−ma(W)); a scalar cannot equal an influence-curve random variable, and the argument switches from Approach 2 to Approach 1 and uses W as the conditioning variable where X is required. As written, the second-order remainder is not derived from the explicit influence curve (7). The authors should redo this proof from φ_a^ALT and show that the only remainder terms are ||m̂_a−m_a||·||ĝ_a−g_a|| and ||φ̂_a−φ_a||·||κ̂−κ||, with all other cross-products such as ||λ̂_a−λ_a||·||μ̂_a−μ_a|| being higher order or otherwise controlled.
- [Theorems 3-4 and Section 6] The asymptotic efficiency claims are conditional on the rate conditions in Theorems 3 and 4, and the paper does not verify these rates in the VCCC application, where phase-two samples can be as small as roughly 131 patients. The application is described as illustrative, so this is not by itself a fatal defect, but the paper should state explicitly that the efficiency guarantee in finite samples is not established and that the reported RMSEs are empirical rather than a verification of the asymptotic variance formula.
minor comments (5)
- [Table 1] The definition of μ_a(z) omits conditioning on A=a: it should be E_P[Y|Z=z,A=a,R=1], not E_P[Y|Z=z,R=1]. The proof and estimators use the conditioned version, so the table should be corrected.
- [Appendix A.4, Eq. (A3)] In addition to the issues raised in the major comments, the proof uses ψ̂_a in the second term where φ̂_a is intended, and the final display writes m_a(W) where m_a(X) is required.
- [Appendix A.5, proof of Theorem 5] The displayed equation after defining ŵ writes 'ŵ ψ̂OS,W_a + (1−ŵ)ψ̂OS,2_a'; this should be 'ŵ ψ̂OS,1_a + (1−ŵ)ψ̂OS,2_a'.
- [Theorem 3, conditions] Condition 2, ||π̂_a−π_a||=oP(n^{−1/4}), is not by itself sufficient for the product condition 1; the two conditions together imply a rate for η̂_a only if both are assumed. The statement would be clearer if the conditions were expressed directly as product rate conditions.
- [Section 4.3] The EEM weighted regression uses weights ((R_i/κ̂(Z_i))−1)^2, which can be very large when κ̂(Z_i) is small; the paper could note that practical implementations may need to stabilize these weights, particularly in the small phase-two samples considered.
Circularity Check
Central semiparametric derivation is self-contained; only internal circularity is a self-referential step in the proof of Theorem 3.
-
other
[Appendix A.3, Proof of Theorem 3, after equation (A1)]
"Above, the first term T1 is OP(1/√n) by the Central Limit Theorem, while the second term can be controlled by analogous arguments made in the proof of Theorem 3."
The proof is establishing Theorem 3 itself, so 'the proof of Theorem 3' does not exist as a prior argument. The control of T2=(Pn−P)(φ̂a−φa), which is needed for the oP(1/√n) remainder and hence for the efficiency claim, is deferred to the very theorem being proved. No independent bound for T2 is supplied in the appendix. The rest of the proof imports a bias decomposition from Levis et al. (2024a), so this circular sentence is the only stated justification for the T2 term. This is a genuine logical gap, though it does not make the theorem's conclusion equivalent to the estimator's definition by construction.
full rationale
The paper's core derivation is self-contained in the sense relevant to circularity: Theorem 2 derives the efficient influence curve ϕa via the Robins et al. mapping and iterated expectations, Proposition 1 follows from the same display, and Theorems 3-5 are standard one-step expansions with explicit rate conditions. No parameter is fitted to produce the claimed ATE, and the EEM step legitimately targets variance rather than the estimand. The only circular passage is the self-referential sentence in A.3, where control of T2 is deferred to 'the proof of Theorem 3'—the theorem being proved. This is a real proof gap, but the central result is not forced by definition: the rate conditions stated in Theorem 3 are substantive and not implied solely by the estimators' construction. Self-citations (e.g., Barnatchez et al. 2024 for the VCCC procedure, Hejazi et al. 2021 for Approach 2 background) are not load-bearing; Proposition 1 is proven in-house, and the data-application citation is incidental to the theoretical claims. Score 2 reflects one internal circular proof step with otherwise independent content.
Assumptions & free parameters
free parameters (1)
- δ (ensemble stabilization constant) =
small positive constant (value not specified)
assumptions (7)
- domain assumption Assumption 1 (SUTVA): Y = A Y(1) + (1-A) Y(0); no interference between units.
- domain assumption Assumption 2 (Positivity of treatment): 0 < P(A=1|X) < 1 for all X.
- domain assumption Assumption 3 (Unconfoundedness): Y(a) ⊥ A | X.
- domain assumption Assumption 4 (Outcome and exposure MAR): (Y,A) ⊥ R | Z.
- domain assumption Assumption 5 (Positivity of validation selection): 0 < κ(z) < 1 for all z.
- domain assumption Nuisance convergence rate conditions in Theorems 3 and 4 (e.g., product of L2 errors oP(1/√n)).
- standard math Regularity conditions for empirical process terms (sample splitting or Donsker class).
Cite this review
Pith. "Pith review of Efficient Estimation of Causal Effects Under Two-Phase Sampling with Error-Prone Outcome and Treatment Measurements." pith.science (2026). https://pith.science/paper/COVYMGUW
@misc{pith2026250621777,
author = {Pith},
title = {Pith review of: Efficient Estimation of Causal Effects Under Two-Phase Sampling with Error-Prone Outcome and Treatment Measurements},
year = {2026},
howpublished = {\url{https://pith.science/paper/COVYMGUW}},
note = {Machine review of arXiv:2506.21777}
}
read the original abstract
Measurement error is a common challenge for causal inference studies using electronic health record (EHR) data, where clinical outcomes and treatments are frequently mismeasured. Researchers often address measurement error by conducting manual chart reviews to validate measurements in a subset of the full EHR data -- a form of two-phase sampling. To improve efficiency, phase-two samples are often collected in a biased manner dependent on the patients' initial, error-prone measurements. In this work, motivated by our aim of performing causal inference with error-prone outcome and treatment measurements under two-phase sampling, we develop solutions applicable to both this specific problem and the broader problem of causal inference with two-phase samples. For our specific measurement error problem, we construct two asymptotically equivalent doubly-robust estimators of the average treatment effect and demonstrate how these estimators arise from two previously disconnected approaches to constructing efficient estimators in general two-phase sampling settings. We document various sources of instability affecting estimators from each approach and propose modifications that can considerably improve finite sample performance in any two-phase sampling context. We demonstrate the utility of our proposed methods through simulation studies and an illustrative example assessing effects of antiretroviral therapy on occurrence of AIDS-defining events in patients with HIV from the Vanderbilt Comprehensive Care Clinic.
Figures
Forward citations
Cited by 2 Pith papers
-
Causal Inference with Multiple Misclassified Exposures: A Control Variate-Adjusted Calibration Weighting Approach
New calibration weighting and control variate estimators for causal inference with multiple misclassified binary exposures achieve consistency and double robustness without modeling the misclassification process, with...
-
Targeted maximum likelihood estimation for longitudinal two-stage designs with outcome subsampling
IPCW-LTMLE with targeted sampling weights and a plug-in LTMLE treating the stage-two indicator as an intervention node substantially outperform weighted Kaplan–Meier, and cross-fitted variance restores nominal coverage.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...
-
[2]
A., Lumley, T., and Shepherd, B
Amorim, G., Tao, R., Lotspeich, S., Shaw, P. A., Lumley, T., and Shepherd, B. E. (2021). Two-phase sampling designs for data validation in settings with covariate measurement error and continuous outcome. Journal of the Royal Statistical Society Series A: Statistics in Society , 184(4):1368--1389
work page 2021
-
[3]
Antonelli, J. and Cefalu, M. (2020). Averaging causal estimators in high dimensions. Journal of Causal Inference , 8(1):92--107
work page 2020
-
[4]
and Robins, J
Bang, H. and Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics , 61(4):962--973
2005
-
[5]
Barnatchez, K., Nethery, R., Shepherd, B. E., Parmigiani, G., and Josey, K. P. (2024). Flexible and efficient estimation of causal effects with error-prone exposures: A control variates approach for measurement error. arXiv preprint arXiv:2410.12590
work page Pith review arXiv 2024
-
[6]
Berk, R., Brown, L., Buja, A., Zhang, K., and Zhao, L. (2013). Valid post-selection inference. The Annals of Statistics , pages 802--837
2013
-
[7]
Breslow, N. E. and Chatterjee, N. (1999). Design and analysis of two-phase studies with binary outcome applied to wilms tumour prognosis. Journal of the Royal Statistical Society: Series C (Applied Statistics) , 48(4):457--468
work page 1999
-
[8]
Carroll, R. J., Ruppert, D., Stefanski, L. A., and Crainiceanu, C. M. (2006). Measurement error in nonlinear models: a modern perspective . Chapman and Hall/CRC
work page 2006
Show all 40 references
-
[9]
R., and Zubizarreta, J
Chattopadhyay, A., Cohn, E. R., and Zubizarreta, J. R. (2024). One-step weighting to generalize and transport treatment effect estimates to a target population. The American Statistician , 78(3):280--289
2024
-
[10]
B., Vansteelandt, S., and Moreno-Betancur, M
Ellul, S., Carlin, J. B., Vansteelandt, S., and Moreno-Betancur, M. (2024). Causal machine learning methods and use of sample splitting in settings with high-dimensional confounding. arXiv preprint arXiv:2405.15242
2024 arXiv
-
[11]
J., Shaw, P
Giganti, M. J., Shaw, P. A., Chen, G., Bebawy, S. S., Turner, M. M., Sterling, T. R., and Shepherd, B. E. (2020). Accounting for dependent errors in predictors and time-to-event outcomes using electronic health records, validation samples, and multiple imputation. The annals o...
2020
-
[12]
and Tibshirani, R
Hastie, T. and Tibshirani, R. (1986). Generalized additive models. Statistical science , 1(3):297--310
1986
-
[13]
S., van der Laan , M
Hejazi, N. S., van der Laan , M. J., Janes, H. E., Gilbert, P. B., and Benkeser, D. C. (2021). Efficient nonparametric inference on the effects of stochastic interventions under two-phase sampling, with applications to vaccine efficacy trials. Biometrics , 77(4):1241--1253
2021
-
[14]
and Robins, J
Hernan, M. and Robins, J. (2024). Causal Inference: What If . Chapman & Hall/CRC Monographs on Statistics & Applied Probab. CRC Press
2024
-
[15]
Hou, J., Mukherjee, R., and Cai, T. (2025). Efficient and robust semi-supervised estimation of average treatment effect with partially annotated treatment and response. Journal of Machine Learning Research , 26(40):1--77
2025
-
[16]
Initiation of antiretroviral therapy in early asymptomatic HIV infection
Insight Start Study Group (2015). Initiation of antiretroviral therapy in early asymptomatic HIV infection. New England Journal of Medicine , 373(9):795--807
2015
-
[17]
and Mao, X
Kallus, N. and Mao, X. (2024). On the role of surrogates in the efficient estimation of treatment effects with limited outcome data. Journal of the Royal Statistical Society Series B: Statistical Methodology , page qkae099
2024
-
[18]
H., and Dahabreh, I
Karlsson, R., Wang, G., Krijthe, J. H., and Dahabreh, I. J. (2024). Robust integration of external control data in randomized trials. arXiv preprint arXiv:2406.17971
2024 arXiv
-
[19]
Kennedy, E. H. (2016). Semiparametric theory and empirical processes in causal inference. Statistical causal inferences and their applications in public health research , pages 141--167
2016
-
[20]
Kennedy, E. H. (2020). Efficient nonparametric causal inference with missing exposure information. The international journal of biostatistics , 16(1)
2020
-
[21]
Kennedy, E. H. (2024). Semiparametric doubly robust targeted double machine learning: a review. Handbook of Statistical Methods for Precision Medicine , pages 207--236
2024
-
[22]
W., Mukherjee, R., Wang, R., Fischer, H., and Haneuse, S
Levis, A. W., Mukherjee, R., Wang, R., Fischer, H., and Haneuse, S. (2024a). Double sampling for informatively missing data in electronic health record-based comparative effectiveness research. Statistics in Medicine
2024
-
[23]
W., Mukherjee, R., Wang, R., and Haneuse, S
Levis, A. W., Mukherjee, R., Wang, R., and Haneuse, S. (2024b). Robust causal inference for point exposures with missing confounders. Canadian Journal of Statistics , page e11832
2024
-
[24]
C., Shepherd, B
Lotspeich, S. C., Shepherd, B. E., Amorim, G. G., Shaw, P. A., and Tao, R. (2022). Efficient odds ratio estimation under two-phase sampling using error-prone data from a multi-national hiv research cohort. Biometrics , 78(4):1674--1685
2022
-
[25]
and Nelder, J
McCullagh, P. and Nelder, J. A. (1989). Generalized linear models . Routledge
1989
-
[26]
Neyman, J. (1938). Contribution to the theory of sampling human populations. Journal of the American Statistical Association , 33(201):101--116
1938
-
[27]
J., Shepherd, B
Oh, E. J., Shepherd, B. E., Lumley, T., and Shaw, P. A. (2021). Raking and regression calibration: Methods to address bias from correlated covariate and time-to-event error. Statistics in Medicine , 40(3):631--649
2021
-
[28]
and Wefelmeyer, W
Pfanzagl, J. and Wefelmeyer, W. (1985). Contributions to a general asymptotic statistical theory. Statistics & Risk Modeling , 3(3-4):379--388
1985
-
[29]
Polley, E., LeDell, E., Kennedy, C., and van der Laan , M. (2024). SuperLearner: Super Learner Prediction . R package version 2.0-29
2024
-
[30]
R : A Language and Environment for Statistical Computing
R Core Team (2025). R : A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria
2025
-
[31]
M., Rotnitzky, A., and Zhao, L
Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association , 89(427):846--866
1994
-
[32]
and van der Laan , M
Rose, S. and van der Laan , M. J. (2011). A targeted maximum likelihood estimator for two-stage designs. The International Journal of Biostatistics , 7(1):0000102202155746791217
2011
-
[33]
Rubin, D. B. and van der Laan , M. J. (2008). Empirical efficiency maximization: improved locally efficient covariate adjustment in randomized experiments and survival analysis. The International Journal of Biostatistics , 4(1)
2008
-
[34]
E., Han, K., Chen, T., Bian, A., Pugh, S., Duda, S
Shepherd, B. E., Han, K., Chen, T., Bian, A., Pugh, S., Duda, S. N., Lumley, T., Heerman, W. J., and Shaw, P. A. (2023). Multiwave validation sampling for error-prone electronic health records. Biometrics , 79(3):2649--2663
2023
-
[35]
Stitelman, O. M. and van der Laan , M. J. (2010). Collaborative targeted maximum likelihood for time to event data. The International Journal of Biostatistics , 6(1)
2010
-
[36]
Tsiatis, A. A. (2006). Semiparametric theory and missing data
2006
-
[37]
Valeri, L. (2021). Measurement error in causal inference. In Handbook of Measurement Error Models , pages 453--480. Chapman and Hall/CRC
2021
-
[38]
J., Polley, E
van der Laan , M. J., Polley, E. C., and Hubbard, A. E. (2007). Super learner. Statistical applications in genetics and molecular biology , 6(1)
2007
-
[39]
van der Laan , M. J. and Robins, J. M. (2003). Unified methods for censored longitudinal data and causality . Springer
2003
-
[40]
Wang, R., Wang, Q., and Miao, W. (2023). A maximin optimal approach for model-free sampling designs in two-phase studies. arXiv preprint arXiv:2312.10596
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.