REVIEW 2 major objections 5 minor 53 references
Targeted maximum likelihood estimation for longitudinal two-stage designs with outcome subsampling
T0 review · 2 major / 5 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read Two new estimators recover large efficiency gains for survival parameters when outcomes are only measured on a second-stage subsample.
desk verdict Solid longitudinal TMLE toolkit for outcome-subsampling two-stage designs; efficiency and coverage claims hold under the reported DGP, with the single-DGP caveat already flagged by the authors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The LTMLE that inserts the second-stage sampling indicator Δ as an intervention node inside the usual sequential-regression recursion, so the target is the counterfactual mean under the joint intervention that both prevents censoring and sets Δ = 1 for everyone; this returns to plug-in estimation and never multiplies by inverse sampling weights.
What would settle it
Re-run the same Monte Carlo design at N = 500, 1 000 and 3 000 and check whether the LTMLE still shows 30–70 percent lower empirical variance than weighted Kaplan–Meier with known weights and whether the cross-fitted influence-curve intervals actually cover at the nominal 95 percent rate; a clear failure on either metric would refute the central claims.
Extended reading notes
Core claim
In longitudinal two-stage designs with outcome subsampling, an LTMLE that treats the second-stage sampling indicator as an additional intervention node (and therefore never inverse-weights) achieves up to 73 percent lower empirical variance than weighted Kaplan–Meier with known sampling probabilities, with 30–50 percent reductions common; an IPCW-LTMLE that estimates and targets those known weights still yields consistent 20–35 percent gains. Cross-fitted influence-curve variance is required for valid 95 percent intervals when flexible nuisance estimators are used.
Load-bearing premise
The identification and efficiency arguments require that both the censoring process and the second-stage sampling decision are independent of the counterfactual outcome given the observed past (sequential randomization), plus positivity; the authors also work under a stronger model than pure coarsening-at-random for the bivariate censoring, so closed-form full efficiency is not claimed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes longitudinal resampling designs (HIV mortality with LTFU tracing) as two-stage designs with outcome subsampling, O = (V, Δ, ΔX). It develops two estimators for marginal means under joint intervention on right-censoring and stage-two sampling: (i) IPCW-LTMLE, a longitudinal extension of Rose & van der Laan IPCW-TMLE that applies full-data LTMLE among Δ = 1 subjects and shows that estimating and targeting known sampling weights Π_Δ yields further variance reductions; (ii) a plug-in LTMLE that treats Δ itself as an intervention node in sequential regression, avoiding inverse weighting. Identification is given under sequential randomization of censoring and of Δ (plus positivity). Simulations (N ∈ {500, 1000, 3000}, 1000 MC replications, ten time points) under a single HIV-like DGP report negligible bias, ordered efficiency gains (LTMLE up to 73 % lower variance than known-weight wKM; IPCW-LTMLE 20–35 %), and that standard influence-curve variance undercovers (coverage as low as ~76 %) while the proposed hybrid cross-fitted variance restores near-nominal coverage. The authors note that bivariate censoring by (τ, Δ) precludes closed-form full efficiency under CAR and that they work under the larger SRA model.
Significance. If the reported efficiency hierarchy and the necessity of cross-fitted variance hold more generally, the paper supplies practically useful closed-form alternatives to the dominant weighted Kaplan–Meier analyses of resampling designs and, more broadly, to inverse-weighted estimators for longitudinal two-stage designs with outcome subsampling. Strengths include a clear identification argument, explicit algorithms for targeting Π_Δ and for hybrid cross-fitted variance, careful handling of deterministic outcomes, and transparent Monte Carlo evidence that standard IC variance fails under flexible Super Learner nuisance estimation. The work also correctly situates itself relative to bivariate-censoring theory (no claim of nonparametric efficiency under CAR). These contributions expand the methodological toolkit for a design class that arises in HIV cascade studies, EHR long-term outcomes, and related settings.
major comments (2)
- All headline numerical claims (LTMLE 30–73 % variance reduction vs known-weight wKM; IPCW-LTMLE 20–35 %; cross-fit restoring coverage from ~76 % to nominal) rest on a single carefully constructed DGP (Section 3.1: fixed visit/death/reporting rates, τ ∈ {5,7,9,10}, resampling probability 0.2 among LTFU, Super Learner library matched to the generative process, substantial deterministic information from V(t) and Id(t)). The Discussion itself flags this limitation. Without at least one qualitatively different regime (e.g., weaker positivity, lower resampling fraction, reduced determinism, or deliberate outcome-model misspecification), it remains unclear whether the efficiency hierarchy and the necessity of cross-fitting are general properties or artifacts of this sparsity/determinism pattern. A second simulation regime is needed to support the central empirical claims.
- Sections 2.3.2 and Discussion correctly note that the authors work under sequential randomization (SRA) treating censoring like a treatment node, whereas the honest missing-data model is coarsening-at-random (CAR) for bivariate censoring by (τ, Δ); estimators efficient for the larger SRA model need not be efficient for the smaller CAR model, and closed-form full efficiency is not claimed. This is an important caveat for applied readers who may interpret “highly efficient closed-form alternatives” as near-optimal. The manuscript should state more prominently (abstract or early methods) that the efficiency gains are relative to wKM and to untargeted IPCW, not relative to the nonparametric CAR bound, and should clarify what practical efficiency loss relative to a fully efficient (non-closed-form) estimator might be expected.
minor comments (5)
- Tables 1–4 report variance to six decimals and coverage to three; a compact relative-efficiency column (vs known-weight wKM) would make the hierarchy easier to read.
- Notation for the full-data structure X = (W, τ, L-bar(τ), Y-bar(τ)) deliberately includes right-censoring; a short remark early in Section 2.2 that this is a convenience definition (not the usual complete-data X) would reduce confusion for readers coming from the two-stage literature.
- The hybrid cross-fitting design (full-data point estimate, cross-fitted variance only) is well motivated by deterministic sparsity, but the fallback rule to non-cross-fitted variance when Super Learner fails inside a fold should be stated more precisely (how often it triggers at N = 500, t = 1).
- Several self-citations to the authors’ related hazard-TMLE preprint and TMLE monographs are appropriate background; a brief sentence distinguishing the present plug-in LTMLE from that hazard-based estimator would help readers who encounter both papers.
- Minor typographical issues: “substatially” (p. 17), “crosffit” (p. 17), and occasional missing spaces around mathematical operators.
Circularity Check
No significant circularity: standard TMLE targeting and Monte-Carlo evaluation against a known DGP, with self-citations only as background method references.
full rationale
The paper constructs two estimators (IPCW-LTMLE with optional targeting of known sampling weights, and LTMLE treating the stage-two indicator as an intervention node) from the longitudinal g-computation formula under sequential randomization, then evaluates finite-sample bias/variance/coverage by simulation under an explicitly stated data-generating process whose true survival curves are known. Targeting steps enforce the mean-zero efficient-influence-curve equation by design (the defining property of TMLE), which is not a circular prediction of an independent quantity. Variance reductions versus weighted Kaplan–Meier and the necessity of cross-fitted influence-curve variance are empirical Monte-Carlo findings, not algebraic identities forced by the inputs. Self-citations (to the authors’ related hazard-TMLE preprint, TMLE monographs, and Rose & van der Laan IPCW-TMLE) supply background algorithms and software; none is invoked as a uniqueness theorem that forces the reported efficiency hierarchy. Identification assumptions (SRA plus positivity) and the acknowledged SRA-versus-CAR efficiency gap are stated openly and do not reduce the simulation claims to tautologies. The derivation chain is therefore self-contained against external benchmarks.
Assumptions & free parameters
free parameters (5)
- Stage-two resampling probability among LTFU =
0.2
- Natural death-reporting probability Id(t) =
0.2
- Administrative censoring time distribution τ∈{5,7,9,10} =
discrete support {5,7,9,10}
- Super Learner library and screening choices
- Cross-fitting fold count V and convergence tolerance for Π targeting
assumptions (5)
- domain assumption Sequential randomization of censoring: Y(t0)_{C̄(t0)=1} ⊥ C(t) | history for t≤t0.
- domain assumption Randomization of stage-two sampling: Y(t0)_{Δ=1,C̄(t0)=1} ⊥ Δ | stage-one history (by design in resampling).
- domain assumption Positivity of censoring and of stage-two sampling conditional on history.
- ad hoc to paper Working under sequential randomization (SRA) for censoring treated like a treatment node, rather than coarsening-at-random only.
- standard math Standard TMLE regularity / Donsker or cross-fit conditions for asymptotic linearity of sequential regression estimators.
Cite this review
Pith. "Pith review of Targeted maximum likelihood estimation for longitudinal two-stage designs with outcome subsampling." pith.science (2026). https://pith.science/paper/CSBNWUVX
@misc{pith2026260702702,
author = {Pith},
title = {Pith review of: Targeted maximum likelihood estimation for longitudinal two-stage designs with outcome subsampling},
year = {2026},
howpublished = {\url{https://pith.science/paper/CSBNWUVX}},
note = {Machine review of arXiv:2607.02702}
}
read the original abstract
We consider efficient estimation of causal parameters in longitudinal two-stage designs with outcome subsampling, motivated by resampling designs in HIV-related mortality studies. In these studies, many participants become lost to follow-up; resampling designs address this by tracing a subset of lost individuals to ascertain their outcomes. Analyses often use inverse-probability-weighted Kaplan-Meier (wKM) estimators that discard longitudinal covariate information and suffer from efficiency losses. We note that resampling designs are an instance of a broader class: two-stage designs with outcome subsampling, in which a first stage collects some data on all participants and a second stage collects outcome information on a selected subset. This connection motivates two estimators. First, drawing on inverse probability of censoring weighted targeted maximum likelihood estimation (IPCW-TMLE) for two-stage designs, we develop its longitudinal extension, IPCW longitudinal TMLE (IPCW-LTMLE) and show that estimating and targeting the known second-stage sampling weights yields variance reductions of up to 36% over the use of known sampling probabilities. Second, given that inverse weighting sacrifices efficiency, we propose an LTMLE that incorporates the second-stage sampling indicator as an intervention node in the sequential regression framework, returning to plug-in estimation and avoiding inverse weighting entirely. Simulations across sample sizes show that LTMLE achieves up to 73% lower variance than wKM with known sampling weights, with reductions of 30-50% common across settings, while IPCW-LTMLE achieves consistent gains of 20-35%. We further demonstrate that cross-fitted variance estimation is essential for valid inference: standard variance estimators yield confidence interval coverage as low as 76%, while our cross-fitted variants consistently restore coverage to nominal levels.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
E., Musick, B
An, M.-W., Frangakis, C. E., Musick, B. S., and Yiannoutsos, C. T. (2009). The need for double-sampling designs in survival studies: an application to monitor pepfar.Biometrics, 65(1):301–306
2009
-
[2]
E., Yiannoutsos, C
An, M.-W., Frangakis, C. E., Yiannoutsos, C. T., et al. (2015). Choosing profile double-sampling designs for survival estimation with application to pepfar evaluation
2015
-
[3]
K., and Yiannoutsos, C
Bakoyannis, G., Diero, L., Mwangi, A., Wools-Kaloustian, K. K., and Yiannoutsos, C. T. (2020). A semiparametric method for the analysis of outcomes during a gap in hiv care under incomplete outcome ascertainment.Statistical communications in infectious diseases, 12(s1):20190013
2020
-
[4]
and Robins, J
Bang, H. and Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models.Biometrics, 61(4):962–973
2005
-
[5]
Barnatchez, K., Josey, K. P., Hejazi, N. S., Shepherd, B. E., Parmigiani, G., and Nethery, R. (2025). Efficient estimation of causal effects under two-phase sampling with error-prone outcome and treatment measurements.arXiv preprint arXiv:2506.21777
arXiv 2025
-
[6]
Breslow, N., McNeney, B., and Wellner, J. A. (2003). Large sample theory for semiparametric regression models with two-phase, outcome dependent sampling.The annals of Statistics, 31(4):1110–1139
2003
-
[7]
Breslow, N. E. and Holubkov, R. (1997). Maximum likelihood estimation of logistic regression param- eters under two-phase, outcome-dependent sampling.Journal of the Royal Statistical Society: Series B (Statistical Methodology), 59(2):447–461
1997
-
[8]
E., Lubin, J
Breslow, N. E., Lubin, J. H., Marek, P., and Langholz, B. (1983). Multiplicative models and cohort analysis.Journal of the American Statistical Association, 78(381):1–12
1983
Show all 53 references
-
[9]
and van der Laan, M
Cai, W. and van der Laan, M. J. (2020). One-step targeted maximum likelihood estimation for time-to- event outcomes.Biometrics, 76(3):722–733
2020
-
[10]
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters
2018
-
[11]
Cochran, W. G. (1977). Sampling techniques.Johan Wiley & Sons Inc
1977
-
[12]
S., Green, D
Coppock, A., Gerber, A. S., Green, D. P., and Kern, H. L. (2017). Combining double sampling and bounds to address nonignorable missing outcomes in randomized experiments.Political Analysis, 25(2):188–206
2017
-
[13]
M., Haber, M., Ferdinands, J
Foppa, I. M., Haber, M., Ferdinands, J. M., and Shay, D. K. (2013). The case test-negative design for studies of the effectiveness of influenza vaccine.Vaccine, 31(30):3104–3109
2013
-
[14]
E., Qian, T., Wu, Z., and Diaz, I
Frangakis, C. E., Qian, T., Wu, Z., and Diaz, I. (2015). Deductive derivation and turing-computerization of semiparametric efficient estimation.Biometrics, 71(4):867–874
2015
-
[15]
Frangakis, C. E. and Rubin, D. B. (2001). Addressing an idiosyncrasy in estimating survival curves using double sampling in the presence of self-selected right censoring.Biometrics, 57(2):333–342
2001
-
[16]
H., Bwana, M
Geng, E. H., Bwana, M. B., Muyindike, W., Glidden, D. V., Bangsberg, D. R., Neilands, T. B., Bern- heimer, I., Musinguzi, N., Yiannoutsos, C. T., and Martin, J. N. (2013). Failure to initiate antiretroviral therapy, loss to follow-up and mortality among hiv-infected patients d...
2013
-
[17]
H., Glidden, D
Geng, E. H., Glidden, D. V., Bangsberg, D. R., Bwana, M. B., Musinguzi, N., Nash, D., Metcalfe, J. Z., Yiannoutsos, C. T., Martin, J. N., and Petersen, M. L. (2012). A causal framework for understanding the effect of losses to follow-up on epidemiologic analyses in clinic-base...
2012
-
[18]
H., Odeny, T
Geng, E. H., Odeny, T. A., Lyamuya, R. E., Nakiwogga-Muwanga, A., Diero, L., Bwana, M., Muyindike, W., Braitstein, P., Somi, G. R., Kambugu, A., et al. (2015). Estimation of mortality among hiv-infected people on antiretroviral treatment in east africa: a sampling based approa...
2015
-
[19]
and van der Laan, M
Gruber, S. and van der Laan, M. (2025).twoStageDesignTMLE: Targeted Maximum Likelihood Esti- mation for Two-Stage Study Design. R package version 1.0.1.2
2025
-
[20]
B., Sikazwe, I., Sikombe, K., Eshun-Wilson, I., Czaicki, N., Beres, L
Holmes, C. B., Sikazwe, I., Sikombe, K., Eshun-Wilson, I., Czaicki, N., Beres, L. K., Mukamba, N., Simbeza, S., Bolton Moore, C., Hantuba, C., et al. (2018). Estimated mortality on hiv treatment among active patients and patients lost to follow-up in 4 provinces of zambia: fin...
2018
-
[21]
´A., et al
J Smith, M., V Phillips, R., Maringe, C., Luque Fern´ andez, M. ´A., et al. (2025). Performance of cross-validated targeted maximum likelihood estimation
2025
-
[22]
Laan, M. J. and Robins, J. M. (2003).Unified methods for censored longitudinal data and causality. Springer
2003
-
[23]
E., Phillips, R
Landsiedel, K. E., Phillips, R. V., Petersen, M. L., and van der Laan, M. J. (2025). Hazard-based tar- geted maximum likelihood estimation for survival in resampling designs.arXiv preprint arXiv:2511.15045. arXiv:2511.15045 [stat.ME]
2025
-
[24]
D., Schwab, J., Petersen, M
Lendle, S. D., Schwab, J., Petersen, M. L., and van der Laan, M. J. (2017). ltmle: an r package imple- menting targeted minimum loss-based estimation for longitudinal data.Journal of Statistical Software, 81:1–21
2017
-
[25]
W., Mukherjee, R., Wang, R., and Haneuse, S
Levis, A. W., Mukherjee, R., Wang, R., and Haneuse, S. (2022). Double sampling and semiparametric methods for informatively missing data.arXiv preprint arXiv:2204.02432
2022 arXiv
-
[26]
and Tseng, C.-h
Li, G. and Tseng, C.-h. (2008). Non-parametric estimation of a survival function with two-stage design studies.Scandinavian journal of statistics, 35(2):193–211
2008
-
[27]
C., Shepherd, B
Lotspeich, S. C., Shepherd, B. E., Amorim, G. G., Shaw, P. A., and Tao, R. (2022). Efficient odds ratio estimation under two-phase sampling using error-prone data from a multi-national hiv research cohort. Biometrics, 78(4):1674–1685
2022
-
[28]
Neyman, J. (1938). Contribution to the theory of sampling human populations.Journal of the American Statistical Association, 33(201):101–116
1938
-
[29]
Petersen, M., Schwab, J., Gruber, S., Blaser, N., Schomaker, M., and van der Laan, M. (2014). Targeted maximum likelihood estimation for dynamic and static longitudinal marginal structural working models. Journal of causal inference, 2(2):147–185
2014
-
[30]
(2025).SuperLearner: Super Learner Prediction
Polley, E., LeDell, E., Kennedy, C., and van der Laan, M. (2025).SuperLearner: Super Learner Prediction. R package version 2.0-40
2025
-
[31]
Polley, E. C. and van der Laan, M. J. (2023).SuperLearner: Super Learner Prediction. R package version 2.1-6
2023
-
[32]
Qian, T., Frangakis, C., and Yiannoutsos, C. (2020). Deductive semiparametric estimation in double- sampling designs with application to pepfar.Statistics in Biosciences, 12(3):417–445
2020
-
[33]
A., Williamson, B
Qiu, S., Gruber, S., Shaw, P. A., Williamson, B. D., and van der Laan, M. J. (2026). Efficient targeted maximum likelihood estimators for two-phase design problems. arXiv:2602.24131v1 [stat.ME]
2026
-
[34]
M., van der Laan, M
Quale, C. M., van der Laan, M. J., and Robins, J. M. (2002). Locally efficient estimation with bivariate right censored data. 30
2002
-
[35]
Robins, J., Rotnitzky, A., and Bonetti, M. (2001). Discussion of the frangakis and rubin article. Biometrics, 57(2):343–347
2001
-
[36]
M., Rotnitzky, A., and Zhao, L
Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed.Journal of the American statistical Association, 89(427):846–866
1994
-
[37]
and van der Laan, M
Rose, S. and van der Laan, M. J. (2011). A targeted maximum likelihood estimator for two-stage designs.The international journal of biostatistics, 7(1):0000102202155746791217
2011
-
[38]
(2023).ltmle: Longitudinal Targeted Maximum Likelihood Estimation
Schwab, J., Lendle, S., Petersen, M., van der Laan, M., and Gruber, S. (2023).ltmle: Longitudinal Targeted Maximum Likelihood Estimation. R package version 1.3-0
2023
-
[39]
Shirakawa, T., Li, Y., Wu, Y., Qiu, S., Li, Y., Zhao, M., Iso, H., and Van Der Laan, M. (2024). Longitudinal targeted minimum loss-based estimation with temporal-difference heterogeneous transformer. Proceedings of machine learning research, 235:45097
2024
-
[40]
W., Lee, C., Arterburn, D
Sun, S., Haneuse, S., Levis, A. W., Lee, C., Arterburn, D. E., Fischer, H., Shortreed, S., and Mukherjee, R. (2025). Estimating weighted quantile treatment effects with missing outcome data by double sampling. Biometrics, 81(2):ujaf038
2025
-
[41]
Tao, R., Zeng, D., and Lin, D.-Y. (2020). Optimal designs of two-phase studies.Journal of the American Statistical Association, 115(532):1946–1959
2020
-
[42]
Therneau, T. M. (2024).A Package for Survival Analysis in R. R package version 3.8-3
2024
-
[43]
Van Der Laan, M. J. (1996). Efficient estimation in the bivariate censoring model and repairing npmle. The Annals of Statistics, 24(2):596–627
1996
-
[44]
van der Laan, M. J. and Gruber, S. (2011). Targeted minimum loss based estimation of causal effects of multiple time point interventions.International Journal of Biostatistics, 8(1)
2011
-
[45]
J., Polley, E
Van der Laan, M. J., Polley, E. C., and Hubbard, A. E. (2007). Super learner.Statistical applications in genetics and molecular biology, 6(1)
2007
-
[46]
van der Laan, M. J. and Robins, J. M. (2003).Unified Methods for Censored Longitudinal Data and Causality. Springer, New York
2003
-
[47]
van der Laan, M. J. and Rose, S. (2018).Targeted Learning in Data Science: Causal Inference for Complex Longitudinal Studies. Springer Series in Statistics. Springer
2018
-
[48]
J., Rose, S., et al
Van der Laan, M. J., Rose, S., et al. (2011).Targeted learning: causal inference for observational and experimental data, volume 4. Springer
2011
-
[49]
Wang, W., Scharfstein, D., Tan, Z., and MacKenzie, E. J. (2009). Causal inference in outcome-dependent two-phase sampling designs.Journal of the Royal Statistical Society Series B: Statistical Methodology, 71(5):947–969
2009
-
[50]
D., Krakauer, C., Johnson, E., Gruber, S., Shepherd, B
Williamson, B. D., Krakauer, C., Johnson, E., Gruber, S., Shepherd, B. E., van der Laan, M. J., Lumley, T., Lee, H., Hern´ andez-Mu˜ noz, J. J., Zhao, F., et al. (2026). Assessing treatment effects in observational data with missing confounders: A comparative study of practica...
2026
-
[51]
T., An, M.-W., Frangakis, C
Yiannoutsos, C. T., An, M.-W., Frangakis, C. E., Musick, B. S., Braitstein, P., Wools-Kaloustian, K., Ochieng, D., Martin, J. N., Bacon, M. C., Ochieng, V., et al. (2008). Sampling-based approaches to improve estimation of mortality among patient dropouts: experience from a la...
2008
-
[52]
and Van Der Laan, M
Zheng, W. and Van Der Laan, M. J. (2010). Asymptotic theory for cross-validated targeted maximum likelihood estimation
2010
-
[53]
Zivich, P. N. and Breskin, A. (2021). Machine learning for causal inference: on the use of cross-fit estimators.Epidemiology, 32(3):393–401. 31 A Supplementary material A.1 Table of results for cross-fitted variance estimation A.1.1 IPCW-LTMLE,N= 500 Figure 8: Coverage, mean s...
2021
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.