Pith. sign in

REVIEW 2 major objections 4 minor 54 references

This paper proves that AIPW confidence intervals in randomized controlled trials reach nominal coverage at essentially the classical root-n rate when the outcome model is estimated at root-n speed with sub-Weibull tails, and that cross-fitt

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 14:50 UTC pith:45SZ4UNV

load-bearing objection Real first result on AIPW Wald-CI coverage with estimated variance, but the near-root-n rate is conditional on sub-Weibull assumptions that are only partially verified. the 2 major comments →

arxiv 2512.18898 v3 pith:45SZ4UNV submitted 2025-12-21 math.ST stat.MEstat.TH

Model-Agnostic Bounds for Augmented Inverse Probability Weighted Estimators' Wald-Confidence Interval Coverage in Randomized Controlled Trials

classification math.ST stat.MEstat.TH MSC 62G2062G1562F1262G05
keywords AIPW estimatorWald confidence intervalBerry-Esseen boundscross-fittingrandomized controlled trialvariance estimation biasnonparametric inferencecausal inference
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Flexible estimators of average treatment effects in randomized trials all share the same asymptotic normal distribution, so asymptotic normality gives practitioners little reason to prefer one. This paper sharpens the comparison by studying how fast Wald confidence intervals built from these estimators reach their nominal coverage. The author proves non-asymptotic Berry-Esseen-type bounds showing that, with a known propensity score, the cross-fit AIPW interval reaches nominal coverage at essentially the root-n rate up to log factors, matching the classical sample-mean rate, provided the black-box outcome model estimator converges at root-n speed with approximately sub-Weibull tails. The paper also shows the cross-fit plug-in variance estimator overestimates the oracle variance when the outcome model converges slowly, which pushes coverage upward, while the non-cross-fit variance estimator can underestimate, pushing coverage downward. These results give a theoretical explanation for the empirical preference for cross-fitting and identify a possible efficiency-coverage trade-off.

Core claim

The central claim is that the error in coverage probability of a Wald interval based on an AIPW estimator with an estimated variance can be bounded non-asymptotically by terms governed by the L2 error of the outcome-model estimator, the complexity of the function class containing it, and the bias of the plug-in variance estimator. Under the paper's sub-Weibull tail conditions, the bound for the cross-fit estimator is O(K^{1/2} r_hat_1(n)(log n)^{...} + ... + K sqrt(log n/n)), which is root-n up to log factors when r_hat_1(n)=n^{-1/2}. A separate theorem shows |sigma_dagger,a - sigma_#,a| is quadratic in the L2 error when the deterministic approximation Q_#,a is chosen as the mean of the esti

What carries the argument

The argument runs through the transformed outcome T_a(Q)(x,a,y)=1(a=a')/pi*(a'|x)(y-Q(x))+Q(x), whose sample mean is the AIPW estimator and whose variance defines the oracle variance sigma^2_#,a. A sample-size-dependent deterministic function Q_#,a, typically the mean of the black-box estimator, absorbs finite-sample bias. Coverage is then bounded by combining a delta-method lemma for coverage probabilities with the classical Berry-Esseen theorem on the linearized sample mean, empirical-process bounds that are sharp for VC and VC-hull classes, and tail bounds on the nuisance estimator's L2 error and on the plug-in variance estimator's deviation from its expectation.

Load-bearing premise

The headline near-root-n rate rests on Conditions 7 and 8, which require the L2 error of the black-box outcome-model estimator (and a related moment) to have approximate sub-Weibull tails; the paper verifies these only for series regression, not for the flexible pipelines used in its simulation.

What would settle it

Run the simulation in the paper with an outcome model whose L2 error is known to decay at n^{-1/4} and has heavy-tailed fluctuations; if the cross-fit Wald interval's coverage does not exceed nominal while the plug-in variance exceeds the oracle variance, the variance-bias mechanism fails. Alternatively, compute empirical tail probabilities of ||Q_hat_{k,a}-Q_#,a||_2 / r_hat_1(n) for a chosen learner; if they decay more slowly than sub-Weibull, the claimed root-n log coverage rate cannot hold.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the bounds are right, cross-fitting improves Wald-CI coverage even when Donsker conditions hold, because its variance estimator tends to overestimate the oracle variance.
  • For outcome model estimators converging at root-n speed with sub-Weibull tails, the coverage error of cross-fit AIPW intervals is comparable to the best-known rate for the sample mean, up to log factors.
  • Non-cross-fit AIPW intervals built on highly flexible outcome models can undercover because the plug-in variance underestimates, even though the estimator is asymptotically normal.
  • The variance-bias term is nearly second order in the L2 nuisance error for known propensity scores, so it does not dominate the coverage rate.
  • There is a potential trade-off: using a more flexible outcome model improves efficiency but slows the convergence of variance bias and hence of coverage.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: A practitioner could use these bounds as a diagnostic: if a black-box outcome estimator's L2 error and tail behavior can be estimated on splits, the formulas give a concrete prediction of whether cross-fit or single-fit intervals will over- or undercover.
  • Inference: In observational studies, where the propensity score must be estimated, the near-quadratic variance-bias cancellation may break down, so the qualitative conclusions about cross-fitting could differ.
  • Inference: The sign predictions for variance bias are testable in a small simulation: compute the oracle variance for Q_#,a as the Monte Carlo mean of the outcome estimator and compare with the average plug-in variance; if the sign pattern disagrees, the mechanism proposed here is incomplete.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper develops Berry-Esseen-type bounds for the coverage of Wald confidence intervals based on augmented inverse probability weighted (AIPW) estimators of counterfactual means in randomized controlled trials, with the asymptotic variance estimated by plug-in. The cross-fit version (Theorem 1 / Theorem S1) gives a bound of order (K^{2/3} E||Qhat_k,a−Q#,a||_2^2/n)^{1/3} + sqrt(K log n/n), which improves to K^{1/2} rhat_1(n) polylog + rhat_2(n) polylog + K sqrt(log n/n) under approximate sub-Weibull tail Conditions 7–8. The non-cross-fit version (Theorem 3 / Theorem S3) has an additional dependence on the uniform entropy integral of the nuisance function class. Theorems 2 and 4 bound the bias of the plug-in variance estimator relative to an oracle variance based on a deterministic approximation Q#,a, showing the bias is second-order when Q#,a is the mean of the nuisance estimator (Condition 9). Proposition 1 and Section 4 argue that cross-fitting overestimates and non-cross-fitting may underestimate the oracle variance; a simulation with Super Learner and highly adaptive lasso nuisance estimators illustrates the theory, including severe undercoverage of a non-cross-fit flexible AIPW estimator.

Significance. If correct, this is a useful addition to the finite-sample inference literature for semiparametric estimators: to this reviewer's knowledge it is the first treatment of Wald-CI coverage convergence with an estimated variance for flexible AIPW estimators in RCTs. The paper's division of labor between a proved non-asymptotic supplement (Theorems S1–S4) and readable asymptotic versions in the main text is a strength, as is the variance-bias decomposition in Theorem 2: the quadratic-rate bias under Condition 9 is a non-obvious consequence of the known propensity score. The simulation is transparent, and the Discussion candidly lists open issues (conservative empirical-process bounds, the heuristic status of the non-cross-fit under-estimation argument, and the open question of choosing K). The main limitation — the unverified scope of Conditions 7–8 for the black-box estimators used in the simulation, and hence the conditional nature of the headline root-n rate — is only partly acknowledged; the revisions below should make the scope of the headline claims precise.

major comments (2)
  1. [Section 3.1, Condition 8 (and Supplement S2)] The passage after Condition 8 claims that, by Lemma S6, the numerator in Condition 8 'has a tail similar to the L2(P*)-distance' and that Condition 8 'may hold with rhat_2 = rhat_1' under Condition 7. Lemma S6 bounds only the second moment: E[{P*(T_a(Qhat)^2 − E_Qhat[T_a(Qhat)^2])}^2] ≤ 16M^2((1−τπ)/τπ)^2 E||Qhat−Q#,a||^2. A second-moment bound yields at most a polynomial (Chebyshev) tail, not the sub-Weibull decay required by Condition 8; moreover the expression in Lemma S6 involves pointwise evaluations of Qhat, not merely its L2 distance. The same unsupported inference is used in Section S2 to assert Conditions 8/12 for series regression from Proposition S1. Because the 'root-n up to a log' version of Theorem 1 requires Conditions 7 and 8 jointly, Condition 8 is a genuinely additional high-level assumption. Please either prove a real tail bound under Condition 7 (which would require m
  2. [Section 3.3, second paragraph (and Discussion)] The text claims the bias term φ(zα)zα(σ†,a−σ#,a)/σ#,a 'has a faster convergence rate than the right-hand side ... regardless of ... whether Q#,a is chosen to yield a faster rate.' This is not supported by Theorems 1–2. Put r = E||Qhat_k,a−Q#,a||^2 ≍ n^{−2s}. Theorem 1's RHS is O((r/n)^{1/3}+sqrt(log n/n)); Theorem 2 without Condition 9 gives |σ†,a−σ#,a| = O(sqrt(r)+n^{−1}). For s < 1/2, n^{−s} dominates both n^{−(2s+1)/3} and n^{−1/2}, so the bias term is the leading term in the coverage expansion — precisely in the slow-convergence regime emphasized by Proposition 1 for cross-fit variance overestimation. Even under Condition 9, the bias n^{−2s} dominates (r/n)^{1/3} when s < 1/4. The stated theorems remain valid, but the quoted claim and the Discussion's assertion that 'this bias is not a first-order term' require qualification by Condition 9 and an explicit rate condition (e.g., r = o(
minor comments (4)
  1. [Condition 11] The displayed tail bound reads exp(˜c1 t^{˜q1}) in the text; as written, the tail probability grows with t. This is evidently a typo for exp(−˜c1 t^{˜q1}); Conditions 8 and 12 show the minus sign.
  2. [Section 3.1, before Theorem 1] 'The proof of of this theorem' — duplicate 'of'.
  3. [Proposition 1, Eqs. (3)–(4)] The Θ+(...) notation is used for two separate nonnegative-order terms, one of which is subtracted. The sign and order of the difference are determined only when one term dominates. Please clarify, since Eqs. (3)–(4) as written are easy to misread as an order statement for the full difference.
  4. [Figure S1 caption] 'Subfigures (A) and (B) are for the cases with and without approximate sub-Weibull conditions 7 and 8, respectively' appears reversed: panel (A) corresponds to the case without the sub-Weibull conditions and panel (B) to the case with them, matching the rates derived in Section S3.

Circularity Check

0 steps flagged

No circularity: the coverage bounds follow from Berry-Esseen and empirical-process arguments, and the oracle Q# is a post-hoc decomposition target rather than a fitted prediction.

full rationale

The central derivation is self-contained. Theorem 1/3 are proved from the Berry-Esseen theorem, Hall's delta method for CI coverage, Taylor expansion, and Chernozhukov--Chetverikov--Kato empirical-process bounds; the variance-bias results in Theorems 2/4 follow from explicit projection identities (S2) and (S4), not from the target coverage statement. Q# is defined after the fact as E[Qhat] or Q* (Section 2), so E||Qhat-Q#||^2 is an input convergence rate, and the bound legitimately inherits it; the sign of sigma^2_dagger - sigma^2_# is a computed decomposition (Proposition 1), not an assumed conclusion. The only self-citation, Qiu (2024), is a passing counterexample in the introduction and is not load-bearing. The real caveats are scope/verification issues, not circularity: Condition 8 is stated as 'may hold' from Condition 7 via Lemma S6, but Lemma S6 only bounds a variance and not a sub-Weibull tail, so the fast root-n-up-to-log rate is not verified for the Super Learner/HAL pipelines in Section 5; and Section 6 concedes the empirical-process bounds may be conservative and that the variance-bias term is not first-order. These affect the strength of the model-agnostic claim but do not make any derivation equivalent to its inputs.

Axiom & Free-Parameter Ledger

3 free parameters · 7 axioms · 0 invented entities

Core results rest on standard Berry-Esseen and empirical-process facts plus high-level rate/tail assumptions on a black-box nuisance estimator. The main free input is the nuisance convergence rate/tail behavior; the paper does not estimate it. No new physical or statistical entities are postulated.

free parameters (3)
  • r̂_1(n), r̂_2(n) (and tildes) — nuisance tail-rate functions = unspecified; O(n^{-1/2}) in the best case
    Appear in Conditions 7/8/11/12; the coverage bound is only root-n up to logs if these are O(n^{-1/2}); otherwise the bound is slower. They are inputs/properties of the black-box estimator, not derived from the theory.
  • δ, δ′ balance parameters = chosen per Corollary 1 as powers of E||Q̃_a−Q#,a||_2^2 or r̃_1(n)(log n)^{1/q}
    In Theorem 3/S3, δ is a hand-chosen cut-off for the empirical-process term; the final rate depends on this choice.
  • Q# deterministic oracle approximation = either Q* or E[Q̂]
    A free 'target' function the nuisance estimator is compared against. The tightness of the bounds depends on how close Q̂ is to Q#; the paper recommends E[Q̂].
axioms (7)
  • standard math Berry-Esseen theorem for i.i.d. sample means
    Used in Lemma S3 and throughout to approximate √n P_n D_a(Q#, ψ*) by a normal distribution.
  • standard math Gaussian approximation bounds for empirical processes (Chernozhukov et al. 2014, Theorems 5.1/5.2)
    Used in Lemma S12 for non-cross-fit empirical process terms; not reproved in the paper.
  • domain assumption Condition 10: bounded uniform entropy integral / Donsker-type class F with J(1,F,M)<∞
    Needed for non-cross-fit Theorem 3/Corollary 1 to control empirical process terms.
  • ad hoc to paper Conditions 7/8 (and 11/12): approximate sub-Weibull tails of nuisance estimator
    High-level tail assumptions made to get root-n-up-to-log rates; verified only for series regression example.
  • domain assumption Condition 6: exchangeable equal-size sample splitting with identically distributed fold estimators
    Used to define Q# as common mean and to simplify cross-fit bounds.
  • domain assumption Conditions 1–5: nonzero variance, finite third moment, positivity, bounded outcome model, higher-moment bounds
    Technical conditions for Berry-Esseen, variance bounding, and boundedness of transformed outcomes.
  • domain assumption Condition 9/13: Q# is Q* or the mean of the nuisance estimator
    Yields the fast second-order bound on variance-estimator bias; Part 2 can always be satisfied by definition.

pith-pipeline@v1.3.0-alltime-deepseek · 45668 in / 16111 out tokens · 158418 ms · 2026-08-03T14:50:55.252578+00:00 · methodology

0 comments
read the original abstract

Nonparametric estimators, such as the augmented inverse probability weighted (AIPW) estimator, have become increasingly popular in causal inference. Numerous nonparametric estimators have been proposed, but they are all asymptotically normal with the same asymptotic variance under similar conditions, leaving little guidance for practitioners to choose an estimator. In this paper, I focus on another important perspective of their asymptotic behaviors beyond asymptotic normality, the convergence of the Wald-confidence interval (CI) coverage to the nominal coverage. Such results have been established for simpler estimators (e.g., the Berry-Esseen Theorem), but are lacking for nonparametric estimators. I consider a simple but practical setting where the AIPW estimator based on a black-box nuisance estimator, with or without cross-fitting, is used to estimate the average treatment effect in randomized controlled trials. I derive non-asymptotic Berry-Esseen-type bounds on the difference between Wald-CI coverage and the nominal coverage. I also analyze the bias of variance estimators, showing that the cross-fit variance estimator might overestimate while the non-cross-fit variance estimator might underestimate, which might explain why cross-fitting has been empirically observed to improve Wald-CI coverage even if both estimators converge to the same asymptotic normal distribution.

Figures

Figures reproduced from arXiv: 2512.18898 by Hongxiang Qiu.

Figure 1
Figure 1. Figure 1: Sampling distribution of ATE estimators. The horizontal blue line is the true ATE. [PITH_FULL_IMAGE:figures/full_fig_p025_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Sampling distribution of estimated influence function-based asymptotic variance for AIPW [PITH_FULL_IMAGE:figures/full_fig_p026_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: QQ-plot of AIPW estimators. The y-axis is the Monte Carlo sample quantile. The x-axis is [PITH_FULL_IMAGE:figures/full_fig_p027_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: CI coverage with 95% Wilson confidence intervals of coverage. The horizontal blue line is [PITH_FULL_IMAGE:figures/full_fig_p028_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

54 extracted references · 12 linked inside Pith

  1. [1]

    and Bearth, N

    Ballinari, D. and Bearth, N. (2024). Improving the Finite Sample Performance of Double/Debiased Machine Learning with Propensity Score Calibration . arXiv preprint arXiv:2409.04874v1

  2. [2]

    and Van Der Laan , M

    Benkeser, D. and Van Der Laan , M. (2016). The highly adaptive lasso estimator . In 2016 IEEE international conference on data science and advanced analytics (DSAA) , pages 689--696. IEEE

  3. [3]

    Bentkus, V., Gotze, F., and Tikhomirov, A. (1997). Berry-Esseen bounds for statistics of weakly dependent samples . Bernoulli , 3(3):329--349

  4. [4]

    Bhattacharya, R. N. and Ghosh, J. K. (2007). On the Validity of the Formal Edgeworth Expansion . The Annals of Statistics , 6(2):434--451

  5. [5]

    J., Gotze, F., and van Zwet, W

    Bickel, P. J., Gotze, F., and van Zwet, W. R. (2007). The Edgeworth Expansion for U -Statistics of Degree Two . The Annals of Statistics , 14(4):1463--1484

  6. [6]

    and Yu, B

    B \" u hlmann, P. and Yu, B. (2003). Boosting with the L2 loss: Regression and classification . Journal of the American Statistical Association , 98(462):324--339

  7. [7]

    Callaert, H., Janssen, P., and Veraverbeke, N. (2007). An Edgeworth Expansion for U -Statistics . The Annals of Statistics , 8(2):299--312

  8. [8]

    and White, H

    Chen, X. and White, H. (1999). Improved rates and asymptotic normality for nonparametric neural network estimators . IEEE Transactions on Information Theory , 45(2):682--691

  9. [9]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., and Newey, W. (2017). Double/debiased/Neyman machine learning of treatment effects . American Economic Review , 107(5):261--265

  10. [10]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters . Econometrics Journal , 21(1):C1--C68

  11. [11]

    Chernozhukov, V., Chetverikov, D., and Kato, K. (2014). Gaussian approximation of suprema of empirical processes . Annals of Statistics , 42(4):1564--1597

  12. [12]

    K., and Singh, R

    Chernozhukov, V., Newey, W. K., and Singh, R. (2023). A simple and general debiased machine learning theorem with finite-sample guarantees . Biometrika , 110(1):257--264

  13. [13]

    and Van Der Laan , M

    Gruber, S. and Van Der Laan , M. J. (2010). A targeted maximum likelihood estimator of a causal effect on a bounded continuous outcome . International Journal of Biostatistics , 6(1)

  14. [14]

    and Van Der Laan , M

    Gruber, S. and Van Der Laan , M. J. (2014). Targeted minimum loss based estimation of a causal effect on an outcome with known conditional bounds . International Journal of Biostatistics , 8(1)

  15. [15]

    Guo, A., Benkeser, D., and Nabi, R. (2023). Targeted Machine Learning for Average Causal Effect Estimation Using the Front-Door Functional . arXiv preprint arXiv:2312.10234v1

  16. [16]

    Hall, P. (1992). The Bootstrap and Edgeworth Expansion . Springer Series in Statistics. Springer New York, New York, NY

  17. [17]

    Hejazi, N., Coyle, J., and van der Laan, M. (2020). hal9001: Scalable highly adaptive lasso regression in R . Journal of Open Source Software , 5(53):2526

  18. [18]

    Jirak, M. (2016). Berry-esseen theorems under weak dependence . Annals of Probability , 44(3):2024--2063

  19. [19]

    Jirak, M. (2023). A Berry-Esseen bound with (almost) sharp dependence conditions . Bernoulli , 29(2):1219--1245

  20. [20]

    Kennedy, E. H. (2022). Semiparametric doubly robust targeted double machine learning: a review . arXiv preprint arXiv:2203.06469v2

  21. [21]

    Kosorok, M. R. (2008). Introduction to Empirical Processes and Semiparametric Inference , volume 77 of Springer Series in Statistics . Springer New York

  22. [22]

    Levy, J. (2018). An Easy Implementation of CV-TMLE . arXiv preprint arXiv:1811.04573v2

  23. [23]

    V., Hejazi, N

    Li, H., Rosete, S., Coyle, J., Phillips, R. V., Hejazi, N. S., Malenica, I., Arnold, B. F., Benjamin-Chung, J., Mertens, A., Colford, J. M., van der Laan, M. J., and Hubbard, A. E. (2022). Evaluating the robustness of targeted maximum likelihood estimators via realistic simulations in nutrition intervention trials . Statistics in Medicine , 41(12):2132--2165

  24. [24]

    Liu, L., Mukherjee, R., and Robins, J. M. (2020). On Nearly Assumption-Free Tests of Nominal Confidence Interval Coverage for Causal Parameters Estimated by Machine Learning . Statistical Science , 35(3):518--539

  25. [25]

    Liu, L., Mukherjee, R., and Robins, J. M. (2023). Can we falsify the justification of the validity of Wald confidence intervals of doubly robust functionals, without assumptions? arXiv preprint arXiv:2306.10590v1

  26. [26]

    Long, J. S. and Ervin, L. H. (2000). Using Heteroscedasticity Consistent Standard Errors in the Linear Regression Model . American Statistician , 54(3):217--224

  27. [27]

    Moore, K. L. and van der Laan, M. J. (2009). Covariate adjustment in randomized trials with binary outcomes: Targeted maximum likelihood estimation . Statistics in Medicine , 28(1):39--64

  28. [28]

    I., Yu, Y

    Naimi, A. I., Yu, Y. H., and Bodnar, L. M. (2024). Pseudo-Random Number Generator Influences on Average Treatment Effect Estimates Obtained with Machine Learning . Epidemiology

  29. [29]

    Neyman, J. (1923). Sur les applications de la th \' e orie des probabilit \' e s aux exp \' e riences agricoles: Essay des principles. (Excerpts reprinted and translated to English, 1990) . Statistical Science , 5:463--472

  30. [30]

    V., van der Laan, M

    Phillips, R. V., van der Laan, M. J., Lee, H., and Gruber, S. (2023). Practical considerations for specifying a super learner . International Journal of Epidemiology , 52(4):1276--1285

  31. [31]

    Qiu, H. (2024). Non-plug-in estimators could outperform plug-in estimators: a cautionary note and a diagnosis . Epidemiologic Methods , 13(1):1--11

  32. [32]

    Quintas-Martinez, V. (2022). Finite-Sample Guarantees for High-Dimensional DML . arXiv preprint arXiv:2206.07386v1

  33. [33]

    Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period-application to control of the healthy worker survivor effect . Mathematical Modelling , 7(9-12):1393--1512

  34. [34]

    Robins, J. M. and Ritov, Y. (1997). Toward a curse of dimensionality appropriate (CODA) asymptotic theory for semi-parametric models . Statistics in Medicine , 16(1-3):285--319

  35. [35]

    M., Rotnitzky, A., and Zhao, L

    Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed . Journal of the American Statistical Association , 89(427):846--866

  36. [36]

    M., Rotnitzky, A., and Zhao, L

    Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1995). Analysis of semiparametric regression models for repeated outcomes in the presence of missing data . Journal of the American Statistical Association , 90(429):106--121

  37. [37]

    and Van Der Laan , M

    Rosenblum, M. and Van Der Laan , M. J. (2009). Using regression models to analyze randomized trials: Asymptotically valid hypothesis tests despite incorrectly specified models . Biometrics , 65(3):937--945

  38. [38]

    Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies . Technical Report 5

  39. [39]

    Rytgaard, H. C. W., Eriksson, F., and Van der Laan , M. (2021). Estimation of time-specific intervention effects on continuously distributed time-to-event outcomes by targeted maximum likelihood estimation . arXiv preprint arXiv:2106.11009v1

  40. [40]

    Saco, G. (2025). Finite-Sample Failures and Condition-Number Diagnostics in Double Machine Learning . arXiv preprint arXiv:2512.07083v1

  41. [41]

    M., Song, W., Kempker, R., and Benkeser, D

    Schader, L. M., Song, W., Kempker, R., and Benkeser, D. (2024). Don't let your analysis go to seed: on the impact of random seed on machine learning-based causal inference . Epidemiology

  42. [42]

    and Ding, P

    Shi, L. and Ding, P. (2022). Berry-Esseen bounds for design-based causal inference with possibly diverging treatment levels and varying group sizes . arXiv preprint arXiv:2209.12345v4

  43. [43]

    J., Phillips, R

    Smith, M. J., Phillips, R. V., Maringe, C., and Fernandez, M. A. L. (2024). Performance of Cross-Validated Targeted Maximum Likelihood Estimation . arXiv preprint arXiv:2409.11265v1

  44. [44]

    Tran, L., Petersen, M., Schwab, J., and van der Laan, M. J. (2023). Robust variance estimation and inference for causal effect estimation . Journal of Causal Inference , 11(1)

  45. [45]

    Tran, L., Yiannoutsos, C., Wools-Kaloustian, K., Siika, A., Van Der Laan , M., and Petersen, M. (2019). Double Robust Efficient Estimators of Longitudinal Treatment Effects: Comparative Performance in Simulations and a Case Study . International Journal of Biostatistics , 15(2)

  46. [46]

    van der Laan, L., Lin, Z., Carone, M., and Luedtke, A. (2024a). Stabilized Inverse Probability Weighting via Isotonic Calibration . arXiv preprint arXiv:2411.06342v1

  47. [47]

    van der Laan, L., Luedtke, A., and Carone, M. (2024b). Automatic doubly robust inference for linear functionals via calibrated debiased machine learning . arXiv preprint arXiv:2411.02771v1

  48. [48]

    van der Laan, M. (2023). Higher Order Spline Highly Adaptive Lasso Estimators of Functional Parameters: Pointwise Asymptotic Normality and Uniform Convergence Rates . arXiv preprint arXiv:2301.13354v1

  49. [49]

    J., Polley, E

    Van Der Laan , M. J., Polley, E. C., and Hubbard, A. E. (2007). Super learner . Statistical Applications in Genetics and Molecular Biology , 6(1)

  50. [50]

    Van der Laan , M. J. and Rose, S. (2018). Targeted Learning in Data Science . Springer

  51. [51]

    and Wellner, J

    van der Vaart, A. and Wellner, J. (1996). Weak Convergence and Empirical Processes: With Applications to Statistics . Springer Series in Statistics. Springer, New York, NY

  52. [52]

    Wainwright, M. J. (2019). High-dimensional statistics: A non-asymptotic viewpoint . Cambridge University Press

  53. [53]

    Zhilova, M. (2020). Nonclassical Berry–Esseen inequalities and accuracy of the bootstrap . Annals of Statistics , 48(4):1922--1939

  54. [54]

    Zhilova, M. (2022). New Edgeworth-Type Expansions With Finite Sample Guarantees . Annals of Statistics , 50(5):2545--2561