Pith. sign in

REVIEW 2 major objections 4 minor 57 references

Towards Best Practices for Covariate Adjustment in Regulatory Trials: From Fixed to Data-Adaptive Approaches

T0 review · 2 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read The paper argues that pre-specified, machine-learning-guided covariate adjustment is ready for primary analysis in confirmatory regulatory trials.

desk verdict A useful, well-scoped position statement on data-adaptive covariate adjustment for regulatory trials; no new technical content, and the abstract overstates the precision guarantee, but the body is careful and the paper deserves review as a perspective piece. read the letter →

arxiv 2607.27542 v1 pith:PPBGEJT3 submitted 2026-07-30 stat.ME

classification stat.ME
keywords covariateadjustmentrandomizedcontrolledtrialsdata-adaptiveestimationmachinelearningprecisionTypeIerrorcontrolpre-specificationcausalinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that randomized trials can safely gain precision by letting the data choose how to adjust for baseline covariates, as long as the selection algorithm is fully pre-specified and the unadjusted estimator remains a fallback candidate. Focusing on marginal intention-to-treat effects with minimal outcome missingness, it contends that such data-adaptive adjustment — implemented with cross-validation, doubly robust estimation, and cross-fitting — is at least as precise as the unadjusted analysis and preserves Type I error control under stated regularity conditions. The stakes: current regulatory guidance endorses only fixed parametric adjustment, so accepting these methods would let sponsors leverage machine learning in primary efficacy analyses. The paper offers practical best practices — locked analysis plans, reproducible code, transparent reporting — to make such analyses acceptable to regulators.

What carries the argument

The load-bearing mechanism is cross-validated selection among candidate estimators. A pre-specified algorithm splits the data, fits each candidate adjustment strategy — the unadjusted estimator plus covariate-adjusted or doubly robust estimators — on training folds, evaluates variance on validation folds, and picks the candidate with the smallest cross-validated variance estimate. Because the unadjusted estimator is always a candidate, the procedure defaults to it when adjustment does not help, yielding a per-trial no-precision-loss property. Cross-fitting, where nuisance functions are estimated on separate folds from the effect estimate, then supports influence-curve-based variance estimati

What would settle it

A simulation study or re-analysis of completed trials with a known true effect, using a fully pre-specified data-adaptive selection algorithm with the unadjusted estimator in the candidate set, where the resulting estimator's sampling variance exceeds that of the unadjusted estimator, or where 95% confidence intervals under-cover by more than simulation error, would refute the paper's central claim.

Watch

Extended reading notes

Core claim

The paper's central claim, on its own terms, is that data-adaptive covariate adjustment is a natural extension of established trial practice. If the adjustment strategy is chosen by a fully pre-specified algorithm that includes the unadjusted estimator as a candidate and selects among candidates by minimizing cross-validated variance, the final estimator is at least as precise as the unadjusted one, defaults to unadjusted analysis when adjustment does not help, and retains consistency and asymptotically valid inference. Four such procedures developed by the authors' working group are reviewed, unified by three requirements: data-driven selection of covariates without precision loss, flexible

Load-bearing premise

That picking the candidate with the smallest cross-validated variance estimate makes the final analysis at least as precise as the unadjusted estimator in the actual finite sample, and that the standard regularity conditions invoked for data-adaptive inference hold in the specific trial setting.

Editorial extensions

If this is right

  • If adopted, trials could pre-specify a data-adaptive primary analysis and gain power without widening confidence intervals or inflating Type I error.
  • Regulators could extend guidance beyond fixed parametric covariate adjustment to include doubly robust, machine-learning-based estimators for the same marginal estimand.
  • Routine reporting standards — locked statistical analysis plans, fixed random seeds, archived containerized code, public code release — would become the norm for adjusted analyses.
  • Trials with rich baseline or historical data would have a principled way to convert that information into narrower effect estimates.
  • The paper's focus on marginal effects with minimal missingness implies the approach applies directly to confirmatory efficacy analyses with intention-to-treat estimands.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 'at least as precise' guarantee is stated per-trial based on cross-validated selection, but the theoretical results cited for several of the procedures are asymptotic; a finite-sample theorem showing cross-validated variance selection does not select a worse candidate than unadjusted would make the guarantee unconditional.
  • The approach's validity relies on the 'standard regularity conditions' for data-adaptive inference; before broad regulatory acceptance, those conditions would need to be verified in concrete trial settings such as small samples, rare outcomes, or many covariates.
  • The paper treats minimal outcome missingness and a target population sampled from a larger population; extending these guarantees to substantial missingness or to intercurrent-event estimands would require separate development, which the paper explicitly leaves open.
  • A testable extension: run the proposed pre-specified adaptive procedures on a large set of completed trial datasets where the unadjusted analysis is the reference, and check empirically whether the adaptive estimator's finite-sample variance ever exceeds the unadjusted estimator's variance, as a stress test of the 'no bad bets' property.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This perspective paper argues that fully pre-specified, data-adaptive covariate adjustment—via TMLE with Adaptive Pre-specification, Super Learner-based selection, COADVISE, and H-AIPW—can improve precision in confirmatory randomized trials while preserving validity, estimand, and Type I error control. It reviews unadjusted and fixed-adjustment estimators, explains data-adaptive alternatives in non-technical terms, and offers practical recommendations on estimand alignment, pre-specification, reproducibility, cross-fitting, simulations, and communication. The paper is written for a broad regulatory and clinical-trials audience and defers technical details to appendices and cited references.

Significance. If the claims are accepted, the paper could serve as a bridge between current FDA/EMA guidance on fixed covariate adjustment and modern machine-learning-based approaches, encouraging broader regulatory acceptance. Its strengths include a clear scope, an explicit estimand-focused framework, and concrete practical guidance on pre-specification, code locking, and cross-fitting. However, the central 'guaranteed to improve precision' claim is overstated relative to the body's own caveats, and the evidence base for the four recommended procedures comes largely from the working group's own methodological papers, with no independent simulation or empirical validation at regulatory trial scales. These issues make the paper's strongest conclusion—that regulators should accept these methods as primary analyses—not yet fully supported.

major comments (2)
  1. [Abstract; 'Avoid Bad Bets' section; 'Data-Adaptive Adjustment' section] The abstract states that the recommended procedures are 'guaranteed to improve precision relative to unadjusted analyses.' The only formal guarantee offered in 'Avoid Bad Bets' is that, if the unadjusted estimator is a candidate, the selected estimator's cross-validated variance estimate is no larger than the unadjusted estimator's. This is a property of the selection rule, not a guarantee about the true finite-sample variance (or MSE). Indeed, the 'Data-Adaptive Adjustment' section explicitly concedes that such approaches 'are still not guaranteed to yield an estimator with lower variance than the unadjusted estimator.' The abstract and conclusion should be reworded to say that the procedures are designed to avoid precision loss on the selection criterion, or that they are guaranteed not to increase the cross-validated variance estimate, but not that they guarantee improved precision in
  2. ['Data-Adaptive Adjustment' first paragraph; Appendix B; Conclusion] The paper states that data-adaptive adjustment 'does not inflate Type I error, provided the above regularity conditions are met,' but the manuscript never states those conditions. Appendix B refers to 'asymptotically exact inference' under 'very general conditions,' and the conclusion asserts unconditionally that these methods 'preserve ... control of Type I error.' Because the recommendation for regulatory acceptance rests on this claim, the manuscript should either list the key conditions (e.g., cross-fitting, Donsker-type or convergence conditions, and conditions on the selection step) or explicitly qualify the conclusion as holding only when those asymptotic conditions are satisfied. As written, a regulator could read the conclusion as a finite-sample guarantee that is not established.
minor comments (4)
  1. ['Avoid Bad Bets' section] The sentence 'the approaches using Adaptive Pre-specification and Super Learner are guaranteed for each trial analysis to be at least as precise as the unadjusted approach for the chosen effect and variance criterion' is tautological; consider rephrasing to avoid implying more than the criterion-based guarantee.
  2. ['Pre-specification or Bust!' section] The phrase 'not influenced by knowledge of the treatment effect' appears to be a typo; the intended meaning is likely 'not influenced by knowledge of the treatment assignment' or 'by the unblinded treatment data.'
  3. ['Don't Double-Dip; Cross-fit' section] The suggestion to 'limit the candidates to working GLMs adjusting for one covariate' cites references [21,22], which concern prognostic scores; the connection between a single-covariate GLM and the cited references should be clarified.
  4. [End of manuscript] There is an unmatched closing parenthesis in the sentence 'The authors report generative AI was not used in their research or preparation of this manuscript).'

Circularity Check

1 steps flagged · score 6.0 of 10

The advertised precision guarantee is the selection rule restated; the abstract drops the variance-criterion qualifier, making a load-bearing claim reduce by construction.

  1. self definitional [Abstract; 'Avoid Bad Bets' section (p. 12); cf. 'Data-Adaptive Adjustment' section (p. 10)]
    "As long as the unadjusted estimator is included as a candidate, the approaches using Adaptive Pre-specification and Super Learner are guaranteed for each trial analysis to be at least as precise as the unadjusted approach for the chosen effect and variance criterion; specifically, they default to the unadjusted approach if none of the candidates using covariate adjustment reduce the cross-validated variance estimate."

    The guarantee is the algorithm's selection rule: by choosing the candidate with the smallest cross-validated variance estimate and including the unadjusted estimator among the candidates, the selected candidate's cross-validated variance estimate cannot exceed the unadjusted candidate's. Thus 'at least as precise for the chosen ... variance criterion' is true by construction, not an independent finite-sample guarantee about actual estimator variance. The paper itself concedes this two pages earlier: 'they are still not guaranteed to yield an estimator with lower variance than the unadjusted estimator.' The abstract nevertheless advertises analyses 'guaranteed to improve precision relative to unadjusted analyses' without the qualifier, so the central precision claim reduces to a restatement

full rationale

The paper is a perspective/review rather than a derivation, so most of its content is not circular: fixed and data-adaptive estimators, influence-curve inference, and the regulatory discussion are grounded in external literature and FDA/EMA guidance. The principal circular element is the precision guarantee in the abstract and the 'Avoid Bad Bets' section. The guarantee holds only for the cross-validated variance estimate used as the selection criterion; the paper explicitly acknowledges there is no guarantee of lower true variance. Because the abstract and the practical recommendation that regulators should accept these methods depend on the unconditional 'guaranteed to improve precision' language, this is a partial reduction of a load-bearing claim to the construction of the selection rule. There is also substantial self-citation (refs 6, 8, 9, 10, 41, 43, 51) for the four recommended procedures and for the 'no instances as harm' and 'asymptotically exact inference' claims; I do not score these as independent circularity because they cite prior method papers rather than deriving the current conclusion from itself, but they do mean the manuscript's supportive evidence is largely internal to the author group. The paper also flags its own limitation in the Data-Adaptive Adjustment section, which I weigh in the score rather than treating the abstract's unconditional guarantee as fully supported.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities or fitted parameters. The main hidden load is the unstated regularity condition for data-adaptive inference and the assumption that minimizing a cross-validated variance estimate provides an actual precision guarantee.

assumptions (5)
  • domain assumption Randomization ensures identification of marginal effects by observed group means and makes covariate adjustment a precision-only tool.
    The entire framework depends on treatment assignment being randomized and unconfounded; stated in the 'Unadjusted Analysis' section as design-based justification.
  • domain assumption Outcome missingness is minimal and missing outcomes are representative (ignorable).
    The scope section explicitly limits the paper to settings with minimal outcome missingness and assumes measured outcomes represent missing outcomes; this is load-bearing for the validity claims.
  • domain assumption Data-adaptive nuisance estimators satisfy 'standard regularity conditions' (no overfitting, sufficient convergence) so that influence-curve inference is asymptotically valid.
    Invoked in 'Data-Adaptive Adjustment' and Appendix A; these conditions are not verified for specific trials and are necessary for the Type I error claims.
  • domain assumption Cross-fitting is sufficient to prevent overfitting and keep the influence-curve variance estimator valid.
    The cross-fitting recommendation in 'Don't Double-Dip; Cross-fit' assumes the selected learners converge to some limit, which the paper acknowledges is not guaranteed in practice when many covariates or very adaptive learners are used.
  • standard math The prediction-unbiasedness condition suffices for consistency of G-computation with a working GLM.
    Correct for canonical-link GLMs fit by MLE or quasi-likelihood in randomized trials; stated in the 'Fixed Adjustment' section.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Best Practices for Covariate Adjustment in Regulatory Trials: From Fixed to Data-Adaptive Approaches." pith.science (2026). https://pith.science/paper/PPBGEJT3

@misc{pith2026260727542,
  author       = {Pith},
  title        = {Pith review of: Towards Best Practices for Covariate Adjustment in Regulatory Trials: From Fixed to Data-Adaptive Approaches},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PPBGEJT3}},
  note         = {Machine review of arXiv:2607.27542}
}
read the original abstract

While randomization justifies the use of unadjusted effect estimators in randomized trials, there is growing interest in covariate adjustment to improve precision. Adjusting for baseline variables that are prognostic of the outcome can reduce estimator variance, resulting in narrower confidence intervals and increased statistical power. Recent guidance by the U.S. Food and Drug Administration supports fixed adjustment for prognostic covariates using parametric regression models. However, this guidance does not address more flexible approaches using data-adaptive or machine learning methods. We offer our perspectives on covariate adjustment to improve analytic precision. We focus on estimating the average effect for the target population in trials with minimal outcome missingness. We provide a non-technical overview of effect estimators that are unadjusted and effect estimators using fixed versus data-adaptive adjustment. We offer practical suggestions for conducting adjusted analyses that are data-adaptive, fully pre-specified, transparently and reproducibly implemented, robust to model misspecification, and guaranteed to improve precision relative to unadjusted analyses --- all while preserving statistical validity and the causal effect of interest. We hope that sharing our perspectives will foster broader discussion and eventual acceptance of principled, pre-specified, data-adaptive covariate adjustment in randomized trials.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

57 extracted references · 17 canonical work pages

  1. [1]

    Regress the outcome 𝑌 on randomized assignment 𝐴 and the prognostic covariate 𝑉 according to a pre-specified model for the conditional mean outcome 𝔼(𝑌|𝐴, 𝑉). Common choices include generalized linear models (GLMs) fit by maximum likelihood estimation, such as linear regression for continuous outcomes, logistic regression for binary outcomes, and a log-li...

  2. [2]

    Using the regression fit, predict the outcome for all 𝑁 participants under assignment to the intervention group and under assignment to the control group: 𝔼̂(𝑌|𝐴 = 1, 𝑉𝑖) = 𝑔−1(𝛽̂0 + 𝛽̂1 + 𝛽̂2𝑉𝑖) and 𝔼̂(𝑌|𝐴 = 0, 𝑉𝑖) = 𝑔−1(𝛽̂0 + 𝛽̂2𝑉𝑖) for 𝑖 = 1, . . . , 𝑁

  3. [3]

    locally efficient

    Average these predictions across all 𝑁 participants to obtain marginal mean estimates under each assignment strategy: 1 𝑁 ∑ 𝔼̂(𝑌|𝐴 = 1, 𝑉𝑖)𝑁 𝑖=1 versus 1 𝑁 ∑ 𝔼̂(𝑌|𝐴 = 0, 𝑉𝑖)𝑁 𝑖=1 . Then contrast on the scale of interest. We refer to Appendix A for an overview of obtaining statistical inference. The regression model in Step 1 is viewed as a “working” model...

  4. [4]

    Covariate adjustment for two-sample treatment comparisons in randomized clinical trials: A principled yet flexible approach

    Tsiatis AA, Davidian M, Zhang M, Lu X. Covariate adjustment for two-sample treatment comparisons in randomized clinical trials: A principled yet flexible approach. Stat Med. 2008;27(23):4658-4677. doi:10.1002/sim.3113

  5. [5]

    Simple, Efficient Estimators of Treatment Effects in Randomized Trials Using Generalized Linear Models to Leverage Baseline Variables

    Rosenblum M, van der Laan MJ. Simple, Efficient Estimators of Treatment Effects in Randomized Trials Using Generalized Linear Models to Leverage Baseline Variables. Int J Biostat. 2010;6(1):Article 13. doi:10.2202/1557-4679.1138

  6. [6]

    Improving Precision and Power in Randomized Trials for COVID-19 Treatments Using Covariate Adjustment, for Binary, Ordinal, and Time-to-Event Outcomes

    Benkeser D, Díaz I, Luedtke A, Segal J, Scharfstein D, Rosenblum M. Improving Precision and Power in Randomized Trials for COVID-19 Treatments Using Covariate Adjustment, for Binary, Ordinal, and Time-to-Event Outcomes. Biometrics. 2021;77(4):1467-1481

  7. [7]

    Covariate adjustment in randomized controlled trials: General concepts and practical considerations

    Van Lancker K, Bretz F, Dukes O. Covariate adjustment in randomized controlled trials: General concepts and practical considerations. Clin Trials. 2024;21(4):399-411. doi:10.1177/17407745241251568

  8. [8]

    Statistical Methods for Research Workers

    Fisher RA. Statistical Methods for Research Workers. 4th ed. Oliver and Boyd Ltd.; 1932

Show all 57 references
  1. [9]

    Machine learning methods for leveraging baseline covariate information to improve the efficiency of clinical trials

    Zhang Z, Ma S. Machine learning methods for leveraging baseline covariate information to improve the efficiency of clinical trials. Stat Med. 2019;38(10):1703-1714. doi:10.1002/sim.8054

  2. [10]

    Optimising precision and power by machine learning in randomised trials with ordinal and time-to-event outcomes with an application to COVID-19

    Williams N, Rosenblum M, Díaz I. Optimising precision and power by machine learning in randomised trials with ordinal and time-to-event outcomes with an application to COVID-19. J R Stat Soc Ser A Stat Soc. Published online September 23, 2022. doi:10.1111/rssa.12915

  3. [11]

    Adaptive selection of the optimal strategy to improve precision and power in randomized trials

    Balzer LB, Cai E, Godoy Garraza L, Amaranath P. Adaptive selection of the optimal strategy to improve precision and power in randomized trials. Biometrics. 2024;80(1):ujad034. doi:10.1093/biomtc/ujad034 arXiv v1 (29-July 2026) 20

  4. [12]

    COADVISE: covariate adjustment with variable selection in randomized controlled trials

    Liu Y, Zhu K, Han L, Yang S. COADVISE: covariate adjustment with variable selection in randomized controlled trials. J R Stat Soc Ser A Stat Soc. Published online November 4, 2025:qnaf171. doi:10.1093/jrsssa/qnaf171

  5. [13]

    Efficient Randomized Experiments Using Foundation Models

    De Bartolomeis P, Abad J, Wang G, et al. Efficient Randomized Experiments Using Foundation Models. Adv Neural Inf Process Syst. 2025;38

  6. [14]

    Statistical Principles for Clinical Trials E9

    ICH. Statistical Principles for Clinical Trials E9. 1998

  7. [15]

    Guideline on Adjustment for Baseline Covariates in Clinical Trials

    EMA. Guideline on Adjustment for Baseline Covariates in Clinical Trials. 2015

  8. [16]

    ICH E9 (R1) Addendum on Estimands and Sensitivity Analysis in Clinical Trials to the Guideline on Statistical Principles for Clinical Trials

    ICH. ICH E9 (R1) Addendum on Estimands and Sensitivity Analysis in Clinical Trials to the Guideline on Statistical Principles for Clinical Trials. European Medicines Agency; 2020

  9. [17]

    Adjusting for Covariates in Randomized Clinical Trials for Drugs and Biological Products Guidance for Industry

    FDA. Adjusting for Covariates in Randomized Clinical Trials for Drugs and Biological Products Guidance for Industry. 2023. https://www.fda.gov/media/148910/download

  10. [18]

    Lasso adjustments of treatment effect estimates in randomized experiments

    Bloniarz A, Liu H, Zhang CH, Sekhon JS, Yu B. Lasso adjustments of treatment effect estimates in randomized experiments. Proc Natl Acad Sci. 2016;113(27):7383-7390. doi:10.1073/pnas.1510506113

  11. [19]

    The LOOP Estimator: Adjusting for Covariates in Randomized Experiments

    Wu E, Gagnon-Bartsch JA. The LOOP Estimator: Adjusting for Covariates in Randomized Experiments. Eval Rev. 2018;42(4):458-488. doi:10.1177/0193841X18808003

  12. [20]

    The Generalized Oaxaca-Blinder Estimator

    Guo K, Basse G. The Generalized Oaxaca-Blinder Estimator. J Am Stat Assoc. 2023;118(541):524-536. doi:10.1080/01621459.2021.1941053

  13. [21]

    A family of Bayesian prognostic and predictive covariate-adjusted response-adaptive randomization designs

    Pei X, Zhao Y, Yu J, Wang L, Zhu H. A family of Bayesian prognostic and predictive covariate-adjusted response-adaptive randomization designs. Stat Methods Med Res. 2025;34(9):1838-1850. doi:10.1177/09622802251335150

  14. [22]

    Making apples from oranges: Comparing noncollapsible effect estimators and their standard errors after adjustment for different covariate sets

    Daniel R, Zhang J, Farewell D. Making apples from oranges: Comparing noncollapsible effect estimators and their standard errors after adjustment for different covariate sets. Biom J. 2021;63(3):528-557. doi:10.1002/bimj.201900297

  15. [23]

    Asymptotic Statistics

    van der Vaart AW. Asymptotic Statistics. Cambridge University Press; 1998

  16. [24]

    Stratification by a multivariate confounder score

    Miettinen OS. Stratification by a multivariate confounder score. Am J Epidemiol. 1976;104(6):609-620. doi:10.1093/oxfordjournals.aje.a112339

  17. [25]

    Increasing the efficiency of randomized trial estimates via linear adjustment for a prognostic score

    Schuler A, Walsh D, Hall D, et al. Increasing the efficiency of randomized trial estimates via linear adjustment for a prognostic score. Int J Biostat. 2022;18(2):329-356. doi:10.1515/ijb- 2021-0072

  18. [26]

    A new approach to causal inference in mortality studies with sustained exposure periods–application to control of the healthy worker survivor effect

    Robins JM. A new approach to causal inference in mortality studies with sustained exposure periods–application to control of the healthy worker survivor effect. Math Model. 1986;7:1393-1512. doi:10.1016/0270-0255(86)90088-6

  19. [27]

    Adjusting for Nonignorable Drop-Out Using Semiparametric Nonresponse Models (with Rejoiner)

    Scharfstein DO, Rotnitzky A, Robins JM. Adjusting for Nonignorable Drop-Out Using Semiparametric Nonresponse Models (with Rejoiner). J Am Stat Assoc. 1999;94(448):1096- 1120 (1135-1146). doi:10.2307/2669930 arXiv v1 (29-July 2026) 21

  20. [28]

    Inverse probability weighted estimation for general missing data problems

    Wooldridge JM. Inverse probability weighted estimation for general missing data problems. J Econom. 2007;141(2):1281-1301. doi:10.1016/j.jeconom.2007.02.002

  21. [29]

    A general form of covariate adjustment in clinical trials under covariate-adaptive randomization

    Bannick MS, Shao J, Liu J, Du Y, Yi Y, Ye T. A general form of covariate adjustment in clinical trials under covariate-adaptive randomization. Biometrika. 2025;112(3):asaf029. doi:10.1093/biomet/asaf029

  22. [30]

    Adjusting for partially missing baseline measurements in randomized trials

    White IR, Thompson SG. Adjusting for partially missing baseline measurements in randomized trials. Stat Med. 2005;24(7):993-1007. doi:10.1002/sim.1981

  23. [31]

    To Adjust or not to Adjust? Estimating the Average Treatment Effect in Randomized Experiments with Missing Covariates

    Zhao A, Ding P. To Adjust or not to Adjust? Estimating the Average Treatment Effect in Randomized Experiments with Missing Covariates. J Am Stat Assoc. 2024;119(545):450-

  24. [32]

    On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects

    Hahn J. On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects. Econometrica. 1998;2:315-331. doi:10.2307/2998560

  25. [33]

    Estimation of regression coefficients when some regressors are not always observed

    Robins JM, Rotnitzky A, Zhao LP. Estimation of regression coefficients when some regressors are not always observed. J Am Stat Assoc. 1994;89(427):846-866

  26. [34]

    Unified Methods for Censored Longitudinal Data and Causality

    van der Laan MJ, Robins JM. Unified Methods for Censored Longitudinal Data and Causality. Springer-Verlag; 2003

  27. [35]

    Variance reduction in randomised trials by inverse probability weighting using the propensity score

    Williamson EJ, Forbes A, White IR. Variance reduction in randomised trials by inverse probability weighting using the propensity score. Stat Med. 2014;33(5):721-737. doi:10.1002/sim.5991

  28. [36]

    Cross-Validated Targeted Minimum-Loss-Based Estimation

    Zheng W, van der Laan MJ. Cross-Validated Targeted Minimum-Loss-Based Estimation. In: van der Laan MJ, Rose S, eds. Targeted Learning: Causal Inference for Observational and Experimental Data. Springer Series in Statistics. Springer; 2011:459-474

  29. [37]

    Targeted Maximum Likelihood Learning

    van der Laan MJ, Rubin DB. Targeted Maximum Likelihood Learning. Int J Biostat. 2006;2(1):Article 11. doi:10.2202/1557-4679.1043

  30. [38]

    Targeted Learning: Causal Inference for Observational and Experimental Data

    van der Laan M, Rose S. Targeted Learning: Causal Inference for Observational and Experimental Data. Springer; 2011

  31. [39]

    Machine learning in the estimation of causal effects: targeted minimum loss-based estimation and double/debiased machine learning

    Díaz I. Machine learning in the estimation of causal effects: targeted minimum loss-based estimation and double/debiased machine learning. Biostatistics. 2019;21(2):353-358. doi:10.1093/biostatistics/kxz042

  32. [40]

    Empirical efficiency maximization: improved locally efficient covariate adjustment in randomized experiments and survival analysis

    Rubin DB, van der Laan MJ. Empirical efficiency maximization: improved locally efficient covariate adjustment in randomized experiments and survival analysis. Int J Biostat. 2008;4(1):Article 5. doi:10.2202/1557-4679.1084

  33. [41]

    Double/debiased machine learning for treatment and structural parameters

    Chernozhukov V, Chetverikov D, Demirer M, et al. Double/debiased machine learning for treatment and structural parameters. Published online 2018

  34. [42]

    Efficient and Adaptive Estimation for Semiparametric Models

    Bickel PJ, Klaassen CAJ, Ritov Y, Wellner JA. Efficient and Adaptive Estimation for Semiparametric Models. Johns Hopkins University Press; 1993

  35. [43]

    Semiparametric Theory and Missing Data

    Tsiatis AA. Semiparametric Theory and Missing Data. Springer; 2006. arXiv v1 (29-July 2026) 22

  36. [44]

    Some Surprising Results about Covariate Adjustment in Logistic Regression Models

    Robinson LD, Jewell NP. Some Surprising Results about Covariate Adjustment in Logistic Regression Models. Int Stat Rev. 1991;59(2):227-240

  37. [45]

    Adaptive Pre-specification in Randomized Trials With and Without Pair-Matching

    Balzer L, van der Laan MJ, Petersen M, SEARCH Collaboration. Adaptive Pre-specification in Randomized Trials With and Without Pair-Matching. Stat Med. 2016;35(10):4528-4545

  38. [46]

    Super Learner

    van der Laan MJ, Polley EC, Hubbard AE. Super Learner. Stat Appl Genet Mol Biol. 2007;6(1):Article 25. doi:10.2202/1544-6115.1309

  39. [47]

    Machine learning to optimize precision in the analysis of randomized trials: A journey in pre-specified, yet data-adaptive learning

    Balzer LB, van der Laan MJ, Petersen ML. Machine learning to optimize precision in the analysis of randomized trials: A journey in pre-specified, yet data-adaptive learning. Clin Trials. Published online February 28, 2026. doi:10.1177/17407745261417227

  40. [48]

    Trial Emulation, Simulation, and Augmentation Using Electronic Health Records and Generative AI

    Dahabreh IJ, Yeh RW, De Bartolomeis P. Trial Emulation, Simulation, and Augmentation Using Electronic Health Records and Generative AI. NEJM AI. 2025;2(10):AIe2500894. doi:10.1056/AIe2500894

  41. [49]

    Coadvise

    Liu Y. Coadvise. Published online 2024. Accessed February 15, 2026. https://github.com/yiliu1998/Coadvise

  42. [50]

    RobinCar2: ROBust INference for Covariate Adjustment in Randomized Clinical Trials

    Li L, Bannick M, Bove DS, et al. RobinCar2: ROBust INference for Covariate Adjustment in Randomized Clinical Trials. Published online January 9, 2026. Accessed February 15,

  43. [51]

    Automated, efficient and model-free inference for randomized clinical trials via data-driven covariate adjustment

    Van Lancker K, Díaz I, Vansteelandt S. Automated, efficient and model-free inference for randomized clinical trials via data-driven covariate adjustment. arXiv. Preprint posted online April 17, 2024:arXiv:2404.11150. doi:10.48550/arXiv.2404.11150

  44. [52]

    The Causal Roadmap and Simulations to Improve the Rigor and Reproducibility of Real-data Applications

    Nance N, Petersen ML, van der Laan M, Balzer LB. The Causal Roadmap and Simulations to Improve the Rigor and Reproducibility of Real-data Applications. Epidemiol Camb Mass. 2024;35(6):791-800. doi:10.1097/EDE.0000000000001773

  45. [54]

    Considerations for the Integration of Randomized Controlled Trials and Real-World Data

    Qiu S, Barr C, Dang L, et al. Considerations for the Integration of Randomized Controlled Trials and Real-World Data. Preprint posted online April 11, 2026:arXiv:2604.10308. Accessed July 29, 2026. https://arxiv.org/abs/2604.10308v1

  46. [55]

    Robust integration of external control data in randomized trials

    Karlsson R, Wang G, De Bartolomeis P, Krijthe JH, Dahabreh IJ. Robust integration of external control data in randomized trials. arXiv.org. June 25, 2024. Accessed July 29,

  47. [56]

    https://arxiv.org/abs/2406.17971v4

  48. [460]

    doi:10.1080/01621459.2022.2123814

  49. [2026]

    https://cran.r-project.org/web/packages/RobinCar2/index.html

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.