Pith. sign in

REVIEW 2 cited by

Causal machine learning methods and use of cross-fitting in settings with high-dimensional confounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15242 v3 pith:4ODHHUQ2 submitted 2024-05-24 stat.ME

classification stat.ME
keywords confoundingcross-fittingtmleaipwcausalhigh-dimensionalimportantmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Observational epidemiological studies commonly seek to estimate the causal effect of an exposure on an outcome. Adjustment for potential confounding bias in modern studies is challenging due to the presence of high-dimensional confounding, which occurs when there are many confounders relative to sample size or complex relationships between continuous confounders and exposure and outcome. Doubly robust methods such as Augmented Inverse Probability Weighting (AIPW) and Targeted Maximum Likelihood Estimation (TMLE) have the potential to address these challenges, using data-adaptive approaches and cross-fitting, but despite recent advances limited evaluation and guidance are available on their implementation in realistic settings where high-dimensional confounding is present. Motivated by an early-life cohort study, we conducted an extensive simulation study to compare the relative performance of AIPW and TMLE using data-adaptive approaches for estimating the average causal effect (ACE). We evaluated the benefits of using cross-fitting with a varying number of folds, as well as the impact of using a reduced versus full (larger, more diverse) library in the Super Learner ensemble learning approach used for implementation. We found that AIPW and TMLE performed similarly in most cases for estimating the ACE, but TMLE was more stable. Cross-fitting improved the performance of both methods, but was more important for variance estimation and coverage than for point estimates, with the number of folds a less important consideration. Using a full Super Learner library was important to reduce bias and variance in complex scenarios typical of modern health research studies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Marginal generalized raking with parametric working models

    stat.ME 2026-07 conditional novelty 6.0 of 10

    Marginal generalized raking directly targets treatment-specific means, ATE, and RR by raking weights on the marginal efficient influence function, and is asymptotically equivalent to optimal AIPCW estimators.

  2. Efficient Estimation of Causal Effects Under Two-Phase Sampling with Error-Prone Outcome and Treatment Measurements

    stat.ME 2025-06 conditional novelty 6.0 of 10

    The authors construct and compare two asymptotically equivalent doubly robust estimators of average treatment effects under two-phase sampling with error-prone outcome and treatment, then improve finite-sample perform...

Pith tools