Pith. sign in

REVIEW 2 major objections 1 minor 46 references

On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data

T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read Distinguishing mediator from non-mediator auxiliary variables prevents bias in multiple imputation for average causal effect estimation.

desk verdict Paper shows mediator vs non-mediator auxiliaries matter for bias in MI for ACE but its m-DAGs may not cover typical cases. read the letter →

arxiv 2606.22016 v1 pith:QEHBBXNR submitted 2026-06-20 stat.ME

classification stat.ME
keywords multipleimputationauxiliaryvariablesaveragecausaleffectmissingdatainferencemissingnessDAGsg-computationrecoverability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper examines how auxiliary variables can help recover the average causal effect when data are missing. It uses missingness DAGs to classify univariable and multivariable mechanisms. Simulations compare multiple imputation strategies that incorporate these variables differently. The results show that mediator auxiliary variables and non-mediator ones affect bias differently, and only compatible flexible imputation models succeed across settings.

What carries the argument

Missingness directed acyclic graphs (m-DAGs) that depict the mechanisms, used to derive recoverability results for the ACE with auxiliary variables in MI.

What would settle it

A simulation or empirical analysis in which mediator auxiliary variables are included in standard MI without compatibility checks produces persistent bias in the ACE estimate relative to a gold-standard complete-data analysis.

Watch

Extended reading notes

Core claim

For a range of missingness mechanisms, the average causal effect is recoverable using multiple imputation only when auxiliary variables are incorporated in a manner compatible with the analysis model, with mediator auxiliary variables requiring particular care to prevent bias in g-computation estimates.

Load-bearing premise

The missingness directed acyclic graphs considered in the paper represent the typical missingness mechanisms encountered when estimating average causal effects.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript claims that distinguishing between mediator and non-mediator auxiliary variables is crucial to avoid bias when using multiple imputation (MI) to estimate the average causal effect (ACE) with missing data, and that compatible and flexible non-parametric MI methods incorporating these variables are necessary. It derives recoverability results for a range of univariable and multivariable missingness mechanisms using m-DAGs and evaluates them through simulation studies comparing MI-based and complete-case methods with g-computation.

Significance. If the results hold, the paper provides important practical guidance on the use of auxiliary variables in MI for causal inference, filling a gap in the literature regarding auxiliaries necessary for identifiability of the ACE. Strengths include the derivation of recoverability results across multiple m-DAGs and the evaluation of different MI strategies under correctly specified g-computation. This could improve the reliability of ACE estimates in observational data with missingness.

major comments (2)
  1. [m-DAGs and recoverability derivations] m-DAGs and recoverability section: The recoverability results and bias-avoidance guidance are derived exclusively under the specific set of m-DAGs considered; the manuscript does not provide evidence or discussion that these m-DAGs are representative of typical missingness mechanisms in applied causal inference settings (e.g., those involving unmeasured common causes of missingness indicators or time-varying mechanisms). This is load-bearing for the generalizability of the recommendations.
  2. [Simulation studies] Simulation studies section: The simulations rely on the chosen m-DAGs and correctly specified models; no sensitivity analyses are presented for misspecified MI models or deviations from the assumed missingness structures, which could affect the conclusion that non-parametric MI is required to avoid bias.
minor comments (1)
  1. [Abstract] The abstract could more explicitly state the number of m-DAGs considered and summarize the key quantitative findings from the simulations.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive comments on our manuscript. We address the two major comments point by point below, focusing on the scope of our m-DAGs and the design of the simulation studies. We propose targeted revisions to improve clarity on limitations while preserving the core contributions on recoverability and MI compatibility.

read point-by-point responses
  1. Referee: m-DAGs and recoverability section: The recoverability results and bias-avoidance guidance are derived exclusively under the specific set of m-DAGs considered; the manuscript does not provide evidence or discussion that these m-DAGs are representative of typical missingness mechanisms in applied causal inference settings (e.g., those involving unmeasured common causes of missingness indicators or time-varying mechanisms). This is load-bearing for the generalizability of the recommendations.

    Authors: We selected the m-DAGs to cover a range of univariable and multivariable mechanisms commonly discussed in the missing data and causal inference literature (e.g., missingness depending on observed covariates or outcomes). The manuscript already frames these as 'typical' based on prior work. We agree that mechanisms with unmeasured common causes of missingness indicators or time-varying structures fall outside this scope and could affect generalizability. We will add a dedicated limitations paragraph in the discussion section explicitly noting the selected m-DAGs, their motivation from existing literature, and the value of future extensions to more complex mechanisms. revision: partial

  2. Referee: Simulation studies section: The simulations rely on the chosen m-DAGs and correctly specified models; no sensitivity analyses are presented for misspecified MI models or deviations from the assumed missingness structures, which could affect the conclusion that non-parametric MI is required to avoid bias.

    Authors: The simulations were intentionally conducted under correctly specified g-computation and the assumed m-DAGs to isolate the impact of MI compatibility and flexibility on bias when recoverability holds. This design directly supports the recoverability derivations. We acknowledge that misspecification of MI models or deviations from the m-DAGs could alter performance and that sensitivity analyses would provide additional practical insight. We will expand the discussion to address this limitation and note that the current results demonstrate bias avoidance under compatible flexible MI when the substantive model is correct. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: recoverability results and simulations are independent of fitted quantities

full rationale

The paper derives recoverability results from the structure of specified m-DAGs using standard graphical criteria for missing data, then evaluates MI and complete-case estimators via simulation studies that apply correctly specified g-computation on data generated from those same DAGs. No prediction or result reduces by construction to a parameter fitted from the target data; the simulations serve as external benchmarks rather than self-referential fits. No load-bearing self-citation chains or ansatz smuggling are present. The derivation chain remains self-contained against the paper's stated mechanisms.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claims rest on domain assumptions about the m-DAGs representing real missingness processes and on the correctness of the g-computation specification; no free parameters or invented entities are described in the abstract.

assumptions (1)
  • domain assumption The m-DAGs considered correctly represent typical univariable and multivariable missingness mechanisms for ACE estimation.
    Invoked to derive recoverability results across the range of settings examined.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data." pith.science (2026). https://pith.science/paper/QEHBBXNR

@misc{pith2026260622016,
  author       = {Pith},
  title        = {Pith review of: On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QEHBBXNR}},
  note         = {Machine review of arXiv:2606.22016}
}
read the original abstract

Estimating the average causal effect (ACE) using observational data is a key focus in causal inference for which missing data present an important challenge. Multiple imputation (MI) is a widely used method for handling missing data and can yield unbiased estimates when the imputation is compatible with the substantive analysis. One of the advantages of MI is its scope to include so-called "auxiliary variables", defined as variables associated with incomplete variables that are excluded from the substantive analysis. Although many studies have looked at the use of auxiliary variables in MI for improving precision, the study of auxiliary variables that are necessary for the identifiability (or "recoverability") of the ACE in the presence of missing data has been scant. In this work, we investigate the use of auxiliary variables, both mediators and non-mediators, across a range of typical univariable and multivariable missingness mechanisms depicted by missingness directed acyclic graphs (m-DAGs). For each setting, we derive recoverability results, then evaluate MI-based and complete-case methods for estimating the ACE using correctly specified g-computation, considering different strategies for incorporating auxiliary variables and varying degrees of compatibility for MI models. Based on findings from the simulation studies, we provide practical guidance, highlighting that distinguishing appropriately between mediator and non-mediator auxiliary variables is important to avoid bias as is the use of compatible and flexible (non-parametric) MI methods that incorporate these variables.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references

  1. [1]

    Chapman & Hall/CRC, Boca Ratonn, 2020

    Robins JM Hernán MA.Causal Inference: What If. Chapman & Hall/CRC, Boca Ratonn, 2020

  2. [2]

    A new approach to causal inference in mortality studies with a sus- tained exposure period—application to control of the healthy worker survivor effect

    James Robins. A new approach to causal inference in mortality studies with a sus- tained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling, 7(9-12):1393–1512, 1986

  3. [3]

    Implementation of g- computation on a simulated data set: demonstration of a causal inference technique

    Jonathan M Snowden, Sherri Rose, and Kathleen M Mortimer. Implementation of g- computation on a simulated data set: demonstration of a causal inference technique. American journal of epidemiology, 173(7):731–738, 2011

  4. [4]

    Graphical models for recovering probabilistic and causal queries from missing data.Advances in Neural Information Processing Systems, 27, 2014

    Karthika Mohan and Judea Pearl. Graphical models for recovering probabilistic and causal queries from missing data.Advances in Neural Information Processing Systems, 27, 2014

  5. [5]

    Canonical causal diagrams to guide the treatment of missing data in epidemiologic studies.American journal of epidemiology, 187(12): 2705–2715, 2018

    Margarita Moreno-Betancur, Katherine J Lee, Finbarr P Leacy, Ian R White, Julie A Simpson, and John B Carlin. Canonical causal diagrams to guide the treatment of missing data in epidemiologic studies.American journal of epidemiology, 187(12): 2705–2715, 2018

  6. [6]

    Recoverability and estimation of causal effects under typical multivariable missingness mechanisms.Biometrical Journal, 66(3):2200326, 2024

    Jiaxin Zhang, S Ghazaleh Dashti, John B Carlin, Katherine J Lee, and Margarita Moreno-Betancur. Recoverability and estimation of causal effects under typical multivariable missingness mechanisms.Biometrical Journal, 66(3):2200326, 2024

  7. [7]

    canonical causal diagrams to guide the treatment of missing data in epidemiologic studies

    Margarita Moreno-Betancur, Katherine J Lee, Finbarr P Leacy, Julie A Simpson, and John B Carlin. Correction to:“canonical causal diagrams to guide the treatment of missing data in epidemiologic studies”.American Journal of Epidemiology, page kwae406, 2025

  8. [8]

    Mediation analysis with the mediator and outcome missing not at random.Journal of the American Statistical Association, 120(550):794–804, 2025

    Shuozhi Zuo, Debashis Ghosh, Peng Ding, and Fan Yang. Mediation analysis with the mediator and outcome missing not at random.Journal of the American Statistical Association, 120(550):794–804, 2025

Show all 46 references
  1. [9]

    Katherine J Lee, John B Carlin, Julie A Simpson, and Margarita Moreno-Betancur. Assumptions and analysis planning in studies with missing data in multiple vari- ables: moving beyond the mcar/mar/mnar classification.International Journal of Epidemiology, 52(4):1268–1275, 2023

  2. [10]

    John Wiley & Sons, 2004

    Donald B Rubin.Multiple imputation for nonresponse in surveys, volume 81. John Wiley & Sons, 2004

  3. [11]

    John Wiley & Sons, 2023

    James R Carpenter, Jonathan W Bartlett, Tim P Morris, Angela M Wood, Matteo Quartagno, and Michael G Kenward.Multiple imputation and its application. John Wiley & Sons, 2023

  4. [12]

    Handling missing data when estimating causal effects with targeted maximum likelihood estimation.American Journal of Epidemi- ology, 193(7):1019–1030, 2024

    S Ghazaleh Dashti, Katherine J Lee, Julie A Simpson, Ian R White, John B Car- lin, and Margarita Moreno-Betancur. Handling missing data when estimating causal effects with targeted maximum likelihood estimation.American Journal of Epidemi- ology, 193(7):1019–1030, 2024

  5. [13]

    Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values.Statistics in medicine, 29 (28):2920–2931, 2010

    Ian R White and John B Carlin. Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values.Statistics in medicine, 29 (28):2920–2931, 2010. 23

  6. [14]

    John Wiley & Sons, 2019

    Roderick JA Little and Donald B Rubin.Statistical analysis with missing data, volume 793. John Wiley & Sons, 2019

  7. [15]

    Imputation without nightmars: Graphical criteria for valid imputation of missing data

    Maya B Mathur and Ilya Shpitser. Imputation without nightmars: Graphical criteria for valid imputation of missing data. Technical report, Center for Open Science, 2024

  8. [16]

    A cautious note on auxiliary variables that can increase bias in missing data problems.Multivariate Behavioral Research, 49(5): 443–459, 2014

    Felix Thoemmes and Norman Rose. A cautious note on auxiliary variables that can increase bias in missing data problems.Multivariate Behavioral Research, 49(5): 443–459, 2014

  9. [17]

    Elinor Curnow, Kate Tilling, Jon E Heron, Rosie P Cornish, and James R Carpenter. Multiple imputation of missing data under missing at random: including a collider as an auxiliary variable in the imputation model can induce bias.Frontiers in epi- demiology, 3:1237447, 2023

  10. [18]

    Elinor Curnow, Rosie P Cornish, Jon E Heron, James R Carpenter, and Kate Tilling. Multiple imputation using auxiliary imputation variables that only predict missing- ness can increase bias due to data missing not at random.BMC Medical Research Methodology, 24(1):231, 2024

  11. [19]

    A common-cause principle for eliminating selection bias in causal estimands through covariate adjustment.The Annals of Statistics, 53(6):2303–2328, 2025

    Maya B Mathur, Tyler J VanderWeele, and Ilya Shpitser. A common-cause principle for eliminating selection bias in causal estimands through covariate adjustment.The Annals of Statistics, 53(6):2303–2328, 2025

  12. [20]

    Multiple-imputation inferences with uncongenial sources of input

    Xiao-Li Meng. Multiple-imputation inferences with uncongenial sources of input. Statistical Science, pages 538–558, 1994

  13. [21]

    Multiple imputation of covariates by fully conditional specification: accommodating the substantive model.Statisti- cal methods in medical research, 24(4):462–487, 2015

    Jonathan W Bartlett, Shaun R Seaman, Ian R White, James R Carpenter, and Alzheimer’s Disease Neuroimaging Initiative*. Multiple imputation of covariates by fully conditional specification: accommodating the substantive model.Statisti- cal methods in medical research, 24(4):462...

  14. [22]

    Cannabis use and mental health in young people: cohort study.Bmj, 325(7374):1195–1198, 2002

    George C Patton, Carolyn Coffey, John B Carlin, et al. Cannabis use and mental health in young people: cohort study.Bmj, 325(7374):1195–1198, 2002

  15. [23]

    The manual of cis-r.London: Institute of Psychiatry, 1992

    G Lewis and AJ Pelosi. The manual of cis-r.London: Institute of Psychiatry, 1992

  16. [24]

    Natalia A Goriounova and Huibert D Mansvelder. Short-and long-term conse- quences of nicotine exposure during adolescence for prefrontal cortex neuronal net- work function.Cold Spring Harbor perspectives in medicine, 2(12):a012120, 2012

  17. [25]

    Cannabis use in adolescence and young adult- hood: a review of findings from the victorian adolescent health cohort study.The Canadian Journal of Psychiatry, 61(6):318–327, 2016

    Carolyn Coffey and George C Patton. Cannabis use in adolescence and young adult- hood: a review of findings from the victorian adolescent health cohort study.The Canadian Journal of Psychiatry, 61(6):318–327, 2016

  18. [26]

    Karis Colyer-Patel, Lauren Kuhns, Alix Weidema, Heidi Lesscher, and Janna Cousijn. Age-dependent effects of tobacco smoke and nicotine on cognition and the brain: A systematic review of the human and animal literature comparing ado- lescents and adults.Neuroscience & Biobehavi...

  19. [27]

    Graphical models for inference with missing data.Advances in neural information processing systems, 26, 2013

    Karthika Mohan, Judea Pearl, and Jin Tian. Graphical models for inference with missing data.Advances in neural information processing systems, 26, 2013. 24

  20. [28]

    On the stationary distribution of iterative imputations.Biometrika, 101(1):155–173, 2014

    Jingchen Liu, Andrew Gelman, Jennifer Hill, Yu-Sung Su, and Jonathan Kropko. On the stationary distribution of iterative imputations.Biometrika, 101(1):155–173, 2014

  21. [29]

    Appropriate inclusion of interactions was needed to avoid bias in multiple imputation.Journal of clinical epidemiology, 80:107–115, 2016

    Kate Tilling, Elizabeth J Williamson, Michael Spratt, Jonathan AC Sterne, and James R Carpenter. Appropriate inclusion of interactions was needed to avoid bias in multiple imputation.Journal of clinical epidemiology, 80:107–115, 2016

  22. [30]

    Multiple imputation of discrete and continuous data by fully con- ditional specification.Statistical methods in medical research, 16(3):219–242, 2007

    Stef Van Buuren. Multiple imputation of discrete and continuous data by fully con- ditional specification.Statistical methods in medical research, 16(3):219–242, 2007

  23. [31]

    Multiple imputation using chained equations: issues and guidance for practice.Statistics in medicine, 30(4): 377–399, 2011

    Ian R White, Patrick Royston, and Angela M Wood. Multiple imputation using chained equations: issues and guidance for practice.Statistics in medicine, 30(4): 377–399, 2011

  24. [32]

    Routledge, 2017

    Leo Breiman, Jerome Friedman, Richard A Olshen, and Charles J Stone.Classifi- cation and regression trees. Routledge, 2017

  25. [33]

    Recursive partitioning for missing data imputation in the presence of interaction effects.Computational statis- tics & data analysis, 72:92–104, 2014

    Lisa L Doove, Stef Van Buuren, and Elise Dusseldorp. Recursive partitioning for missing data imputation in the presence of interaction effects.Computational statis- tics & data analysis, 72:92–104, 2014

  26. [34]

    Principles of confounder selection: Tj vanderweele.European journal of epidemiology, 34(3):211–219, 2019

    Tyler J VanderWeele. Principles of confounder selection: Tj vanderweele.European journal of epidemiology, 34(3):211–219, 2019

  27. [35]

    Bootstrap inference for multiple im- putation under uncongeniality and misspecification.Statistical methods in medical research, 29(12):3533–3546, 2020

    Jonathan W Bartlett and Rachael A Hughes. Bootstrap inference for multiple im- putation under uncongeniality and misspecification.Statistical methods in medical research, 29(12):3533–3546, 2020

  28. [36]

    R Founda- tion for Statistical Computing, Vienna, Austria, 2019

    R Core Team.R: A Language and Environment for Statistical Computing. R Founda- tion for Statistical Computing, Vienna, Austria, 2019. URL https://www.R-project. org/

  29. [37]

    mice: Multivariate imputation by chained equations in r.Journal of statistical software, 45:1–67, 2011

    Stef Van Buuren and Karin Groothuis-Oudshoorn. mice: Multivariate imputation by chained equations in r.Journal of statistical software, 45:1–67, 2011

  30. [38]

    rpart: Recursive partitioning and regression trees.R package version, 4:1–9, 2015

    Terry Therneau, Beth Atkinson, Brian Ripley, et al. rpart: Recursive partitioning and regression trees.R package version, 4:1–9, 2015

  31. [39]

    smcfcs: multiple imputation of co- variates by substantive model compatible fully conditional specification

    JW Bartlett, R Keogh, EF Bonneville, et al. smcfcs: multiple imputation of co- variates by substantive model compatible fully conditional specification. r package version 1.2. 1.URL https://CRAN. R-project. org/package= smcfcs, 2015

  32. [40]

    Using simulation studies to evaluate statistical methods.Statistics in medicine, 38(11):2074–2102, 2019

    Tim P Morris, Ian R White, and Michael J Crowther. Using simulation studies to evaluate statistical methods.Statistics in medicine, 38(11):2074–2102, 2019

  33. [41]

    Generalized adjustment under con- founding and selection biases

    Juan Correa, Jin Tian, and Elias Bareinboim. Generalized adjustment under con- founding and selection biases. InProceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018

  34. [42]

    Jochen Hardt, Max Herke, and Rainer Leonhart. Auxiliary variables in multiple imputation in regression with missing x: a warning against including too many in small sample research.BMC medical research methodology, 12(1):184, 2012. 25

  35. [43]

    A comparison of strategies for selecting auxiliary variables for multiple imputation.Biometrical Journal, 66(1):2200291, 2024

    Rheanna M Mainzer, Cattram D Nguyen, John B Carlin, Margarita Moreno- Betancur, Ian R White, and Katherine J Lee. A comparison of strategies for selecting auxiliary variables for multiple imputation.Biometrical Journal, 66(1):2200291, 2024

  36. [44]

    Pitfalls of imputing using incomplete auxiliary variables.American Journal of Epidemiology, 194(6):1801–1802, 2025

    Maya B Mathur and Ilya Shpitser. Pitfalls of imputing using incomplete auxiliary variables.American Journal of Epidemiology, 194(6):1801–1802, 2025

  37. [45]

    Re- coverability of causal effects under presence of missing data: a longitudinal case study.Biostatistics, 26(1):kxae044, 2025

    Anastasiia Holovchak, Helen McIlleron, Paolo Denti, and Michael Schomaker. Re- coverability of causal effects under presence of missing data: a longitudinal case study.Biostatistics, 26(1):kxae044, 2025

  38. [46]

    Introduction to double robust methods for incomplete data.Statistical science: a review journal of the Institute of Mathemati- cal Statistics, 33(2):184, 2018

    Shaun R Seaman and Stijn Vansteelandt. Introduction to double robust methods for incomplete data.Statistical science: a review journal of the Institute of Mathemati- cal Statistics, 33(2):184, 2018

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.