REVIEW 2 major objections 1 minor 46 references
On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data
T0 review · 2 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read Distinguishing mediator from non-mediator auxiliary variables prevents bias in multiple imputation for average causal effect estimation.
desk verdict Paper shows mediator vs non-mediator auxiliaries matter for bias in MI for ACE but its m-DAGs may not cover typical cases. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Missingness directed acyclic graphs (m-DAGs) that depict the mechanisms, used to derive recoverability results for the ACE with auxiliary variables in MI.
What would settle it
A simulation or empirical analysis in which mediator auxiliary variables are included in standard MI without compatibility checks produces persistent bias in the ACE estimate relative to a gold-standard complete-data analysis.
Extended reading notes
Core claim
For a range of missingness mechanisms, the average causal effect is recoverable using multiple imputation only when auxiliary variables are incorporated in a manner compatible with the analysis model, with mediator auxiliary variables requiring particular care to prevent bias in g-computation estimates.
Load-bearing premise
The missingness directed acyclic graphs considered in the paper represent the typical missingness mechanisms encountered when estimating average causal effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript claims that distinguishing between mediator and non-mediator auxiliary variables is crucial to avoid bias when using multiple imputation (MI) to estimate the average causal effect (ACE) with missing data, and that compatible and flexible non-parametric MI methods incorporating these variables are necessary. It derives recoverability results for a range of univariable and multivariable missingness mechanisms using m-DAGs and evaluates them through simulation studies comparing MI-based and complete-case methods with g-computation.
Significance. If the results hold, the paper provides important practical guidance on the use of auxiliary variables in MI for causal inference, filling a gap in the literature regarding auxiliaries necessary for identifiability of the ACE. Strengths include the derivation of recoverability results across multiple m-DAGs and the evaluation of different MI strategies under correctly specified g-computation. This could improve the reliability of ACE estimates in observational data with missingness.
major comments (2)
- [m-DAGs and recoverability derivations] m-DAGs and recoverability section: The recoverability results and bias-avoidance guidance are derived exclusively under the specific set of m-DAGs considered; the manuscript does not provide evidence or discussion that these m-DAGs are representative of typical missingness mechanisms in applied causal inference settings (e.g., those involving unmeasured common causes of missingness indicators or time-varying mechanisms). This is load-bearing for the generalizability of the recommendations.
- [Simulation studies] Simulation studies section: The simulations rely on the chosen m-DAGs and correctly specified models; no sensitivity analyses are presented for misspecified MI models or deviations from the assumed missingness structures, which could affect the conclusion that non-parametric MI is required to avoid bias.
minor comments (1)
- [Abstract] The abstract could more explicitly state the number of m-DAGs considered and summarize the key quantitative findings from the simulations.
Simulated Author's Rebuttal
We thank the referee for their constructive comments on our manuscript. We address the two major comments point by point below, focusing on the scope of our m-DAGs and the design of the simulation studies. We propose targeted revisions to improve clarity on limitations while preserving the core contributions on recoverability and MI compatibility.
read point-by-point responses
-
Referee: m-DAGs and recoverability section: The recoverability results and bias-avoidance guidance are derived exclusively under the specific set of m-DAGs considered; the manuscript does not provide evidence or discussion that these m-DAGs are representative of typical missingness mechanisms in applied causal inference settings (e.g., those involving unmeasured common causes of missingness indicators or time-varying mechanisms). This is load-bearing for the generalizability of the recommendations.
Authors: We selected the m-DAGs to cover a range of univariable and multivariable mechanisms commonly discussed in the missing data and causal inference literature (e.g., missingness depending on observed covariates or outcomes). The manuscript already frames these as 'typical' based on prior work. We agree that mechanisms with unmeasured common causes of missingness indicators or time-varying structures fall outside this scope and could affect generalizability. We will add a dedicated limitations paragraph in the discussion section explicitly noting the selected m-DAGs, their motivation from existing literature, and the value of future extensions to more complex mechanisms. revision: partial
-
Referee: Simulation studies section: The simulations rely on the chosen m-DAGs and correctly specified models; no sensitivity analyses are presented for misspecified MI models or deviations from the assumed missingness structures, which could affect the conclusion that non-parametric MI is required to avoid bias.
Authors: The simulations were intentionally conducted under correctly specified g-computation and the assumed m-DAGs to isolate the impact of MI compatibility and flexibility on bias when recoverability holds. This design directly supports the recoverability derivations. We acknowledge that misspecification of MI models or deviations from the m-DAGs could alter performance and that sensitivity analyses would provide additional practical insight. We will expand the discussion to address this limitation and note that the current results demonstrate bias avoidance under compatible flexible MI when the substantive model is correct. revision: partial
Circularity Check
No circularity: recoverability results and simulations are independent of fitted quantities
full rationale
The paper derives recoverability results from the structure of specified m-DAGs using standard graphical criteria for missing data, then evaluates MI and complete-case estimators via simulation studies that apply correctly specified g-computation on data generated from those same DAGs. No prediction or result reduces by construction to a parameter fitted from the target data; the simulations serve as external benchmarks rather than self-referential fits. No load-bearing self-citation chains or ansatz smuggling are present. The derivation chain remains self-contained against the paper's stated mechanisms.
Assumptions & free parameters
assumptions (1)
- domain assumption The m-DAGs considered correctly represent typical univariable and multivariable missingness mechanisms for ACE estimation.
Cite this review
Pith. "Pith review of On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data." pith.science (2026). https://pith.science/paper/QEHBBXNR
@misc{pith2026260622016,
author = {Pith},
title = {Pith review of: On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data},
year = {2026},
howpublished = {\url{https://pith.science/paper/QEHBBXNR}},
note = {Machine review of arXiv:2606.22016}
}
read the original abstract
Estimating the average causal effect (ACE) using observational data is a key focus in causal inference for which missing data present an important challenge. Multiple imputation (MI) is a widely used method for handling missing data and can yield unbiased estimates when the imputation is compatible with the substantive analysis. One of the advantages of MI is its scope to include so-called "auxiliary variables", defined as variables associated with incomplete variables that are excluded from the substantive analysis. Although many studies have looked at the use of auxiliary variables in MI for improving precision, the study of auxiliary variables that are necessary for the identifiability (or "recoverability") of the ACE in the presence of missing data has been scant. In this work, we investigate the use of auxiliary variables, both mediators and non-mediators, across a range of typical univariable and multivariable missingness mechanisms depicted by missingness directed acyclic graphs (m-DAGs). For each setting, we derive recoverability results, then evaluate MI-based and complete-case methods for estimating the ACE using correctly specified g-computation, considering different strategies for incorporating auxiliary variables and varying degrees of compatibility for MI models. Based on findings from the simulation studies, we provide practical guidance, highlighting that distinguishing appropriately between mediator and non-mediator auxiliary variables is important to avoid bias as is the use of compatible and flexible (non-parametric) MI methods that incorporate these variables.
Reference graph
Works this paper leans on
-
[1]
Chapman & Hall/CRC, Boca Ratonn, 2020
Robins JM Hernán MA.Causal Inference: What If. Chapman & Hall/CRC, Boca Ratonn, 2020
2020
-
[2]
A new approach to causal inference in mortality studies with a sus- tained exposure period—application to control of the healthy worker survivor effect
James Robins. A new approach to causal inference in mortality studies with a sus- tained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling, 7(9-12):1393–1512, 1986
1986
-
[3]
Implementation of g- computation on a simulated data set: demonstration of a causal inference technique
Jonathan M Snowden, Sherri Rose, and Kathleen M Mortimer. Implementation of g- computation on a simulated data set: demonstration of a causal inference technique. American journal of epidemiology, 173(7):731–738, 2011
2011
-
[4]
Graphical models for recovering probabilistic and causal queries from missing data.Advances in Neural Information Processing Systems, 27, 2014
Karthika Mohan and Judea Pearl. Graphical models for recovering probabilistic and causal queries from missing data.Advances in Neural Information Processing Systems, 27, 2014
2014
-
[5]
Canonical causal diagrams to guide the treatment of missing data in epidemiologic studies.American journal of epidemiology, 187(12): 2705–2715, 2018
Margarita Moreno-Betancur, Katherine J Lee, Finbarr P Leacy, Ian R White, Julie A Simpson, and John B Carlin. Canonical causal diagrams to guide the treatment of missing data in epidemiologic studies.American journal of epidemiology, 187(12): 2705–2715, 2018
2018
-
[6]
Recoverability and estimation of causal effects under typical multivariable missingness mechanisms.Biometrical Journal, 66(3):2200326, 2024
Jiaxin Zhang, S Ghazaleh Dashti, John B Carlin, Katherine J Lee, and Margarita Moreno-Betancur. Recoverability and estimation of causal effects under typical multivariable missingness mechanisms.Biometrical Journal, 66(3):2200326, 2024
2024
-
[7]
canonical causal diagrams to guide the treatment of missing data in epidemiologic studies
Margarita Moreno-Betancur, Katherine J Lee, Finbarr P Leacy, Julie A Simpson, and John B Carlin. Correction to:“canonical causal diagrams to guide the treatment of missing data in epidemiologic studies”.American Journal of Epidemiology, page kwae406, 2025
2025
-
[8]
Mediation analysis with the mediator and outcome missing not at random.Journal of the American Statistical Association, 120(550):794–804, 2025
Shuozhi Zuo, Debashis Ghosh, Peng Ding, and Fan Yang. Mediation analysis with the mediator and outcome missing not at random.Journal of the American Statistical Association, 120(550):794–804, 2025
2025
Show all 46 references
-
[9]
Katherine J Lee, John B Carlin, Julie A Simpson, and Margarita Moreno-Betancur. Assumptions and analysis planning in studies with missing data in multiple vari- ables: moving beyond the mcar/mar/mnar classification.International Journal of Epidemiology, 52(4):1268–1275, 2023
2023
-
[10]
John Wiley & Sons, 2004
Donald B Rubin.Multiple imputation for nonresponse in surveys, volume 81. John Wiley & Sons, 2004
2004
-
[11]
John Wiley & Sons, 2023
James R Carpenter, Jonathan W Bartlett, Tim P Morris, Angela M Wood, Matteo Quartagno, and Michael G Kenward.Multiple imputation and its application. John Wiley & Sons, 2023
2023
-
[12]
Handling missing data when estimating causal effects with targeted maximum likelihood estimation.American Journal of Epidemi- ology, 193(7):1019–1030, 2024
S Ghazaleh Dashti, Katherine J Lee, Julie A Simpson, Ian R White, John B Car- lin, and Margarita Moreno-Betancur. Handling missing data when estimating causal effects with targeted maximum likelihood estimation.American Journal of Epidemi- ology, 193(7):1019–1030, 2024
2024
-
[13]
Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values.Statistics in medicine, 29 (28):2920–2931, 2010
Ian R White and John B Carlin. Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values.Statistics in medicine, 29 (28):2920–2931, 2010. 23
2010
-
[14]
John Wiley & Sons, 2019
Roderick JA Little and Donald B Rubin.Statistical analysis with missing data, volume 793. John Wiley & Sons, 2019
2019
-
[15]
Imputation without nightmars: Graphical criteria for valid imputation of missing data
Maya B Mathur and Ilya Shpitser. Imputation without nightmars: Graphical criteria for valid imputation of missing data. Technical report, Center for Open Science, 2024
2024
-
[16]
A cautious note on auxiliary variables that can increase bias in missing data problems.Multivariate Behavioral Research, 49(5): 443–459, 2014
Felix Thoemmes and Norman Rose. A cautious note on auxiliary variables that can increase bias in missing data problems.Multivariate Behavioral Research, 49(5): 443–459, 2014
2014
-
[17]
Elinor Curnow, Kate Tilling, Jon E Heron, Rosie P Cornish, and James R Carpenter. Multiple imputation of missing data under missing at random: including a collider as an auxiliary variable in the imputation model can induce bias.Frontiers in epi- demiology, 3:1237447, 2023
2023
-
[18]
Elinor Curnow, Rosie P Cornish, Jon E Heron, James R Carpenter, and Kate Tilling. Multiple imputation using auxiliary imputation variables that only predict missing- ness can increase bias due to data missing not at random.BMC Medical Research Methodology, 24(1):231, 2024
2024
-
[19]
A common-cause principle for eliminating selection bias in causal estimands through covariate adjustment.The Annals of Statistics, 53(6):2303–2328, 2025
Maya B Mathur, Tyler J VanderWeele, and Ilya Shpitser. A common-cause principle for eliminating selection bias in causal estimands through covariate adjustment.The Annals of Statistics, 53(6):2303–2328, 2025
2025
-
[20]
Multiple-imputation inferences with uncongenial sources of input
Xiao-Li Meng. Multiple-imputation inferences with uncongenial sources of input. Statistical Science, pages 538–558, 1994
1994
-
[21]
Multiple imputation of covariates by fully conditional specification: accommodating the substantive model.Statisti- cal methods in medical research, 24(4):462–487, 2015
Jonathan W Bartlett, Shaun R Seaman, Ian R White, James R Carpenter, and Alzheimer’s Disease Neuroimaging Initiative*. Multiple imputation of covariates by fully conditional specification: accommodating the substantive model.Statisti- cal methods in medical research, 24(4):462...
2015
-
[22]
Cannabis use and mental health in young people: cohort study.Bmj, 325(7374):1195–1198, 2002
George C Patton, Carolyn Coffey, John B Carlin, et al. Cannabis use and mental health in young people: cohort study.Bmj, 325(7374):1195–1198, 2002
2002
-
[23]
The manual of cis-r.London: Institute of Psychiatry, 1992
G Lewis and AJ Pelosi. The manual of cis-r.London: Institute of Psychiatry, 1992
1992
-
[24]
Natalia A Goriounova and Huibert D Mansvelder. Short-and long-term conse- quences of nicotine exposure during adolescence for prefrontal cortex neuronal net- work function.Cold Spring Harbor perspectives in medicine, 2(12):a012120, 2012
2012
-
[25]
Cannabis use in adolescence and young adult- hood: a review of findings from the victorian adolescent health cohort study.The Canadian Journal of Psychiatry, 61(6):318–327, 2016
Carolyn Coffey and George C Patton. Cannabis use in adolescence and young adult- hood: a review of findings from the victorian adolescent health cohort study.The Canadian Journal of Psychiatry, 61(6):318–327, 2016
2016
-
[26]
Karis Colyer-Patel, Lauren Kuhns, Alix Weidema, Heidi Lesscher, and Janna Cousijn. Age-dependent effects of tobacco smoke and nicotine on cognition and the brain: A systematic review of the human and animal literature comparing ado- lescents and adults.Neuroscience & Biobehavi...
2023
-
[27]
Graphical models for inference with missing data.Advances in neural information processing systems, 26, 2013
Karthika Mohan, Judea Pearl, and Jin Tian. Graphical models for inference with missing data.Advances in neural information processing systems, 26, 2013. 24
2013
-
[28]
On the stationary distribution of iterative imputations.Biometrika, 101(1):155–173, 2014
Jingchen Liu, Andrew Gelman, Jennifer Hill, Yu-Sung Su, and Jonathan Kropko. On the stationary distribution of iterative imputations.Biometrika, 101(1):155–173, 2014
2014
-
[29]
Appropriate inclusion of interactions was needed to avoid bias in multiple imputation.Journal of clinical epidemiology, 80:107–115, 2016
Kate Tilling, Elizabeth J Williamson, Michael Spratt, Jonathan AC Sterne, and James R Carpenter. Appropriate inclusion of interactions was needed to avoid bias in multiple imputation.Journal of clinical epidemiology, 80:107–115, 2016
2016
-
[30]
Multiple imputation of discrete and continuous data by fully con- ditional specification.Statistical methods in medical research, 16(3):219–242, 2007
Stef Van Buuren. Multiple imputation of discrete and continuous data by fully con- ditional specification.Statistical methods in medical research, 16(3):219–242, 2007
2007
-
[31]
Multiple imputation using chained equations: issues and guidance for practice.Statistics in medicine, 30(4): 377–399, 2011
Ian R White, Patrick Royston, and Angela M Wood. Multiple imputation using chained equations: issues and guidance for practice.Statistics in medicine, 30(4): 377–399, 2011
2011
-
[32]
Routledge, 2017
Leo Breiman, Jerome Friedman, Richard A Olshen, and Charles J Stone.Classifi- cation and regression trees. Routledge, 2017
2017
-
[33]
Recursive partitioning for missing data imputation in the presence of interaction effects.Computational statis- tics & data analysis, 72:92–104, 2014
Lisa L Doove, Stef Van Buuren, and Elise Dusseldorp. Recursive partitioning for missing data imputation in the presence of interaction effects.Computational statis- tics & data analysis, 72:92–104, 2014
2014
-
[34]
Principles of confounder selection: Tj vanderweele.European journal of epidemiology, 34(3):211–219, 2019
Tyler J VanderWeele. Principles of confounder selection: Tj vanderweele.European journal of epidemiology, 34(3):211–219, 2019
2019
-
[35]
Bootstrap inference for multiple im- putation under uncongeniality and misspecification.Statistical methods in medical research, 29(12):3533–3546, 2020
Jonathan W Bartlett and Rachael A Hughes. Bootstrap inference for multiple im- putation under uncongeniality and misspecification.Statistical methods in medical research, 29(12):3533–3546, 2020
2020
-
[36]
R Founda- tion for Statistical Computing, Vienna, Austria, 2019
R Core Team.R: A Language and Environment for Statistical Computing. R Founda- tion for Statistical Computing, Vienna, Austria, 2019. URL https://www.R-project. org/
2019
-
[37]
mice: Multivariate imputation by chained equations in r.Journal of statistical software, 45:1–67, 2011
Stef Van Buuren and Karin Groothuis-Oudshoorn. mice: Multivariate imputation by chained equations in r.Journal of statistical software, 45:1–67, 2011
2011
-
[38]
rpart: Recursive partitioning and regression trees.R package version, 4:1–9, 2015
Terry Therneau, Beth Atkinson, Brian Ripley, et al. rpart: Recursive partitioning and regression trees.R package version, 4:1–9, 2015
2015
-
[39]
smcfcs: multiple imputation of co- variates by substantive model compatible fully conditional specification
JW Bartlett, R Keogh, EF Bonneville, et al. smcfcs: multiple imputation of co- variates by substantive model compatible fully conditional specification. r package version 1.2. 1.URL https://CRAN. R-project. org/package= smcfcs, 2015
2015
-
[40]
Using simulation studies to evaluate statistical methods.Statistics in medicine, 38(11):2074–2102, 2019
Tim P Morris, Ian R White, and Michael J Crowther. Using simulation studies to evaluate statistical methods.Statistics in medicine, 38(11):2074–2102, 2019
-
[41]
Generalized adjustment under con- founding and selection biases
Juan Correa, Jin Tian, and Elias Bareinboim. Generalized adjustment under con- founding and selection biases. InProceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018
2018
-
[42]
Jochen Hardt, Max Herke, and Rainer Leonhart. Auxiliary variables in multiple imputation in regression with missing x: a warning against including too many in small sample research.BMC medical research methodology, 12(1):184, 2012. 25
2012
-
[43]
A comparison of strategies for selecting auxiliary variables for multiple imputation.Biometrical Journal, 66(1):2200291, 2024
Rheanna M Mainzer, Cattram D Nguyen, John B Carlin, Margarita Moreno- Betancur, Ian R White, and Katherine J Lee. A comparison of strategies for selecting auxiliary variables for multiple imputation.Biometrical Journal, 66(1):2200291, 2024
2024
-
[44]
Pitfalls of imputing using incomplete auxiliary variables.American Journal of Epidemiology, 194(6):1801–1802, 2025
Maya B Mathur and Ilya Shpitser. Pitfalls of imputing using incomplete auxiliary variables.American Journal of Epidemiology, 194(6):1801–1802, 2025
2025
-
[45]
Re- coverability of causal effects under presence of missing data: a longitudinal case study.Biostatistics, 26(1):kxae044, 2025
Anastasiia Holovchak, Helen McIlleron, Paolo Denti, and Michael Schomaker. Re- coverability of causal effects under presence of missing data: a longitudinal case study.Biostatistics, 26(1):kxae044, 2025
2025
-
[46]
Introduction to double robust methods for incomplete data.Statistical science: a review journal of the Institute of Mathemati- cal Statistics, 33(2):184, 2018
Shaun R Seaman and Stijn Vansteelandt. Introduction to double robust methods for incomplete data.Statistical science: a review journal of the Institute of Mathemati- cal Statistics, 33(2):184, 2018
2018
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.