REVIEW 3 major objections 4 minor 23 references
This simulation study finds that multiple imputation by chained equations is the most reliable method for estimating average treatment effects when treatment or outcome data are missing, and the only one of the four compared methods that ke
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:02 UTC pith:X3J6G3SL
load-bearing objection A competent, confirmatory simulation study of missing-data methods for ATE estimation under time-varying confounding; the MICE conclusion is plausible but the missing target-ATE definition and method-dependent exclusion of extreme replications keep it from being decisive. the 3 major comments →
Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In a full-factorial simulation with 48 scenarios and 500 replications per scenario, binary treatment and survival outcomes over three time points were generated under time-varying confounding, missingness was introduced under MCAR, MAR (in treatment or outcome), and MNAR, and the average treatment effect was estimated by logistic regression with inverse-probability treatment weighting. The paper's central finding is that MICE outperforms complete-case analysis, single mode imputation, and stratified hot deck imputation on absolute bias, 95% confidence interval coverage, and empirical standard error under most conditions, and is the only method to keep coverage at acceptable levels under MNAR
What carries the argument
The central object is a full-factorial simulation design that crosses missingness mechanism (MCAR, MAR in treatment, MAR in outcome, MNAR), missing rate (20% and 50%), missingness location (treatment vs. outcome), confounding strength (rho = -0.1, -0.4, -0.7), and sample size (1000 and 5000), benchmarked against a full-data model. The data-generating process is a longitudinal model with three treatment time points, binary survival outcomes, and time-varying confounding; ATE is estimated by logistic regression with inverse-probability treatment weighting. MICE (iterative chained regression imputation with m = 5 imputations) is the key method under study, and performance is measured by absolut
Load-bearing premise
The entire comparison treats the full-data logistic IPTW model as the unbiased reference; if that model is misspecified under time-varying confounding with sparse binary outcomes, then 'bias' and 'coverage' measure deviation from a model rather than from the true treatment effect, and the MICE-versus-alternatives ranking could change.
What would settle it
In a replication of the simulation, compute the true ATE directly from the generated potential outcomes and compare the full-data IPTW estimate against it; if the full-data estimator is biased in any scenario, the reported absolute bias and coverage are relative to a flawed benchmark. A second check would run MICE with more than 5 imputations under MNAR to see whether its coverage remains acceptable only because of variance inflation.
If this is right
- If the paper's finding is correct, applied researchers estimating average treatment effects with missing treatment or outcome data should prefer MICE over complete-case analysis, single mode imputation, and stratified hot deck imputation.
- MICE can serve as a useful robustness check even when missingness may be not at random, though its acceptable coverage under MNAR likely comes from inflated standard errors rather than reduced bias.
- Routinely checking propensity score distributions is important, because extreme weights are common under MNAR, high missingness, and small samples, and they signal potential positivity violations.
- Strong confounding degrades coverage most for the simpler missing-data methods, so the choice of imputation method interacts with the difficulty of confounding adjustment.
- Missingness in variables with highly skewed distributions, such as the survival outcomes used here, degrades performance even for sophisticated imputation, so variable-specific distributions should be considered when planning analyses.
Where Pith is reading between the lines
- The paper treats the full-data logistic IPTW model as the reference for bias and coverage; since the simulation knows the true potential outcomes, a direct check of whether the full-data estimator is unbiased would confirm whether the reported ranking is against the truth or against a possibly misspecified model.
- MICE's robustness under MNAR may trade bias for uncertainty; in settings where tight intervals matter, the wider confidence intervals produced by MICE could be a real cost rather than a free gain.
- The finding that missingness location matters more when outcomes are skewed suggests that imputation models for rare binary events could benefit from outcome-specific adjustments, such as weighting or transformations, a testable extension the paper leaves implicit.
- The simulation design could be extended to doubly robust estimators and machine-learning-based imputation methods, which the paper itself suggests as future work; such extensions would clarify whether MICE's advantage persists when the outcome or treatment model is less parametric.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a simulation study comparing four missing-data methods—complete-case analysis, single mode imputation, stratified hot deck imputation, and multiple imputation by chained equations (MICE)—for estimating an average treatment effect under time-varying confounding. The data-generating process is a longitudinal binary-treatment/binary-survival setting with four confounders. Missingness is introduced under MCAR, MAR (separately in treatment and outcome variables), and MNAR; missing rate, confounding strength, and sample size are also varied in a full-factorial design (48 scenarios, 500 replications each). ATE estimates are obtained from logistic regression with inverse-probability treatment weighting, and methods are compared using absolute bias, 95% confidence interval coverage, empirical standard error, and the proportion of extreme propensity scores. The paper's central claim is that MICE outperforms the alternatives under most conditions and is the only method to maintain acceptable coverage under MNAR.
Significance. If the results are valid, the study addresses a practically important gap: it jointly varies missingness mechanism, missingness location, missing rate, confounding strength, and sample size, and it compares commonly used imputation approaches under time-varying confounding. The full-factorial design, use of Monte Carlo standard errors, and effect-size-thresholded ANOVA are strengths, as is the explicit attention to positivity violations. However, the credibility of the headline ranking depends on two load-bearing aspects that are not adequately established: the validity of the full-data benchmark as a proxy for the true ATE, and the treatment of method-dependent exclusion of extreme replications. Because both can change the relative ranking of methods, the paper's conclusions are not yet supported as reported.
major comments (3)
- [Sections 2.1 and 2.7] The target estimand is never defined. Absolute bias and coverage are computed relative to the full-data model, i.e., a logistic IPTW model fitted on the complete synthetic data, but the paper does not demonstrate that this model is unbiased for the ATE under the stated DGP. The DGP includes time-varying treatments, survival outcomes, structural missingness, and very sparse binary outcomes; a logistic IPTW model with HC3 standard errors can be misspecified in such settings. If the full-data model is biased, 'absolute bias' and 'coverage' measure deviation from a possibly incorrect model rather than from the truth, and the MICE-versus-alternatives ranking becomes uninterpretable. Please define the true ATE (e.g., via the g-formula or Monte Carlo standardization using the known data-generating probabilities), report the full-data model's bias and coverage against that target, and, if needed
- [Section 2.7 and Section 3.1] The exclusion of replications with |ATE|>100 is method-dependent and can induce a selection effect that directly undermines the headline comparison. The paper reports that MICE exclusion rates are negligible while CCA, SMI, and HD have higher exclusion rates, especially under MNAR, 50% missingness, and n=1000. Because bias, coverage, and empirical SE are computed only on the remaining replications, methods with higher exclusion rates are evaluated on a selected subset that may have systematically different properties. This is a collider/selection effect: it can differentially remove the worst-performing replications for the comparator methods and inflate MICE's apparent coverage advantage. The manuscript reports exclusion rates but does not incorporate them into the performance measures or the interpretation. Please add sensitivity analyses that retain all replications (e.g., through win
- [Section 5, Table 4, Table 5] The conclusion that MICE was 'the only method to maintain acceptable coverage under MNAR' is not directly supported by the reported results. Table 4 shows no statistically significant coverage difference between MCAR and MNAR pooled over methods, and the paper does not present a method-by-mechanism coverage table or figure. Moreover, Table 5 reports MICE coverage as low as 0.094 in a MAR(Y), rho=-0.7, n=5000 scenario, so 'acceptable coverage' needs an explicit definition and the MNAR-specific evidence should be presented. Please provide a table or figure of coverage by method x mechanism x missing rate and define the acceptable-coverage threshold. Without this, the core conclusion is not verifiable from the manuscript.
minor comments (4)
- [Data Availability Statement] The data availability statement contains a placeholder '[ ]'. If the simulation code is genuinely available, the URL or repository identifier should be included; if not, the reproducibility claim should be softened.
- [Page 15, table header] The table header 'Txables' appears to be a typo for 'Tables'.
- [Section 2.3] The description of MCAR generation as 'simply selects random values of the variable which are then set to a null value' is ambiguous for binary variables. Please clarify whether each value is independently set to missing with probability r, or whether a random subset of rows is selected.
- [Section 3.4] The discussion of higher coverage at n=1000 being due to inflated robust standard errors is interesting but is not accompanied by the supporting model-based standard error results. A supplementary figure or table would help readers evaluate this claim.
Circularity Check
No circular derivation chain: simulation compares methods against a synthetic full-data benchmark; no fitted target or self-citation chain is used as evidence for the conclusions.
full rationale
The paper is a simulation study, not a derivation of a target quantity from assumptions that already contain the result. The data-generating process is taken from the authors' prior medRxiv preprint (ref 12), but this is an input assumption that defines the synthetic populations; it is not a conclusion that the paper then feeds back into its own evaluation. Performance measures (absolute bias, coverage, empirical SE) are computed against the full-data benchmark model described in Section 2.7, and no parameter of the missing-data methods is fitted to the outcome measure in a way that would make a 'prediction' tautological. The comparison among CCA, SMI, hot deck, and MICE is an empirical ranking of methods under specified scenarios, not a quantity that is equivalent to its inputs by construction. The manuscript's own limitations (Section 4.3) acknowledge modelling choices and potential numerical instability, but none of these limitations amounts to a circular step. The only author-overlap dependency is the DGP citation, which is standard practice for simulation studies and does not make the conclusions load-bearing on an unverified self-citation chain. The skeptic's concern about differential exclusion of extreme replications (Section 2.7 and Table 2) is a potential validity threat but not a circularity: it concerns conditioning on a collider in simulation summaries, not a reduction of the claimed result to its own inputs. Therefore no circular step can be quoted, and the correct score is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- MICE imputation count m =
5
- Extreme ATE exclusion threshold =
|ATE|>100
- Extreme propensity score cutoff =
PS<0.05 or >0.95
- Scenario factor levels =
rho in {-0.1,-0.4,-0.7}; n in {1000,5000}; missing rate in {20%,50%}; mechanism in {MCAR,MAR(A),MAR(Y),MNAR}
axioms (5)
- domain assumption The data-generating process of Velasco-Pardo et al. (ref 12) produces realistic time-varying confounding and valid potential outcomes.
- domain assumption Under full data, the logistic IPTW model is correctly specified and recovers the true ATE.
- domain assumption Generated missingness mechanisms (single cutoff for MAR and MNAR) are valid and representative of real missingness.
- domain assumption No unmeasured confounding under full data.
- standard math Rubin's rules and HC3 robust standard errors are valid for pooling and inference.
read the original abstract
Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability conditions. We generated synthetic data and conducted a simulation study comparing missing data methods across scenarios varying these factors. Missingness was introduced in both treatment and outcome variables, and we applied stratified hot deck imputation, single mode imputation, multiple imputation with chained equations (MICE), and complete-case analysis. Average treatment effect (ATE) estimates were obtained using logistic regression with propensity score weighting, and we measured coverage, absolute bias and empirical standard errors across 48 scenarios with 500 replications each. Performance was primarily driven by the missingness mechanism and choice of method, with multiple imputation generally achieving better coverage and lower bias than other methods. Missingness location was also important, while missing rate and sample size primarily affected positivity violations, which were most pronounced under MNAR, high missingness and low sample sizes. Exchangeability violations from mild to moderate confounding were adequately controlled for by propensity score models, whereas strong confounding produced a modest decrease in coverage. Further research should examine additional ways identifiability conditions can be violated under missingness, using more advanced methods and more complex missingness scenarios.
Reference graph
Works this paper leans on
-
[1]
Causal Inference
Hernán MA, Robins JM. Causal Inference. CRC Press; 2021
2021
-
[2]
Statistical Analysis with Missing Data
Little RJA, Rubin DB. Statistical Analysis with Missing Data. 2nd ed. John Wiley & Sons; 2002
2002
-
[3]
Much ado about nothing: a comparison of missing data methods and software to fit incomplete data regression models
Horton NJ, Kleinman KP. Much ado about nothing: a comparison of missing data methods and software to fit incomplete data regression models. Am Stat. 2007;61(1):79-90
2007
-
[4]
The central role of the propensity score in observational studies for causal effects
Rosenbaum PR, Rubin DB. The central role of the propensity score in observational studies for causal effects. Biometrika. 1983;70(1):41-55
1983
-
[5]
Comparison of several imputation methods for missing baseline data in propensity scores analysis of binary outcome
Crowe BJ, Lipkovich IA, Wang O. Comparison of several imputation methods for missing baseline data in propensity scores analysis of binary outcome. Pharm Stat. 2010;9(4):269- 279
2010
-
[6]
A simulation study of missing data with multiple missing X's
Rubright JD, Nandakumar R, Glutting JJ. A simulation study of missing data with multiple missing X's. Pract Assess Res Eval. 2014;19(10):1-19
2014
-
[7]
A simulation-based bias analysis to assess the impact of unmeasured confounding when designing non-randomized database studies
Desai RJ, Bradley MC, Lee H, et al. A simulation-based bias analysis to assess the impact of unmeasured confounding when designing non-randomized database studies. Am J Epidemiol. 2024
2024
-
[8]
A comparison of different methods to handle missing data in the context of propensity score analysis
Choi J, Dekkers OM, le Cessie S. A comparison of different methods to handle missing data in the context of propensity score analysis. Eur J Epidemiol. 2019;34(1):23-36
2019
-
[9]
Causal inference with confounders missing not at random
Yang S, Wang L, Ding P. Causal inference with confounders missing not at random. Biometrika. 2019;106(4):875–888
2019
-
[10]
mice: Multivariate imputation by chained equations in R
Van Buuren S, Groothuis-Oudshoorn K. mice: Multivariate imputation by chained equations in R. J Stat Softw. 2011;45(3):1-67
2011
-
[11]
Using simulation studies to evaluate statistical methods
Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Stat Med. 2019;38(11):2074-2102
2019
-
[12]
Velasco-Pardo V, Swallow B, Robertson C, McCowan C. Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy. medRxiv 2026.07.17.26358308; doi: https://doi.org/10.64898/2026.07.17.26358308
-
[13]
Prevalence and demographic variation of cardiovascular, renal, metabolic, and mental health conditions in 12 million English primary care records
Cooper JN, Nirantharakumar K, Crowe FL, et al. Prevalence and demographic variation of cardiovascular, renal, metabolic, and mental health conditions in 12 million English primary care records. BMC Med Inform Decis Mak. 2023;23(1)
2023
-
[14]
How to generate missing data for simulation studies
Zhang X. How to generate missing data for simulation studies. Quant Methods Psychol. 2023;19(2):100-122
2023
-
[15]
Multiple imputation: a primer
Schafer JL. Multiple imputation: a primer. Stat Methods Med Res. 1999;8(1):3-15. Missing Data Methods in Causal Inference 14
1999
-
[16]
A review of hot deck imputation for survey non-response
Andridge RR, Little RJA. A review of hot deck imputation for survey non-response. Int Stat Rev. 2010;78(1):40-64
2010
-
[17]
The prevention and handling of the missing data
Kang H. The prevention and handling of the missing data. Korean J Anesthesiol. 2013;64(5):402-406
2013
-
[18]
Multiple Imputation for Nonresponse in Surveys
Rubin DB. Multiple Imputation for Nonresponse in Surveys. John Wiley & Sons; 1987
1987
-
[19]
Multiple imputation using chained equations: issues and guidance for practice
White IR, Royston P, Wood AM. Multiple imputation using chained equations: issues and guidance for practice. Stat Med. 2011;30(4):377-399
2011
-
[20]
rsimsum: summarise results from Monte Carlo simulation studies
Gasparini A. rsimsum: summarise results from Monte Carlo simulation studies. J Open Source Softw. 2018;3(26):739
2018
-
[21]
Imputation of data missing not at random: artificial generation and benchmark analysis
Pereira RC, Abreu PH, Rodrigues PP, Figueiredo MAT. Imputation of data missing not at random: artificial generation and benchmark analysis. Expert Syst Appl. 2024;249:123654
2024
-
[22]
A simulation study on missing data imputation for dichotomous variables using statistical and machine learning methods
Ge Y, Li Z, Zhang J. A simulation study on missing data imputation for dichotomous variables using statistical and machine learning methods. Sci Rep. 2023;13(1)
2023
-
[23]
Miscellanea
Barnard J, Rubin DB. Miscellanea. Small-sample degrees of freedom with multiple imputation. Biometrika. 1999;86(4):948-955. Missing Data Methods in Causal Inference 15 Txables TABLE 1 Marginal distributions of treatment (A), outcome (Y) and confounder (Z) variables (n=1000, ρ=-0.4) Value A1 A2 Y1 Y2 Y3 0 80.4% 78.9% 2.6% 5.1% 7.6% 1 19.6% 21.1% 97.4% 94.9...
1999
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.