Pith. sign in

REVIEW 3 major objections 4 minor 23 references

This simulation study finds that multiple imputation by chained equations is the most reliable method for estimating average treatment effects when treatment or outcome data are missing, and the only one of the four compared methods that ke

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:02 UTC pith:X3J6G3SL

load-bearing objection A competent, confirmatory simulation study of missing-data methods for ATE estimation under time-varying confounding; the MICE conclusion is plausible but the missing target-ATE definition and method-dependent exclusion of extreme replications keep it from being decisive. the 3 major comments →

arxiv 2607.17775 v1 pith:X3J6G3SL submitted 2026-07-20 stat.ME stat.AP

Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study

classification stat.ME stat.AP
keywords average treatment effectmissing datamultiple imputationMICEpropensity score weightingtime-varying confoundingsimulation studymissing not at random
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks which missing-data method to trust when estimating average treatment effects from longitudinal data with time-varying confounding and missing treatment or outcome values. Through 48 simulated scenarios, it finds that missingness mechanism and method choice dominate performance, with multiple imputation by chained equations (MICE) consistently delivering lower bias and better coverage than complete-case analysis, single mode imputation, or stratified hot deck imputation. MICE is the only method that maintains acceptable coverage under missing-not-at-random, although it does so partly by inflating standard errors rather than reducing bias. The paper also shows that extreme propensity scores, an indicator of positivity violations, concentrate under MNAR, high missing rates, and small samples, while strong confounding erodes coverage especially for the simpler missing-data methods.

Core claim

In a full-factorial simulation with 48 scenarios and 500 replications per scenario, binary treatment and survival outcomes over three time points were generated under time-varying confounding, missingness was introduced under MCAR, MAR (in treatment or outcome), and MNAR, and the average treatment effect was estimated by logistic regression with inverse-probability treatment weighting. The paper's central finding is that MICE outperforms complete-case analysis, single mode imputation, and stratified hot deck imputation on absolute bias, 95% confidence interval coverage, and empirical standard error under most conditions, and is the only method to keep coverage at acceptable levels under MNAR

What carries the argument

The central object is a full-factorial simulation design that crosses missingness mechanism (MCAR, MAR in treatment, MAR in outcome, MNAR), missing rate (20% and 50%), missingness location (treatment vs. outcome), confounding strength (rho = -0.1, -0.4, -0.7), and sample size (1000 and 5000), benchmarked against a full-data model. The data-generating process is a longitudinal model with three treatment time points, binary survival outcomes, and time-varying confounding; ATE is estimated by logistic regression with inverse-probability treatment weighting. MICE (iterative chained regression imputation with m = 5 imputations) is the key method under study, and performance is measured by absolut

Load-bearing premise

The entire comparison treats the full-data logistic IPTW model as the unbiased reference; if that model is misspecified under time-varying confounding with sparse binary outcomes, then 'bias' and 'coverage' measure deviation from a model rather than from the true treatment effect, and the MICE-versus-alternatives ranking could change.

What would settle it

In a replication of the simulation, compute the true ATE directly from the generated potential outcomes and compare the full-data IPTW estimate against it; if the full-data estimator is biased in any scenario, the reported absolute bias and coverage are relative to a flawed benchmark. A second check would run MICE with more than 5 imputations under MNAR to see whether its coverage remains acceptable only because of variance inflation.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the paper's finding is correct, applied researchers estimating average treatment effects with missing treatment or outcome data should prefer MICE over complete-case analysis, single mode imputation, and stratified hot deck imputation.
  • MICE can serve as a useful robustness check even when missingness may be not at random, though its acceptable coverage under MNAR likely comes from inflated standard errors rather than reduced bias.
  • Routinely checking propensity score distributions is important, because extreme weights are common under MNAR, high missingness, and small samples, and they signal potential positivity violations.
  • Strong confounding degrades coverage most for the simpler missing-data methods, so the choice of imputation method interacts with the difficulty of confounding adjustment.
  • Missingness in variables with highly skewed distributions, such as the survival outcomes used here, degrades performance even for sophisticated imputation, so variable-specific distributions should be considered when planning analyses.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper treats the full-data logistic IPTW model as the reference for bias and coverage; since the simulation knows the true potential outcomes, a direct check of whether the full-data estimator is unbiased would confirm whether the reported ranking is against the truth or against a possibly misspecified model.
  • MICE's robustness under MNAR may trade bias for uncertainty; in settings where tight intervals matter, the wider confidence intervals produced by MICE could be a real cost rather than a free gain.
  • The finding that missingness location matters more when outcomes are skewed suggests that imputation models for rare binary events could benefit from outcome-specific adjustments, such as weighting or transformations, a testable extension the paper leaves implicit.
  • The simulation design could be extended to doubly robust estimators and machine-learning-based imputation methods, which the paper itself suggests as future work; such extensions would clarify whether MICE's advantage persists when the outcome or treatment model is less parametric.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript reports a simulation study comparing four missing-data methods—complete-case analysis, single mode imputation, stratified hot deck imputation, and multiple imputation by chained equations (MICE)—for estimating an average treatment effect under time-varying confounding. The data-generating process is a longitudinal binary-treatment/binary-survival setting with four confounders. Missingness is introduced under MCAR, MAR (separately in treatment and outcome variables), and MNAR; missing rate, confounding strength, and sample size are also varied in a full-factorial design (48 scenarios, 500 replications each). ATE estimates are obtained from logistic regression with inverse-probability treatment weighting, and methods are compared using absolute bias, 95% confidence interval coverage, empirical standard error, and the proportion of extreme propensity scores. The paper's central claim is that MICE outperforms the alternatives under most conditions and is the only method to maintain acceptable coverage under MNAR.

Significance. If the results are valid, the study addresses a practically important gap: it jointly varies missingness mechanism, missingness location, missing rate, confounding strength, and sample size, and it compares commonly used imputation approaches under time-varying confounding. The full-factorial design, use of Monte Carlo standard errors, and effect-size-thresholded ANOVA are strengths, as is the explicit attention to positivity violations. However, the credibility of the headline ranking depends on two load-bearing aspects that are not adequately established: the validity of the full-data benchmark as a proxy for the true ATE, and the treatment of method-dependent exclusion of extreme replications. Because both can change the relative ranking of methods, the paper's conclusions are not yet supported as reported.

major comments (3)
  1. [Sections 2.1 and 2.7] The target estimand is never defined. Absolute bias and coverage are computed relative to the full-data model, i.e., a logistic IPTW model fitted on the complete synthetic data, but the paper does not demonstrate that this model is unbiased for the ATE under the stated DGP. The DGP includes time-varying treatments, survival outcomes, structural missingness, and very sparse binary outcomes; a logistic IPTW model with HC3 standard errors can be misspecified in such settings. If the full-data model is biased, 'absolute bias' and 'coverage' measure deviation from a possibly incorrect model rather than from the truth, and the MICE-versus-alternatives ranking becomes uninterpretable. Please define the true ATE (e.g., via the g-formula or Monte Carlo standardization using the known data-generating probabilities), report the full-data model's bias and coverage against that target, and, if needed
  2. [Section 2.7 and Section 3.1] The exclusion of replications with |ATE|>100 is method-dependent and can induce a selection effect that directly undermines the headline comparison. The paper reports that MICE exclusion rates are negligible while CCA, SMI, and HD have higher exclusion rates, especially under MNAR, 50% missingness, and n=1000. Because bias, coverage, and empirical SE are computed only on the remaining replications, methods with higher exclusion rates are evaluated on a selected subset that may have systematically different properties. This is a collider/selection effect: it can differentially remove the worst-performing replications for the comparator methods and inflate MICE's apparent coverage advantage. The manuscript reports exclusion rates but does not incorporate them into the performance measures or the interpretation. Please add sensitivity analyses that retain all replications (e.g., through win
  3. [Section 5, Table 4, Table 5] The conclusion that MICE was 'the only method to maintain acceptable coverage under MNAR' is not directly supported by the reported results. Table 4 shows no statistically significant coverage difference between MCAR and MNAR pooled over methods, and the paper does not present a method-by-mechanism coverage table or figure. Moreover, Table 5 reports MICE coverage as low as 0.094 in a MAR(Y), rho=-0.7, n=5000 scenario, so 'acceptable coverage' needs an explicit definition and the MNAR-specific evidence should be presented. Please provide a table or figure of coverage by method x mechanism x missing rate and define the acceptable-coverage threshold. Without this, the core conclusion is not verifiable from the manuscript.
minor comments (4)
  1. [Data Availability Statement] The data availability statement contains a placeholder '[ ]'. If the simulation code is genuinely available, the URL or repository identifier should be included; if not, the reproducibility claim should be softened.
  2. [Page 15, table header] The table header 'Txables' appears to be a typo for 'Tables'.
  3. [Section 2.3] The description of MCAR generation as 'simply selects random values of the variable which are then set to a null value' is ambiguous for binary variables. Please clarify whether each value is independently set to missing with probability r, or whether a random subset of rows is selected.
  4. [Section 3.4] The discussion of higher coverage at n=1000 being due to inflated robust standard errors is interesting but is not accompanied by the supporting model-based standard error results. A supplementary figure or table would help readers evaluate this claim.

Circularity Check

0 steps flagged

No circular derivation chain: simulation compares methods against a synthetic full-data benchmark; no fitted target or self-citation chain is used as evidence for the conclusions.

full rationale

The paper is a simulation study, not a derivation of a target quantity from assumptions that already contain the result. The data-generating process is taken from the authors' prior medRxiv preprint (ref 12), but this is an input assumption that defines the synthetic populations; it is not a conclusion that the paper then feeds back into its own evaluation. Performance measures (absolute bias, coverage, empirical SE) are computed against the full-data benchmark model described in Section 2.7, and no parameter of the missing-data methods is fitted to the outcome measure in a way that would make a 'prediction' tautological. The comparison among CCA, SMI, hot deck, and MICE is an empirical ranking of methods under specified scenarios, not a quantity that is equivalent to its inputs by construction. The manuscript's own limitations (Section 4.3) acknowledge modelling choices and potential numerical instability, but none of these limitations amounts to a circular step. The only author-overlap dependency is the DGP citation, which is standard practice for simulation studies and does not make the conclusions load-bearing on an unverified self-citation chain. The skeptic's concern about differential exclusion of extreme replications (Section 2.7 and Table 2) is a potential validity threat but not a circularity: it concerns conditioning on a collider in simulation summaries, not a reduction of the claimed result to its own inputs. Therefore no circular step can be quoted, and the correct score is 0.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The central comparison rests on the simulated data-generating process, the assumed unbiasedness of the full-data benchmark, and the chosen missingness generators. These are domain assumptions, not derived results. The study contains no fitted free parameters in the target-estimation sense; the listed free parameters are design and configuration choices that affect the conclusions.

free parameters (4)
  • MICE imputation count m = 5
    Fixed at 5 imputations; more imputations can improve stability of variance estimates, so MI's advantage is partly configuration-dependent.
  • Extreme ATE exclusion threshold = |ATE|>100
    Replications with extreme estimates are excluded as numerically unstable; exclusion rates differ by method and could favor MI.
  • Extreme propensity score cutoff = PS<0.05 or >0.95
    Chosen proxy for positivity violations; not a formal positivity test and sensitive to sample size and model specification.
  • Scenario factor levels = rho in {-0.1,-0.4,-0.7}; n in {1000,5000}; missing rate in {20%,50%}; mechanism in {MCAR,MAR(A),MAR(Y),MNAR}
    Simulation design inputs; conclusions may not generalize beyond these levels or to other missingness generators.
axioms (5)
  • domain assumption The data-generating process of Velasco-Pardo et al. (ref 12) produces realistic time-varying confounding and valid potential outcomes.
    Used as ground truth for all scenarios; not validated within this paper and is from the authors' own medRxiv preprint.
  • domain assumption Under full data, the logistic IPTW model is correctly specified and recovers the true ATE.
    The full-data model serves as the benchmark for bias and coverage, but the true ATE is never explicitly defined; misspecification would make all reported 'bias' relative to a model rather than the truth.
  • domain assumption Generated missingness mechanisms (single cutoff for MAR and MNAR) are valid and representative of real missingness.
    One implementation per mechanism; no sensitivity analysis across different MAR/MNAR generators.
  • domain assumption No unmeasured confounding under full data.
    The four confounders are treated as sufficient; violations from missingness are the focus.
  • standard math Rubin's rules and HC3 robust standard errors are valid for pooling and inference.
    Standard statistical tools adopted without proof, appropriate for this context.

pith-pipeline@v1.3.0-alltime-deepseek · 9662 in / 12942 out tokens · 141821 ms · 2026-08-01T17:02:51.830547+00:00 · methodology

0 comments
read the original abstract

Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability conditions. We generated synthetic data and conducted a simulation study comparing missing data methods across scenarios varying these factors. Missingness was introduced in both treatment and outcome variables, and we applied stratified hot deck imputation, single mode imputation, multiple imputation with chained equations (MICE), and complete-case analysis. Average treatment effect (ATE) estimates were obtained using logistic regression with propensity score weighting, and we measured coverage, absolute bias and empirical standard errors across 48 scenarios with 500 replications each. Performance was primarily driven by the missingness mechanism and choice of method, with multiple imputation generally achieving better coverage and lower bias than other methods. Missingness location was also important, while missing rate and sample size primarily affected positivity violations, which were most pronounced under MNAR, high missingness and low sample sizes. Exchangeability violations from mild to moderate confounding were adequately controlled for by propensity score models, whereas strong confounding produced a modest decrease in coverage. Further research should examine additional ways identifiability conditions can be violated under missingness, using more advanced methods and more complex missingness scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

23 extracted references · 1 canonical work pages

  1. [1]

    Causal Inference

    Hernán MA, Robins JM. Causal Inference. CRC Press; 2021

  2. [2]

    Statistical Analysis with Missing Data

    Little RJA, Rubin DB. Statistical Analysis with Missing Data. 2nd ed. John Wiley & Sons; 2002

  3. [3]

    Much ado about nothing: a comparison of missing data methods and software to fit incomplete data regression models

    Horton NJ, Kleinman KP. Much ado about nothing: a comparison of missing data methods and software to fit incomplete data regression models. Am Stat. 2007;61(1):79-90

  4. [4]

    The central role of the propensity score in observational studies for causal effects

    Rosenbaum PR, Rubin DB. The central role of the propensity score in observational studies for causal effects. Biometrika. 1983;70(1):41-55

  5. [5]

    Comparison of several imputation methods for missing baseline data in propensity scores analysis of binary outcome

    Crowe BJ, Lipkovich IA, Wang O. Comparison of several imputation methods for missing baseline data in propensity scores analysis of binary outcome. Pharm Stat. 2010;9(4):269- 279

  6. [6]

    A simulation study of missing data with multiple missing X's

    Rubright JD, Nandakumar R, Glutting JJ. A simulation study of missing data with multiple missing X's. Pract Assess Res Eval. 2014;19(10):1-19

  7. [7]

    A simulation-based bias analysis to assess the impact of unmeasured confounding when designing non-randomized database studies

    Desai RJ, Bradley MC, Lee H, et al. A simulation-based bias analysis to assess the impact of unmeasured confounding when designing non-randomized database studies. Am J Epidemiol. 2024

  8. [8]

    A comparison of different methods to handle missing data in the context of propensity score analysis

    Choi J, Dekkers OM, le Cessie S. A comparison of different methods to handle missing data in the context of propensity score analysis. Eur J Epidemiol. 2019;34(1):23-36

  9. [9]

    Causal inference with confounders missing not at random

    Yang S, Wang L, Ding P. Causal inference with confounders missing not at random. Biometrika. 2019;106(4):875–888

  10. [10]

    mice: Multivariate imputation by chained equations in R

    Van Buuren S, Groothuis-Oudshoorn K. mice: Multivariate imputation by chained equations in R. J Stat Softw. 2011;45(3):1-67

  11. [11]

    Using simulation studies to evaluate statistical methods

    Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Stat Med. 2019;38(11):2074-2102

  12. [12]

    Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy

    Velasco-Pardo V, Swallow B, Robertson C, McCowan C. Simulation of synthetic health records for assessment of causal inference methods for vaccine efficacy. medRxiv 2026.07.17.26358308; doi: https://doi.org/10.64898/2026.07.17.26358308

  13. [13]

    Prevalence and demographic variation of cardiovascular, renal, metabolic, and mental health conditions in 12 million English primary care records

    Cooper JN, Nirantharakumar K, Crowe FL, et al. Prevalence and demographic variation of cardiovascular, renal, metabolic, and mental health conditions in 12 million English primary care records. BMC Med Inform Decis Mak. 2023;23(1)

  14. [14]

    How to generate missing data for simulation studies

    Zhang X. How to generate missing data for simulation studies. Quant Methods Psychol. 2023;19(2):100-122

  15. [15]

    Multiple imputation: a primer

    Schafer JL. Multiple imputation: a primer. Stat Methods Med Res. 1999;8(1):3-15. Missing Data Methods in Causal Inference 14

  16. [16]

    A review of hot deck imputation for survey non-response

    Andridge RR, Little RJA. A review of hot deck imputation for survey non-response. Int Stat Rev. 2010;78(1):40-64

  17. [17]

    The prevention and handling of the missing data

    Kang H. The prevention and handling of the missing data. Korean J Anesthesiol. 2013;64(5):402-406

  18. [18]

    Multiple Imputation for Nonresponse in Surveys

    Rubin DB. Multiple Imputation for Nonresponse in Surveys. John Wiley & Sons; 1987

  19. [19]

    Multiple imputation using chained equations: issues and guidance for practice

    White IR, Royston P, Wood AM. Multiple imputation using chained equations: issues and guidance for practice. Stat Med. 2011;30(4):377-399

  20. [20]

    rsimsum: summarise results from Monte Carlo simulation studies

    Gasparini A. rsimsum: summarise results from Monte Carlo simulation studies. J Open Source Softw. 2018;3(26):739

  21. [21]

    Imputation of data missing not at random: artificial generation and benchmark analysis

    Pereira RC, Abreu PH, Rodrigues PP, Figueiredo MAT. Imputation of data missing not at random: artificial generation and benchmark analysis. Expert Syst Appl. 2024;249:123654

  22. [22]

    A simulation study on missing data imputation for dichotomous variables using statistical and machine learning methods

    Ge Y, Li Z, Zhang J. A simulation study on missing data imputation for dichotomous variables using statistical and machine learning methods. Sci Rep. 2023;13(1)

  23. [23]

    Miscellanea

    Barnard J, Rubin DB. Miscellanea. Small-sample degrees of freedom with multiple imputation. Biometrika. 1999;86(4):948-955. Missing Data Methods in Causal Inference 15 Txables TABLE 1 Marginal distributions of treatment (A), outcome (Y) and confounder (Z) variables (n=1000, ρ=-0.4) Value A1 A2 Y1 Y2 Y3 0 80.4% 78.9% 2.6% 5.1% 7.6% 1 19.6% 21.1% 97.4% 94.9...