{"id":"dc11c95d-2176-4b08-a6ce-9a799f970f27","arxiv_id":"1908.01333","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"In simulations of randomized experiments with 500 people per arm, using the randomization when imputing missing covariates gives only small accuracy gains, while ignoring the outcome can bias heterogeneous treatment effect estimates.","lead":"This paper tests whether imputing missing baseline data in randomized experiments should pool both treatment groups, since randomization makes the groups similar on average. Pooling barely improved accuracy in simulations with 1000 participants, but imputing without using the outcome biased estimates of treatment effects that vary by subgroup.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The small-gains result is established only at n=1000; the sample-size threshold at which gains become material is unmapped, leaving the 'moderately sized' scope of the claim uncertain.","rationale":"The paper is careful, the simulations are extensive, and the abstract is explicitly hedged to 'our simulation scenarios.' The reader's ACCEPT verdict is reasonable. The only load-bearing gap is that the scope statement 'moderately-sized randomized experiments' is supported by a single sample size. The authors already acknowledge the small-arm regime, but not where it begins. A sample-size sweep would either confirm the conclusion across the moderate range or force a sharper qualification. This is a scope concern, not a correctness flaw, so the verdict need not change.","tokens_in":25841,"tokens_out":26017,"duration_ms":270795,"concrete_test":"Re-run Scenario 1 (MAR, high association) with per-arm sample sizes n_arm = 50, 100, 200, 300, and 500, using 1000 replications and the same data-generating process, for Mean-R versus Mean-NR and MI-R versus MI-NR. Compute relative MC-SD differences, (MC-SD_NR - MC-SD_R)/MC-SD_NR, for both beta_t and beta_tx2. If any per-arm n >= 100 shows a relative gain above 5%, the 'small gains' claim must be qualified to larger sample sizes rather than stated for moderately sized experiments generally.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a null result: respecting randomization offers only small gains. The simulation supports this at exactly one sample size, n=1000 (500 per arm). The mechanism the authors themselves invoke in Section 6 is that with arm sizes of roughly 500, imputation-model parameters are estimated accurately whether pooled or not; with arms 'in the tens' gains would be more substantial. But no simulation varies n, so the boundary between 'small' and 'substantial' is never located. Many randomized experiments in the social and biomedical sciences have per-arm sample sizes of 100-300, not 500. At those sizes, the efficiency loss from estimating arm-specific imputation models could be several times the 1-3% differences seen in Table 2. For example, in Table 2 the MC-SD for beta_t under Mean-NR is 0.249 versus 0.241 under Mean-R; at smaller n this gap would grow as the arm-specific mean estimator becomes noisier. Because the abstract and introduction frame the conclusion around 'our simulation scenarios' but also 'moderately-sized randomized experiments,' the unmeasured sample-size dimension is the most load-bearing uncertainty in the paper's scope.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether, when imputing missing baseline covariates in randomized experiments, the analyst should 'respect' the randomization by pooling imputation-model estimation across treatment arms or should ignore it by estimating separate imputation models within arms. The authors compare mean imputation, stochastic regression imputation, and multiple imputation, each implemented in a design-stage version (covariates only) and an outcome-stage version (covariates plus outcome). They evaluate the methods in a simulation with n=1000 (500 per arm), two binary covariates, three missingness scenarios (pre-treatment missingness, post-treatment missingness, and missingness predictive of the outcome), two identifying assumptions (ICIN and MAR), and two strengths of covariate-outcome association, using 1000 replications and reporting bias, MC-SD, estimated SE, coverage, and CI length. The central finding is that respecting versus ignoring randomization makes only small differences in the quality of treatment-effect and heterogeneous-treatment-effect estimates. The paper also compares design-stage versus outcome-stage imputation, finding a bias-efficiency tradeoff for MI and poor performance of outcome-based single imputation methods, and it applies the methods to a real randomized experiment on political ambition.","tokens_in":26129,"tokens_out":10624,"duration_ms":108145,"significance":"If the result holds, it gives practical guidance that, in moderately large randomized trials with binary covariates, the extra effort of building imputation models that enforce the same covariate distribution across arms is unlikely to materially improve inferences, provided the imputation model is otherwise reasonable. The paper also makes a useful methodological contribution by implementing non-parametrically identified multiple imputation under the ICIN assumption for a covariate with heterogeneous treatment effects, and by carefully comparing design-stage and outcome-stage MI. The simulation is thorough: 1000 replications, three missingness mechanisms, two identifying assumptions, complete performance metrics, and an application to real data. The main weaknesses are that the simulation uses only binary covariates, a single sample size, and does not ship code; nevertheless, the empirical claims are well supported by the reported tables.","major_comments":[{"comment":"The central claim that respecting randomization offers only small gains is established at a single sample size, n=1000 (500 per arm). Section 6 itself states that for treatment arms 'in the tens' the gains 'can be more substantial.' Because no simulation varies n, the paper does not locate the boundary between the 'small gains' regime and the 'substantial gains' regime. Many randomized experiments have per-arm sample sizes of 100-300, and the efficiency loss from estimating arm-specific imputation parameters in that range could be larger than the 1-3% MC-SD differences seen in Tables 2-5. I recommend adding a smaller-sample simulation condition (e.g., n=200 or n=400 total) or, at a minimum, revising the abstract and introduction to restrict the conclusion explicitly to n around 1000 and to explain the expected behavior at other sizes.","section":"Section 4.1 and Section 6"},{"comment":"In scenario 2, the missingness indicators (D1,D2) are generated as post-treatment variables depending on T, yet the design-stage 'respecting randomization' method (MI-R) is implemented by collapsing the contingency table over T, which presumes (D1,D2) ⊥⊥ T. This assumption is violated by design in that scenario. The paper does not acknowledge the misspecification or explain why the comparison between MI-R and MI-NR remains a clean comparison of 'respecting' versus 'not respecting' randomization. As written, the scenario 2 results could be interpreted as showing robustness to a false design-stage assumption rather than as evidence about the value of using the randomization. Please clarify the interpretation, or adjust the implementation so that 'respecting randomization' is defined consistently with the data-generating mechanism.","section":"Section 4.1.2 and Section 3.1"}],"minor_comments":[{"comment":"The simulation code is not provided; making it available would strengthen reproducibility and allow readers to probe the sample-size question raised above.","section":"General"},{"comment":"The abbreviation 'NA' is used without definition; please state that it denotes a missing value.","section":"Table 1"},{"comment":"The loglinear model coefficients in Eq. (14) are given without any explanation of how they were chosen; a brief indication of how they yield the desired marginal distributions would aid reproducibility.","section":"Section 4.1.1"},{"comment":"The application has approximately 408 treated and 204 control units, which is still moderately large; the near-identical estimates between respecting and not respecting randomization there do not provide evidence about small-arm behavior.","section":"Section 5"},{"comment":"The Dirichlet prior specification and the handling of sparse missingness patterns (e.g., when the observed contingency table has zero counts) could be clarified; the text currently does not say how zero counts are treated in the posterior draws.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-executed simulation study with honest reporting, and the substantive conclusion is plausible. The main concern is that the 'small gains' finding is demonstrated only at n=1000; the paper's own discussion suggests the gains can be material for arms of size in the tens, but no simulation maps that boundary. I would be comfortable with acceptance if the authors either add a small-n simulation condition or explicitly narrow the claim to the sample sizes studied. The scenario 2 interpretation issue is secondary but should be addressed because it affects the clean comparison of respecting versus not respecting randomization."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the headline finding—that respecting randomization when imputing missing covariates buys little in moderately-sized experiments—is supported by the simulations as far as they go. Second, that 'moderately-sized' is doing a lot of work: every simulation runs at n=1000 (500 per arm), so the boundary at which pooling starts to matter is never located. The authors know this; Section 6 says gains can be more substantial when arms are in the tens. So the result is honest but narrower than the abstract invites.\n\nWhat is new: they extend the earlier comparisons (White and Thompson; Sullivan et al.) to multiple binary covariates, non-ignorable missingness, and heterogeneous treatment effects, and they contribute a principled factorized design/outcome-stage MI procedure that respects randomization without assuming Y⊥T. The supplement derives the extrapolation distributions under ICIN and MAR, which is real, checkable work. The simulation reporting is thorough: 1000 replications, three missingness scenarios, two identifying assumptions, bias/MC-SD/SE/coverage all in tables, and the claims track the tables. The real-data application is a useful sanity check.\n\nSoft spots, in proportion. The single-n design is the main one; the paper would be stronger with a small n=200 or 250 arm condition, which is where many social and biomedical experiments actually sit. Binary covariates only, and no shipped code—the latter is a minor artifact gap, not a scientific one. The design-stage MI methods show notable bias for βtx2 when covariates are prognostic; the authors discuss this trade-off, but readers should not walk away thinking design-stage imputation is generally safer. Those are caveats, not flaws.\n\nOn the stress-test note: I think it lands, but only as a scope warning. The paper says 'for our simulation scenarios,' and it flags the sample-size dependence explicitly. What is missing is a mapping of where gains become material, which would make the guidance actually actionable.\n\nBottom line: this is a useful, honest simulation study for applied statisticians who impute covariates in RCTs, and it deserves a serious referee. I would accept it with minor revisions—add at least one smaller-n condition, release the code, and soften the abstract's scope claims. Cite it when writing about MI in experiments.","headline":"A careful simulation study whose central null result is real but only established at n=1000; worth publishing with the scope made more explicit.","tokens_in":26623,"tokens_out":2210,"would_cite":true,"duration_ms":23463,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D10","62F15","62K99"],"pacs":[],"model":"deepseek-v4-flash","headline":"Respecting randomization in imputation of missing covariates yields only small accuracy gains in moderately sized experiments.","keywords":["missing covariates","randomized experiments","multiple imputation","mean imputation","regression imputation","heterogeneous treatment effects","identifying restrictions","non-ignorable missingness"],"falsifier":"Reproduce the paper's high-association scenario 1 simulation with n=1,000 and 35–40% missingness: the paper reports Monte Carlo standard deviations for the interaction coefficient $\\beta_{tx2}$ that are essentially identical between the respecting-randomization (MI-R) and ignoring-randomization (MI-NR) multiple imputation routines (0.255 in both cases in Table 2). A replication in which these MC-SDs differ by more than 10% would contradict the paper's small-gains claim for its own settings.","tokens_in":25661,"feed_emoji":"🎲","tokens_out":9927,"duration_ms":89600,"temperature":0.7,"pith_summary":"The paper asks whether analysts imputing missing baseline covariates in randomized experiments should exploit the fact that randomization makes covariate distributions identical across treatment arms. Using simulation studies with 1,000 units, two binary covariates, and missingness rates of 35–40%, it compares imputation routines that pool across arms (respecting randomization) with those that impute separately within arms (ignoring randomization). For mean imputation, stochastic regression imputation, and multiple imputation, and for both covariate-only ('design stage') and covariate-plus-outcome ('outcome stage') routines, the accuracy gains from respecting randomization are small in its scenarios. The paper also finds that using the outcome in imputation can reduce bias but increase variance for estimating heterogeneous treatment effects, and that single imputation methods that use the outcome in the way considered here perform poorly. The authors caution that in very small studies, say treatment arms of size in the tens, gains from leveraging randomization can be more substantial.","feed_headline":"Pooling treatment arms barely improves missing-covariate imputation","feed_subtitle":"Simulations with 1,000 patients show accounting for randomization gives only small accuracy gains.","key_machinery":"The central object is a Bayesian multiple imputation model for the joint distribution of the covariates and their missingness indicators, represented as a multinomial/Dirichlet model on the contingency table formed by covariates, missingness indicators, treatment, and optionally outcome. To identify the full-data distribution when missingness is not at random, the paper uses two identifying restrictions: the itemwise conditionally independent non-response (ICIN) assumption, which posits that each variable's missingness is conditionally independent of its own value given all other variables and their missingness indicators, and the missing at random (MAR) assumption. Respecting randomization in the design stage amounts to collapsing the contingency table over treatment, which is justified by the independence of treatment from covariates and missingness indicators; for the outcome stage, the factorization $f(X_{mis},X_{obs},D,T,Y) = f(X_{mis},X_{obs},D,T)\\,f(Y|X_{mis},X_{obs},D,T)$ allows using the outcome without assuming $Y \\perp\\!\\!\\perp T$. The inference machinery is data augmentation via a Polya-Gamma sampler for the logistic outcome model and Dirichlet draws for the table probabilities.","core_discovery":"The paper's central discovery is that, under its simulation settings, properly accounting for randomization—by pooling data across treatment arms when estimating imputation models—makes little practical difference to the quality of treatment effect estimates. For all three imputation approaches (mean, regression, multiple), the Monte Carlo standard deviations, biases, and confidence interval coverage rates are nearly identical between routines that respect and those that ignore randomization. The explanation offered is that with 500 units per arm, the imputation model parameters are already estimated accurately within each arm, and with binary covariates the extra precision from pooling does not translate into materially different imputations. The same holds when the outcome is included in the imputation model. The paper does find a bias-variance trade-off between design-stage and outcome-stage multiple imputation when estimating effect modification: design-stage MI can be biased for the interaction coefficient but more efficient, while outcome-stage MI is less biased but more variable.","pith_inferences":["The paper's explanation for the small gains—that 500 units per arm already yield stable multinomial estimates for binary covariates—suggests a threshold relationship: as sample size shrinks or missingness tables become sparser, the benefit of pooling should become visible. A systematic sweep of sample sizes would turn that threshold into an actionable rule.","If the same comparison were run with continuous covariates, pooling might matter more because within-arm parametric models can be unstable in small arms; the paper's contingency-table framework does not directly address that case.","The paper's negative result for outcome-based single imputation stems largely from variance underestimation; a version that inflates the imputation variance (for example, through repeated draws rather than a single draw) could restore competitiveness, though the paper does not pursue this.","For extremely rare outcomes or sparse covariate combinations, the design-stage MI's bias in the interaction coefficient could dominate its variance advantage, reversing the paper's favorable assessment of design-stage MI in large samples; this is a direct extrapolation of the paper's own large-n thought experiment."],"forward_implications":["Analysts in moderately sized randomized experiments can impute missing covariates within treatment arms without giving up much accuracy in treatment effect estimates, since pooling across arms barely changes point estimates, variances, or coverage.","For estimating heterogeneous treatment effects with multiple imputation, including the outcome in the imputation model reduces bias in the interaction coefficient but increases its variance, so the choice of design-stage versus outcome-stage imputation may be guided by whether bias or efficiency matters more.","The single imputation strategies that impute within treatment-by-outcome cells perform poorly—producing biased estimates and undercoverage—and should not be used for missing covariates.","Because the small-gains result was derived under an outcome model with a treatment-by-covariate interaction, it does not transfer directly to settings with constant treatment effects, where earlier studies found larger benefits from leveraging randomization."],"supporting_citations":[{"why":"Previous simulation work that found precision gains from pooling under mean imputation; the paper's diverging finding for interaction models is the contrast that frames the central claim.","marker":"White and Thompson (2005)"},{"why":"Compared multiple imputation with treatment indicator versus within-arm for a single binary covariate; this paper extends that comparison to multiple covariates and heterogeneous treatment effects.","marker":"Sullivan et al. (2018)"},{"why":"Provides the itemwise conditionally independent non-response assumption and the identification theorem that the paper's Bayesian MI approach relies on.","marker":"Sadinle and Reiter (2017)"},{"why":"Defines multiple imputation and the combining rules that all MI-based estimators in the simulations use.","marker":"Rubin (1987)"},{"why":"Articulates the rationale that omitting the outcome from imputation can attenuate covariate-outcome associations, which motivates the outcome-stage imputation models.","marker":"Little (1992)"},{"why":"Supplies the Polya-Gamma latent variable sampler used to update the logistic regression parameters within the data augmentation MI.","marker":"Polson et al. (2013)"},{"why":"The randomized experiment used as the application; the paper applies all imputation methods to these data and finds results consistent with its simulations.","marker":"Foos and Gilardi (2019)"}],"fun_headline_variants":["Randomization-aware imputation? Small gains only","Pooling treatment arms for imputation? Barely matters","Randomization in imputation: no real accuracy gain","Imputing missing covariates: pooling arms gains little"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on the simulation design: sample size of 1,000 with 500 per arm, two binary covariates, missingness rates around 35–40%, and a logistic outcome model with a single treatment-by-covariate interaction; if real experiments differ materially from these settings, the size of the gains could change.","fun_headline_variants_meta":{"raw":{"variants":["Randomization-aware imputation? Small gains only","Pooling treatment arms for imputation? Barely matters","Randomization in imputation: no real accuracy gain","Imputing missing covariates: pooling arms gains little"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000612,"raw_usage":{"total_tokens":2833,"prompt_tokens":920,"completion_tokens":1913,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":1850}},"tokens_in":536,"tokens_out":1913,"duration_ms":14891,"temperature":1.0,"reasoning_tokens":1850,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:16:18.605273+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reproduce the paper's high-association scenario 1 simulation with n=1,000 and 35–40% missingness: the paper reports Monte Carlo standard deviations for the interaction coefficient $\\beta_{tx2}$ that are essentially identical between the respecting-randomization (MI-R) and ignoring-randomization (MI-NR) multiple imputation routines (0.255 in both cases in Table 2). A replication in which these MC-SDs differ by more than 10% would contradict the paper's small-gains claim for its own settings.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Previous simulation work that found precision gains from pooling under mean imputation; the paper's diverging finding for interaction models is the contrast that frames the central claim."},{"cited_title":"R., White, I","cited_arxiv_id":null,"evidence_quote":"Compared multiple imputation with treatment indicator versus within-arm for a single binary covariate; this paper extends that comparison to multiple covariates and heterogeneous treatment effects."},{"cited_title":"and Reiter, J","cited_arxiv_id":null,"evidence_quote":"Provides the itemwise conditionally independent non-response assumption and the identification theorem that the paper's Bayesian MI approach relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines multiple imputation and the combining rules that all MI-based estimators in the simulations use."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Articulates the rationale that omitting the outcome from imputation can attenuate covariate-outcome associations, which motivates the outcome-stage imputation models."},{"cited_title":"G., Scott, J","cited_arxiv_id":null,"evidence_quote":"Supplies the Polya-Gamma latent variable sampler used to update the logistic regression parameters within the data augmentation MI."},{"cited_title":"and Gilardi, F","cited_arxiv_id":null,"evidence_quote":"The randomized experiment used as the application; the paper applies all imputation methods to these data and finds results consistent with its simulations."}],"review_version":1}