REVIEW 6 major objections 5 minor 50 references
The optimal estimator for an RCT is a property of the trial family and the analytical goal, not of the estimator alone; a cross-fitting framework identifies it using real historical trials.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-07-31 23:56 UTC pith:UQ2IIXTR
load-bearing objection A correct but narrow cross-fitted estimator comparison framework; the empirical recommendations are in-sample argmins and the abstract contradicts the main results—worth peer review, not acceptance as-is. the 6 major comments →
Towards Optimal Estimators for Randomized Control Trials
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that no estimator is universally optimal; the best estimator belongs to a family of related experiments and an evaluation metric. Given studies that share an outcome and satisfy covariate-conditional study independence, the paper defines expected MSE and expected regret for each estimator and constructs a cross-fitted test statistic: split each study into K folds, fit the candidate estimator on K−1 folds, and score it against a difference-in-means estimate on the held-out fold, which is unbiased for the true effect. Theorem 4.1 proves the MSE comparison is unbiased; the regret comparison has bias bounded by interpretable sign-error terms and vanishes asymptotically. Empi
What carries the argument
Cross-fitted performance comparison on real trial families. For each study the sample is split into K folds; the candidate estimator is trained on the complement and evaluated against a difference-in-means estimate from the held-out fold, using squared error (MSE) or a sign-based regret that captures the cost of wrong launch decisions. The paper's Theorem 4.1 shows the MSE statistic is unbiased, and Theorem 4.2 gives a bias bound for regret that vanishes asymptotically, turning collections of historical RCTs into a testing ground for estimator rankings.
Load-bearing premise
Assumption A.2: conditional on measured covariates, potential outcomes are independent of which study a unit belongs to; if unmeasured study-level differences affect outcomes, the best estimator on past trials may not be best on the next one.
What would settle it
Split a family of historical trials into two non-overlapping halves; use the framework to pick the top estimator on the first half, then measure its MSE and regret on the second half. If the top-ranked estimator is not top-ranked on the held-out half, the transfer assumption fails. A more direct test would show that study membership predicts potential outcomes after adjusting for all observed covariates, e.g., via a placebo test on the trial design.
If this is right
- Organizations running many trials on the same outcome can pre-register the estimator the framework ranks first, replacing post-hoc estimator choice with a data-driven rule.
- Estimator rankings can reverse between MSE and regret, so a single 'best' estimator for both inference and decisions is provably suboptimal on at least one objective.
- The unbiasedness of the cross-fitted MSE statistic means performance gaps observed in the case studies reflect real differences, not overfitting to individual trials.
- Regret comparisons carry a conservative finite-sample bias: sign errors in the benchmark difference-in-means make regret look smaller, a distortion that disappears only as study sizes grow.
Where Pith is reading between the lines
- The framework doubles as a portfolio-level audit: an organization could rerun it periodically to check whether its default estimator still matches its stated objective as trial populations drift.
- A natural extension is to weaken the covariate-conditional study independence assumption with a hierarchical or sensitivity model that allows mechanism differences across studies; without such an extension, rankings are only guaranteed within a tightly controlled family.
- A concrete finite-sample question the paper leaves open is how many studies and how many units per study are needed before the ranking is stable; simulation or resampling from the case-study data could answer it.
- The regret metric could be generalized to asymmetric costs of false positives vs. false negatives; the DM_0 dominance found here may not survive when misses are weighted much more heavily than false alarms.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a framework for choosing treatment-effect estimators within families of RCTs. It defines MSE and regret performance metrics, estimates them via within-study cross-fitting using the difference-in-means estimate on a held-out fold as a proxy for the truth, and builds a permutation test for comparing estimators. The method is applied to 556 Amazon Supply Chain Optimization Technology trials and to 25 Strengthening Democracy Challenge interventions. The paper claims that the optimal estimator is objective-dependent: a covariate-adjusted, winsorized estimator minimizes MSE for inference, while unadjusted difference-in-means minimizes regret for decision-making. Theoretical results include an unbiasedness theorem for the cross-fitted MSE test statistic and bias bounds for the regret-based statistic.
Significance. If the framework is valid, it is a useful step toward principled, domain-specific estimator selection, exploiting the growing availability of large portfolios of real experiments. The use of within-study cross-fitting to separate the candidate estimator from the benchmark is sound for MSE comparisons, and the 556-trial Amazon application is unusually rich. However, the paper's central theoretical claim is not established as stated, the regret comparisons have an acknowledged but unquantified bias that can favor the benchmark estimator, and the key cross-study generalization assumption is untested. The empirical recommendations are also internally inconsistent (WLS vs Linear-T in the abstract/body). The core idea is promising, but the current version does not yet support the headline claims.
major comments (6)
- [Section 4.3 and Appendix B, Eqs. (5)-(8)] Theorem 4.1 is not established as stated because the proof equates the expectation of the cross-fitted squared error with the full-sample performance. Eq. (5) defines epsilon_alpha^(s) using tau_hat_alpha^(s) on the full sample O_s, but the cross-fitted estimate uses tau_hat_(alpha,-k) fitted on O_(s,-k), which has size (1-1/K)m_s. MSE depends on sample size, so E[(tau_hat_(alpha,-k)-tau)^2 | S=s] differs from epsilon_alpha^(s) in finite samples. Unbiasedness would require redefining epsilon_alpha as the split-sample estimator's performance; then the optimal estimator selected by the framework is not the full-sample estimator recommended in Section 5.2.
- [Section 3.1, Assumption A.2] Assumption A.2, which the paper calls 'crucial,' licenses the transition from 'best on the observed 556 studies' to 'best for the family.' The Amazon studies differ by geography, marketplace, product category, and treatment, and the covariates X_i are not described in enough detail to support the claim that they capture all mechanism-relevant differences. Because most treatments are unique, A.2 is largely unfalsifiable from these data. If A.2 fails, the pooled optimum need not generalize to the next trial. The theory simply assumes A.2; the conclusions should be framed as conditional on this assumption or the assumption should be relaxed/tested.
- [Abstract vs. Section 5.2] The abstract states that 'weighted least squares performs best for inference goals,' but Section 5.2 reports that Linear-T_0.005 achieves minimum MSE on the Amazon data, with OLS optimal at winsorization levels below 0.5%. No WLS result is identified as optimal in the body. This contradiction in the paper's headline empirical claim must be resolved before publication.
- [Section 4.3 and Appendix B, Theorem 4.3] For the regret metric, Theorem 4.3 gives E[hat_theta_regret - theta_regret] = (1/N) sum_s (B_s^(beta) - B_s^(alpha)), which is not zero. The bias is not corrected or bounded empirically. Because the difference-in-means estimate hat_tau_DM,k serves as the surrogate truth, the DM estimator's own decision rule is correlated with the benchmark, potentially giving DM_0 a systematic advantage in the regret comparison. The finding that 'DM_0 minimizes regret' may be an artifact of this bias; the paper should quantify the bias or temper the conclusion.
- [Section 4.4, permutation test] The permutation test is justified by claiming that H0: theta(alpha,beta)=0 implies exchangeability of estimator labels on the per-study performance estimates. This is false: equal expected performance does not imply that the two estimators' performance scores are exchangeable. A permutation test of label exchangeability will reject when the estimators have equal means but different variances. A valid test for equality of means, such as a bootstrap or studentized permutation test, is needed.
- [Section 5.1, multi-arm trials] The paper says that for multi-arm trials it 'construct[s] pairwise comparisons between each treatment and control.' If each treatment-control pair is treated as an independent study, then control arms are shared across observations, violating Assumption A.3 (non-overlapping study samples) and inducing dependence in the 556-study aggregate. The analysis should either use the original trial as the unit or account for clustering within multi-arm trials.
minor comments (5)
- [Section 1] Typo: 'the the standard difference-in-means estimator'.
- [Section 7] Typo: 'minimizing regreet' should be 'minimizing regret'.
- [Section A.1] The SDC outcome description says SPV is measured through 'three 7-point scale items' but only SPV 1 and SPV 2 are listed; also Figure 3's caption says 'six estimators' while the legend shows four.
- [Section 4.5] The weighted extension estimates w^(s) on the full data. This breaks the independence between the weights, the candidate estimator, and the DM benchmark used in Theorem 4.1. No theoretical guarantee is given for the weighted test statistic.
- [References] Some references are incomplete or appear mismatched: Parikh et al. (2022) has no venue/arXiv number, and the Losch et al. (2021) citation title concerns optimal transport rather than a comparison of synthetic versus real-world estimator rankings.
Circularity Check
No significant circularity: cross-fitted comparisons are genuine paired tests; empirical recommendations are in-sample summaries, not predictions.
full rationale
The paper's central derivation (Theorems 4.1-4.3) is self-contained. The cross-fitted MSE statistic compares each estimator against a common difference-in-means reference computed on an independent fold; the DM error term cancels in pairwise differences, and the cross term vanishes by sample-splitting independence. This is a substantive paired-comparison argument, not a definitional identity. The regret bounds in Theorems 4.2-4.3 are derived explicitly from the regret definition rather than assumed. The empirical findings (e.g., DM_0 for regret, Linear-T_0.005 for MSE in the Amazon body; OLS_0.1 for some SDC outcomes) are in-sample minimizers of the estimated metrics over the observed 556/25 studies; the paper does not present them as out-of-sample predictions, and its 'actionable guidance' is explicitly conditioned on the observed family. The step from observed best to family-optimal is mediated by Assumption A.2, which is a substantive, potentially untestable premise about cross-study exchangeability, not a circular definition. The only overlapping-author citation (Parikh et al. 2022) is used as context on synthetic-data limitations and is not load-bearing. No estimator is 'predicted' from a parameter fitted to the same target, and no uniqueness theorem or ansatz is imported from the authors' prior work. A separate internal inconsistency (abstract says WLS best for inference while the body says Linear-T_0.005) is a correctness/consistency concern, not circularity.
Axiom & Free-Parameter Ledger
free parameters (5)
- Winsorization grid =
p ∈ {0, 0.001, 0.005, 0.01, 0.025, 0.05, 0.1}
- WLS weight function =
w_i = 1/(Y_pre,i − \bar Y_pre)^2
- Inverse-variance study weights =
w(s) = 1 / Var(\hat τ_DM(s))
- Critical threshold c and regret normalization =
c ∈ R_+; regret normalized at c=1.96 for DM_0
- Number of cross-fitting folds K =
unspecified
axioms (6)
- domain assumption A.1 Randomization within studies: Yi(t) ⟂ Ti | S=s, with overlap
- domain assumption A.2 Covariate-Conditional Study Independence: Yi(t) ⟂ S | Xi
- domain assumption A.3 Non-overlapping study samples: Os ∩ Os' = ∅
- domain assumption Superpopulation two-stage sampling of studies and treatments
- domain assumption Exchangeability of estimator labels under H0 for permutation inference
- domain assumption DM consistency / sign consistency for asymptotic unbiasedness of regret comparisons
read the original abstract
Randomized controlled trials (RCTs) are fundamental tools for causal inference across technology companies, pharmaceutical research, and federal agencies. While the standard difference-in-means estimator provides unbiased treatment effect estimates, it often lacks precision, particularly when treatment effects are heterogeneous or outcomes exhibit heavy-tailed distributions. Although numerous precision-enhancing methods exist---from covariate adjustment techniques to variance reduction strategies---recent research demonstrates that no single estimator performs optimally across all datasets. Rather than seeking the best estimator for individual RCTs, which risks compromising scientific validity through convenient selection, we propose a principled framework for identifying optimal estimators within families of RCTs based on specific analytical goals. Our approach uses sample splitting to estimate the distribution of evaluation metrics (e.g., mean squared error, regret) across RCT families, enabling systematic comparisons between estimators while maintaining asymptotic guarantees. We demonstrate this framework using a sample of Amazon's Supply Chain Optimization Technology trials and the Strengthening Democracy Challenge dataset (25 interventions). Results reveal that optimal estimators vary significantly by analytical objective: weighted least squares performs best for inference goals, while difference-in-means minimizes regret for decision-making contexts. This work provides actionable guidance for estimator selection while preserving methodological rigor across diverse research applications.
Figures
Reference graph
Works this paper leans on
-
[1]
2008 , publisher=
Mostly Harmless Econometrics , author=. 2008 , publisher=
2008
-
[2]
Proceedings of the National Academy of Sciences , volume=
Recursive partitioning for heterogeneous causal effects , author=. Proceedings of the National Academy of Sciences , volume=. 2016 , publisher=
2016
-
[3]
Handbook of Economic Field Experiments , volume=
The Econometrics of Randomized Experiments , author=. Handbook of Economic Field Experiments , volume=. 2017 , publisher=
2017
-
[4]
Athey, Susan and Imbens, Guido and Metzger, Jonas and Munro, Evan , year=. Using. 1909.02210 , archivePrefix=
Pith/arXiv arXiv 1909
-
[5]
The Annals of Statistics , pages=
Bootstrap of the mean in the infinite variance case , author=. The Annals of Statistics , pages=. 1987 , publisher=
1987
-
[6]
arXiv preprint arXiv:2104.00673 , year=
Cross-validation: what does it estimate and how well does it do it? , author=. arXiv preprint arXiv:2104.00673 , year=
-
[7]
2020 , journal=
Cross-validation confidence intervals for test error , author=. 2020 , journal=
2020
-
[8]
Journal of Machine Learning Research , volume=
No unbiased estimator of the variance of k-fold cross-validation , author=. Journal of Machine Learning Research , volume=
-
[9]
Improving precision and power in randomized trials for
Benkeser, David and D. Improving precision and power in randomized trials for. 2020 , month = oct, publisher =. doi:10.1111/biom.13377 , url =
-
[10]
2022 , publisher=
Covariate Adjustment in Randomized Trials , author=. 2022 , publisher=
2022
-
[11]
Proceedings of the National Academy of Sciences , volume=
Lasso Adjustments of Treatment Effect Estimates in Randomized Experiments , author=. Proceedings of the National Academy of Sciences , volume=. 2016 , publisher=
2016
-
[12]
2020 , journal=
Optimal mean estimation without a variance , author=. 2020 , journal=
2020
-
[13]
2018 , publisher=
Double debiased machine learning for treatment and structural parameters , author=. 2018 , publisher=
2018
-
[14]
2023 , eprint=
A First Course in Causal Inference , author=. 2023 , eprint=
2023
-
[15]
How to make a
Drees, Holger and Resnick, Sidney and de Haan, Laurens , journal=. How to make a. 2000 , publisher=
2000
-
[16]
1998 , institution=
Statistical Principles for Clinical Trials (. 1998 , institution=
1998
-
[17]
2014 , journal=
Semiparametric Exponential Families for Heavy-Tailed Data , author=. 2014 , journal=
2014
-
[18]
Advances in Applied Mathematics , volume=
On Regression Adjustments to Experimental Data , author=. Advances in Applied Mathematics , volume=. 2008 , publisher=
2008
-
[19]
American Journal of Epidemiology , volume=
Doubly Robust Estimation of Causal Effects , author=. American Journal of Epidemiology , volume=. 2011 , publisher=
2011
-
[20]
Journal of the American Statistical Association , volume=
The predictive sample reuse method with applications , author=. Journal of the American Statistical Association , volume=. 1975 , publisher=
1975
-
[21]
Advances in Neural Information Processing Systems , volume=
The Case for Evaluating Causal Models Using Interventional Measures and Empirical Data , author=. Advances in Neural Information Processing Systems , volume=. 2019 , url=
2019
-
[22]
Hadad, Vitor , year=
-
[23]
2020 , publisher=
Causal Inference: What If , author=. 2020 , publisher=
2020
-
[24]
2015 , publisher=
Causal Inference in Statistics, Social, and Biomedical Sciences , author=. 2015 , publisher=
2015
-
[25]
and Sollecito, William A
Koch, Gary G. and Sollecito, William A. , title =. Drug Information Journal , volume =. 1984 , doi =
1984
-
[26]
Proceedings of the National Academy of Sciences , volume=
Metalearners for estimating heterogeneous treatment effects using machine learning , author=. Proceedings of the National Academy of Sciences , volume=. 2019 , publisher=
2019
-
[27]
Journal of the American Statistical Association , volume=
Cross-validation with confidence , author=. Journal of the American Statistical Association , volume=. 2020 , publisher=
2020
-
[28]
Agnostic Notes on Regression Adjustments to Experimental Data: Reexamining
Lin, Winston , journal=. Agnostic Notes on Regression Adjustments to Experimental Data: Reexamining. 2013 , publisher=
2013
-
[29]
arXiv preprint arXiv:2112.09266 , year=
Optimal Transport of Data Generates Distributional Robustness , author=. arXiv preprint arXiv:2112.09266 , year=
-
[30]
Foundations of Computational Mathematics , volume=
Mean estimation and regression under heavy-tailed distributions: A survey , author=. Foundations of Computational Mathematics , volume=. 2019 , publisher=
2019
-
[31]
Essay on principles
On the application of probability theory to agricultural experiments. Essay on principles. Section 9 , author=. Statistical Science , pages=. 1990 , publisher=
1990
-
[32]
, author=
Estimating causal effects of treatments in randomized and nonrandomized studies. , author=. Journal of educational Psychology , volume=. 1974 , publisher=
1974
-
[33]
International Conference on Machine Learning , pages=
Orthogonal random forest for causal inference , author=. International Conference on Machine Learning , pages=. 2019 , organization=
2019
-
[34]
2022 , eprint=
Validating Causal Inference Methods , author=. 2022 , eprint=
2022
-
[35]
Statistics in Medicine , volume=
Some methods for heterogeneous treatment effect estimation in high dimensions , author=. Statistics in Medicine , volume=. 2018 , publisher=
2018
-
[36]
Journal of the American Statistical Association , volume=
Estimation of regression coefficients when some regressors are not always observed , author=. Journal of the American Statistical Association , volume=. 1994 , publisher=
1994
-
[37]
Journal of the American Statistical Association , volume=
Causal inference using potential outcomes: Design, modeling, decisions , author=. Journal of the American Statistical Association , volume=. 2005 , publisher=
2005
-
[38]
and Merlin, Vincent R
Saari, Donald G. and Merlin, Vincent R. , journal=. The. 1996 , publisher=
1996
-
[39]
arXiv preprint arXiv:1804.05146 , year=
A comparison of methods for model selection when estimating individual treatment effects , author=. arXiv preprint arXiv:1804.05146 , year=
-
[40]
Journal of the Royal Statistical Society: Series B (Methodological) , volume=
Cross-validatory choice and assessment of statistical predictions , author=. Journal of the Royal Statistical Society: Series B (Methodological) , volume=. 1974 , publisher=
1974
-
[41]
Conference on Learning Theory , pages=
Estimation and inference with trees and forests in high dimensions , author=. Conference on Learning Theory , pages=. 2020 , organization=
2020
-
[42]
2016 , journal=
Scalable semiparametric inference for the means of heavy-tailed distributions , author=. 2016 , journal=
2016
-
[43]
Statistics in Medicine , volume=
Covariate Adjustment for Two-Sample Treatment Comparisons in Randomized Clinical Trials: A Principled yet Flexible Approach , author=. Statistics in Medicine , volume=. 2008 , publisher=
2008
-
[44]
and Stagnaro, Michael N
Voelkel, Jan G. and Stagnaro, Michael N. and Chu, James Y. and Pink, Sophia L. and Mernyk, Joseph S. and Redekopp, Chrystal and Ghezae, Isaias and Cashman, Matthew and Adjodah, Dhaval and Allen, Levi G. and Allis, L. Victor and Baleria, Gina and Ballantyne, Nathan and Van Bavel, Jay J. and Blunden, Hayley and Braley, Alia and Bryan, Christopher J. and Cel...
2024
-
[45]
Proceedings of the National Academy of Sciences , volume=
High-Dimensional Regression Adjustments in Randomized Experiments , author=. Proceedings of the National Academy of Sciences , volume=. 2016 , publisher=
2016
-
[46]
Econometrica , volume=
Efficient estimation of average treatment effects using the estimated propensity score , author=. Econometrica , volume=. 2003 , publisher=
2003
-
[47]
2018 , journal=
Estimation and Inference of Heterogeneous Treatment Effects using Random Forests , author=. 2018 , journal=
2018
-
[48]
Wager, Stefan , year=
-
[49]
Econometrica: Journal of the Econometric Society , pages=
A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity , author=. Econometrica: Journal of the Econometric Society , pages=. 1980 , publisher=
1980
-
[50]
The American Statistician , volume=
Efficiency Study of Estimators for a Treatment Effect in a Pretest-Posttest Trial , author=. The American Statistician , volume=. 2001 , publisher=
2001
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.