REVIEW 4 major objections 5 minor 32 references
Multiply Robust Inference of Average Treatment Effects by High-dimensional Empirical Likelihood
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A high-dimensional weighting scheme for average treatment effects stays valid when only the outcome model—or any one of several propensity-score models—is correct.
desk verdict A novel and promising multiply robust ATE estimator, but the OR-only-correct branch rests on unstated sparsity conditions and the key proofs are in the SM. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the high-dimensional empirical likelihood weight, obtained from a constrained optimization that forces the weighted sums of each working propensity score to match the sample proportion while requiring the weighted means of an augmented covariate vector—original basis functions plus derivatives of all propensity models—to stay within a small tolerance of the sample mean. This soft calibration avoids the impossibility of exact covariate balancing when p >> n. The second mechanism is a regularized augmented outcome regression, fit with inverse EL weights, that includes the same augmented covariates and the working propensity scores as regressors. Together, through the KKT
What would settle it
Simulate n = 500, p = 1000 with a true propensity score that is a dense linear function of many covariates (so the population calibration vector is dense), use working propensity models that are all misspecified, and keep the outcome model correct; if the proposed confidence interval's coverage drops noticeably below 95%, the sparsity assumption is carrying the result.
Extended reading notes
Core claim
The central claim is Theorem 1: the proposed estimator for the mean potential outcome admits the same influence function under three alternative correct specifications—any working propensity-score model or a linear combination of them, the outcome regression model, or both. Consequently the resulting confidence interval for the average treatment effect has asymptotically nominal coverage under p >> n whenever at least one of those models is correct, and attains the semiparametric efficiency bound when both a propensity model (or combination) and the outcome model are correct. The estimator is built from empirical likelihood weights with a soft balancing constraint that includes derivatives o
Load-bearing premise
When all propensity-score models are wrong, the proof supposes that an implicitly defined population calibration vector is sparse and that the limiting propensity weights stay bounded away from 0 and 1; if that sparsity fails, the outcome-only robustness branch has no guarantee.
Editorial extensions
If this is right
- If any one working propensity model, any linear combination of them, or the outcome model is correctly specified, the ATE confidence interval has asymptotically nominal coverage even when the covariate dimension grows faster than the sample size.
- When both a propensity model (or combination) and the outcome model are correct, the estimator reaches the semiparametric efficiency bound, matching the best attainable precision.
- Because the construction extends to generalized linear outcomes and to cluster-specific propensity models, multi-source data with unknown subgroup labels can still get valid ATE inference if at least one cluster-based propensity model or the outcome model is right.
- In the right heart catheterization analysis, the multiply robust estimate gives a small, statistically insignificant ATE, in contrast to two doubly robust baselines that report significantly negative effects—showing that the robustness property can change the substantive conclusion.
Reading between the lines
- Editorial inference: the Neyman-orthogonal construction could plausibly be adapted to other low-dimensional targets, such as quantile treatment effects or weighted policy-relevant estimands, since the soft-calibration-plus-augmented-regression mechanism is not tied to the particular influence function of a mean.
- Editorial inference: the outcome-only robustness branch depends on an unverifiable sparsity condition on the population calibration vector; in practice a sensitivity analysis varying the balancing tolerance and the set of working models would reveal how much of the guarantee is carried by that assumption.
- Editorial inference: treating each clustering method or cluster count as one working propensity model turns clustering uncertainty into a robustness property, an implicit connection to model averaging that the paper does not develop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a multiply robust inference procedure for the average treatment effect (ATE) under high-dimensional covariates (p >> n). The method combines high-dimensional empirical likelihood weighting with internal bias correction constraints and soft covariate balancing, together with a regularized augmented outcome regression. The central claim (Theorem 1) is that the resulting estimator has an asymptotically linear expansion with the same influence function under three scenarios: (i) at least one working propensity score model or their linear combination is correctly specified; (ii) the outcome regression model is correctly specified; or (iii) both are correctly specified, in which case the semiparametric efficiency bound is attained. Confidence intervals with nominal coverage follow (Proposition 2). The methodology is extended to generalized linear outcomes and to clustered data with unknown subgroup structure. Simulation studies compare the proposed estimator with five existing high-dimensional doubly robust methods, and the method is applied to the right heart catheterization dataset.
Significance. If the main theorems are correct, the paper makes a substantial contribution: it extends multiply robust inference from the fixed-dimensional setting to high-dimensional covariates, provides a new Neyman-orthogonal construction via soft calibration with augmented covariates and augmented outcome regression, and addresses a practically important clustered-data setting with heterogeneous treatment assignment mechanisms. The simulation evidence is encouraging and the application to the RHC dataset illustrates a situation where the proposed method changes the substantive conclusion. The main reservations are that the central proofs are relegated to a supplementary file that is not part of the preprint, and that the outcome-regression-only branch of the multiply robust claim relies on a nontrivial sparsity condition on an implicitly defined population quantity. These issues need to be resolved or explicitly qualified before the headline claim can be accepted.
major comments (4)
- [Section 4 and Appendix A] Theorem 1 and Proposition 2, as well as all GLM results, are proved only in the supplementary material; Appendix A is a sketch. In particular, the terms (A.4) and (A.5) are declared negligible 'due to the construction of the AOR' without a quantitative bound. These terms are load-bearing for the multiply robust expansion. I am unable to verify the central claim from the manuscript alone. Please either include the full proofs in a posted supplement or, at minimum, state the explicit rates for (A.4) and (A.5) and show how they are controlled by Conditions 5 and 6.
- [Condition 5(iii) and Theorem 1(ii)] The outcome-regression-only-correct branch of the multiply robust claim requires Condition 5(iii): when all working PS models are misspecified, the probability limit lambda_{2,*} of the soft-calibration multipliers must be s_lambda-sparse with s_1^{1/2} s_lambda log(m) = o(n^{1/2}). This is an assumption on an implicit, data-dependent population quantity. The l1 penalty in (3.6) produces sparse finite-sample estimates, but it does not imply that the population limit is sparse. Thus the abstract's statement that valid inference follows if the outcome regression model is correctly specified is stronger than what Theorem 1 actually establishes. This condition should be stated prominently in the abstract and introduction, or the OR-only branch needs a proof that avoids the sparsity of lambda_{2,*}.
- [Condition 6] Condition 6, (s_0 s_1 s_2)^{1/2} log(m) = o(n^{1/2}), is stronger than the condition max{s_1, s_2} log(m)/n^{1/2} = o(1) used in Ning et al. (2020) and Tan (2020a). The authors acknowledge this, but its implications for the practical high-dimensional regime are not discussed. The coupling of the PS and OR sparsity levels is a real restriction on the multiply robust claim. I ask for a discussion of when Condition 6 holds (e.g., concrete sparsity regimes) and whether the condition can be relaxed, or whether it is an artifact of the proof technique.
- [Proposition 1(ii)] Proposition 1(ii) only establishes feasibility of the high-dimensional EL problem when all PS models are misspecified; it does not provide a rate for (e_lambda_2 - lambda_{2,*}). The root-n expansion in Theorem 1(ii) requires such a rate, and Condition 5(iii) is exactly what supplies it. This should be made explicit: the reader should see that the OR-only branch is not a consequence of the EL construction alone, but of the additional sparsity assumption on lambda_{2,*}.
minor comments (5)
- [Notation] Proposition 1 uses the symbol 'wps' while the methodology and Theorem 1 use 'omega_ps'. Please make the notation consistent.
- [Equation (3.7)] The expression 'Di / e_pi_i^2' is rendered ambiguously. It should be clarified that the regression is weighted by e_pi_i^{-2} = (n e_p_i)^2.
- [Condition 5(i)] The Lipschitz condition '|pi_k(X_i, gamma_k) - pi_k(X_i, gamma_k')| <= C |X_i^T(gamma_{k,*} - gamma_k')|' has a scalar left-hand side and a vector-absolute-value right-hand side. Please use |X_i^T(gamma_{k,*} - gamma_k')| or a norm, as appropriate.
- [Section 5] The definition of e_b_i^(g) is hard to parse because 'b^(g)' and 'bb(X_i)' are used without explicitly stating dimensions. Please define all components clearly.
- [Section 8] The interpretation of the RHC result as 'RHC treatment had an insignificant impact' is based on one confidence interval. A more cautious wording would acknowledge that the conclusion depends on the correctness of the identification assumptions and the estimated clustering.
Circularity Check
No significant circularity: the multiply robust estimator is defined by an explicit EL/AOR construction, and Theorem 1 is a genuine asymptotic result under stated conditions; no load-bearing step reduces to its own inputs by construction, and the paper contains no self-citations.
full rationale
The paper's central claim is a theorem, not a fitted or renamed input. The estimator is explicitly defined via the constrained EL problem (3.2)-(3.4) and the regularized augmented outcome regression (3.7). The influence-function expansions in Theorem 1 are derived in Appendix A using Taylor expansions and standard high-dimensional results; the multiply robust property is a consequence of Neyman orthogonality engineered through the soft calibration and AOR, and the proof is deferred to the supplement rather than assumed. The conditions (1-6) are regularity and sparsity assumptions; in particular, Condition 5(iii) requiring sparsity of the probability limit lambda_{2,*} when all PS models are misspecified is a strong assumption that restricts the OR-only-correct branch, and Condition 6 is admittedly stronger than in Ning et al. (2020) and Tan (2020a). This is an honest limitation and an overclaim in the abstract relative to the theorem, but it is not circular: the theorem does not assume its conclusion. All external citations (Chang et al. 2018/2021 for high-dimensional EL feasibility, van de Geer 2008 for Lasso rates, Jin et al. 2017 and Abbe et al. 2022 for spectral clustering) are independent, non-self-citations and provide standard mathematical support. No equation is defined in terms of the target parameter, and no fitted quantity is relabeled as a prediction. Therefore the derivation chain is self-contained and no circular step is present.
Assumptions & free parameters
free parameters (3)
- soft balancing tolerance ωps =
ωps ≍ {s1 log(m)/n}^{1/2} in theory; selected by 5-fold CV in simulations
- AOR penalty ωor =
ωor ≍ {log(m)/n}^{1/2} in theory; selected by 5-fold CV
- number of clusters v =
v = 2 in Section 7.2 and Section 8
assumptions (8)
- domain assumption Unconfoundedness: Y(d) ⊥ D | X (Condition 1)
- domain assumption Overlap: c0 ≤ π(X) ≤ 1 - c0 (Condition 2)
- domain assumption Sub-Gaussianity of augmented covariates b*(X), OR errors, and bounded errors under a misspecified OR model (Condition 3)
- domain assumption Bounded eigenvalues of Σ = Cov((b*(X), π*)) (Condition 4)
- standard math Convergence rates of PS estimates γ̂ (Condition 5(i)-(ii))
- ad hoc to paper Sparsity of the limit soft-calibration multiplier λ2,* when all PS models are misspecified (Condition 5(iii))
- standard math Coupled sparsity growth (s0 s1 s2)^{1/2} log m = o(n^{1/2}) (Condition 6)
- domain assumption Clustering error rate O(s1 log p / n) (Condition 7)
invented entities (1)
-
Latent subgroup indicator U_i = (U_i1,...,U_iv) in model (6.1)
Cite this review
Pith. "Pith review of Multiply Robust Inference of Average Treatment Effects by High-dimensional Empirical Likelihood." pith.science (2026). https://pith.science/paper/5U2HRI25
@misc{pith2026250900312,
author = {Pith},
title = {Pith review of: Multiply Robust Inference of Average Treatment Effects by High-dimensional Empirical Likelihood},
year = {2026},
howpublished = {\url{https://pith.science/paper/5U2HRI25}},
note = {Machine review of arXiv:2509.00312}
}
read the original abstract
In this paper, we develop a multiply robust inference procedure of the average treatment effect (ATE) for data with high-dimensional covariates. We consider the case where it is difficult to correctly specify a single parametric model for the propensity scores (PS). For example, the target population is formed from heterogeneous sources with different treatment assignment mechanisms. We propose a novel high-dimensional empirical likelihood weighting method under soft covariate balancing constraints to combine multiple working PS models. An extended set of calibration functions is used, and a regularized augmented outcome regression is developed to correct the bias due to non-exact covariate balancing. Those two approaches provide a new way to construct the Neyman orthogonal score of the ATE. The proposed confidence interval for the ATE achieves asymptotically valid nominal coverage under high-dimensional covariates if any of the PS models, their linear combination, or the outcome regression model is correctly specified. The proposed method is extended to generalized linear models for the outcome variable. Specifically, we consider estimating the ATE for data with unknown clusters, where multiple working PS models can be fitted based on the estimated clusters. Our proposed approach enables robust inference of the ATE for clustered data. We demonstrate the advantages of the proposed approach over the existing doubly robust inference methods under high-dimensional covariates via simulation studies. We analyzed the right heart catheterization dataset, initially collected from five medical centers and two different phases of studies, to demonstrate the effectiveness of the proposed method in practice.
Figures
Reference graph
Works this paper leans on
-
[1]
Abbe, E., Fan, J., and Wang, K. (2022). An _p theory of PCA and spectral clustering. Ann. Statist. , 50(4):2359--2385
work page 2022
-
[2]
Athey, S., Imbens, G. W., and Wager, S. (2018). Approximate residual balancing: debiased inference of average treatment effects in high dimensions. Journal of the Royal Statistical Society Series B: Statistical Methodology , 80(4):597--623
work page 2018
-
[3]
Bang, H. and Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics , 61(4):962--973
work page 2005
-
[4]
Belloni, A., Chernozhukov, V., and Hansen, C. (2014). Inference on treatment effects after selection among high-dimensional controls. Review of Economic Studies , 81(2):608--650
work page 2014
-
[5]
Cai, T. T., Ma, J., and Zhang, L. (2019). Chime: Clustering of high-dimensional gaussian mixtures with em algorithm and its optimality. Annals of Statistics , 47(3):1234--1267
work page 2019
-
[6]
Cao, W., Tsiatis, A. A., and Davidian, M. (2009). Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data. Biometrika , 96(3):723--734
work page 2009
-
[7]
Chan, K. C. G. and Yam, S. C. P. (2014). Oracle, Multiple Robust and Multipurpose Calibration in a Missing Response Problem . Statistical Science , 29(3):380 -- 396
work page 2014
-
[8]
Chang, J., Chen, S. X., Tang, C. Y., and Wu, T. T. (2021). High-dimensional empirical likelihood inference. Biometrika , 108(1):127--147
work page 2021
Show all 32 references
-
[9]
Y., and Wu, T
Chang, J., Tang, C. Y., and Wu, T. T. (2018). A new scope of penalized empirical likelihood with high-dimensional estimating equations. The Annals of Statistics , 46(6B):3185--3216
2018
-
[10]
and Haziza, D
Chen, S. and Haziza, D. (2017). Multiply robust imputation procedures for the treatment of item nonresponse in surveys. Biometrika , 104(2):439--453
2017
-
[11]
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters: Double/debiased machine learning. The Econometrics Journal , 21(1)
2018
-
[12]
F., Speroff, T., Dawson, N
Connors, A. F., Speroff, T., Dawson, N. V., Thomas, C., Harrell, F. E., Wagner, D., Desbiens, N., Goldman, L., Wu, A. W., Califf, R. M., et al. (1996). The effectiveness of right heart catheterization in the initial care of critically III patients . JAMA: Journal of the Americ...
1996
-
[13]
Farrell, M. H. (2015). Robust inference on average treatment effects with possibly more covariates than observations. Journal of Econometrics , 189(1):1--23
2015
-
[14]
A., Fleischmann, K
Fleisher, L. A., Fleischmann, K. E., et al. (2014). 2014 ACC/AHA guideline on perioperative cardiovascular evaluation and management of patients undergoing noncardiac surgery: a report of the American College of Cardiology/American Heart Association Task Force on practice guid...
2014
-
[15]
R., Kanwar, M., et al
Garan, A. R., Kanwar, M., et al. (2020). Complete hemodynamic profiling with pulmonary artery catheters in cardiogenic shock is associated with lower in-hospital mortality. JACC: Heart Failure , 8(11):903--913
2020
-
[16]
Hahn, J. (1998). On the role of the propensity score in efficient semiparametric estimation of average treatment effects. Econometrica , pages 315--331
1998
-
[17]
Han, P. (2014). Multiply robust estimation in regression analysis with missing data. Journal of the American Statistical Association , 109(507):1159--1173
2014
-
[18]
and Imbens, G
Hirano, K. and Imbens, G. W. (2001). Estimation of causal effects using propensity score weighting: An application to data on right heart catheterization. Health Services and Outcomes research methodology , 2:259--278
2001
-
[19]
and Ratkovic, M
Imai, K. and Ratkovic, M. (2014). Covariate balancing propensity score. Journal of the Royal Statistical Society Series B: Statistical Methodology , 76(1):243--263
2014
-
[20]
T., and Wang, W
Jin, J., Ke, Z. T., and Wang, W. (2017). Phase transitions for high dimensional clustering and related problems . Ann. Statist. , 45(5):2151--2189
2017
-
[21]
Kang, J. D. and Schafer, J. L. (2007). Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data. Statistical Science , 22(4):523--539
2007
-
[22]
Li, W., Gu, Y., and Liu, L. (2020). Demystifying a class of multiply robust estimators. Biometrika , 107(4):919--933
2020
-
[23]
Ning, Y., Sida, P., and Imai, K. (2020). Robust estimation of causal effects via a high-dimensional covariate balancing propensity score. Biometrika , 107(3):533--554
2020
-
[24]
Rosenbaum, P. R. and Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika , 70(1):41--55
1983
-
[25]
Stevenson, L. W. (2005). Evaluation study of congestive heart failure and pulmonary artery catheterization effectiveness: the escape trial. JAMA: Journal of the American Medical Association , 294(13)
2005
-
[26]
Tan, Z. (2020a). Model-assisted inference for treatment effects using regularized calibrated estimation with high-dimensional data. The Annals of Statistics , 48(2):811--837
-
[27]
Tan, Z. (2020b). Regularized calibrated estimation of propensity scores with model misspecification and high-dimensional data. Biometrika , 107(1):137--158
-
[28]
B., Ritov, Y., and Dezeure, R
van de Geer, Sara, P. B., Ritov, Y., and Dezeure, R. (2014). On asymptotically optimal confidence regions and tests for high-dimensional models. The Annals of Statistics , 42(3):1166--1202
2014
-
[29]
van de Geer, S. A. (2008). High-dimensional generalized linear models and the lasso. The Annals of Statistics , 36(2):614
2008
-
[30]
and Vansteelandt, S
Vermeulen, K. and Vansteelandt, S. (2015). Bias-reduced doubly robust estimation. Journal of the American Statistical Association , 110(511):1024--1036
2015
-
[31]
Wang, Z., B \"u hlmann, P., and Guo, Z. (2023). Distributionally robust machine learning with multi-source data. arXiv preprint arXiv:2309.02211
2023 arXiv
-
[32]
and Zhang, S
Zhang, C.-H. and Zhang, S. S. (2014). Confidence intervals for low dimensional parameters in high dimensional linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology , 76(1):217--242
2014
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.