Pith. sign in

REVIEW 4 major objections 5 minor 32 references

Multiply Robust Inference of Average Treatment Effects by High-dimensional Empirical Likelihood

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A high-dimensional weighting scheme for average treatment effects stays valid when only the outcome model—or any one of several propensity-score models—is correct.

desk verdict A novel and promising multiply robust ATE estimator, but the OR-only-correct branch rests on unstated sparsity conditions and the key proofs are in the SM. read the letter →

arxiv 2509.00312 v1 pith:5U2HRI25 submitted 2025-08-30 stat.ME

classification stat.ME
keywords averagetreatmenteffectmultiplyrobustinferencehigh-dimensionalempiricallikelihoodcovariatebalancingcalibrationclusteringpropensityscoreselectionbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the average treatment effect can be estimated with valid confidence intervals in high-dimensional data (more covariates than observations) even when no single propensity-score model is trustworthy, as long as the true propensity score lies in the span of several working models, or the outcome regression is correct. The method combines high-dimensional empirical likelihood weights that only softly balance covariates with a regularized augmented outcome regression; the two pieces are chosen so that errors in estimating all nuisance models do not leak into the treatment-effect estimate. If true, this bridges fixed-dimension multiply robust methods and high-dimensional doubly robust methods, and gives a practical guarantee for heterogeneous or clustered samples where a single propensity model is likely misspecified.

What carries the argument

The central object is the high-dimensional empirical likelihood weight, obtained from a constrained optimization that forces the weighted sums of each working propensity score to match the sample proportion while requiring the weighted means of an augmented covariate vector—original basis functions plus derivatives of all propensity models—to stay within a small tolerance of the sample mean. This soft calibration avoids the impossibility of exact covariate balancing when p >> n. The second mechanism is a regularized augmented outcome regression, fit with inverse EL weights, that includes the same augmented covariates and the working propensity scores as regressors. Together, through the KKT

What would settle it

Simulate n = 500, p = 1000 with a true propensity score that is a dense linear function of many covariates (so the population calibration vector is dense), use working propensity models that are all misspecified, and keep the outcome model correct; if the proposed confidence interval's coverage drops noticeably below 95%, the sparsity assumption is carrying the result.

Watch

Extended reading notes

Core claim

The central claim is Theorem 1: the proposed estimator for the mean potential outcome admits the same influence function under three alternative correct specifications—any working propensity-score model or a linear combination of them, the outcome regression model, or both. Consequently the resulting confidence interval for the average treatment effect has asymptotically nominal coverage under p >> n whenever at least one of those models is correct, and attains the semiparametric efficiency bound when both a propensity model (or combination) and the outcome model are correct. The estimator is built from empirical likelihood weights with a soft balancing constraint that includes derivatives o

Load-bearing premise

When all propensity-score models are wrong, the proof supposes that an implicitly defined population calibration vector is sparse and that the limiting propensity weights stay bounded away from 0 and 1; if that sparsity fails, the outcome-only robustness branch has no guarantee.

Editorial extensions

If this is right

  • If any one working propensity model, any linear combination of them, or the outcome model is correctly specified, the ATE confidence interval has asymptotically nominal coverage even when the covariate dimension grows faster than the sample size.
  • When both a propensity model (or combination) and the outcome model are correct, the estimator reaches the semiparametric efficiency bound, matching the best attainable precision.
  • Because the construction extends to generalized linear outcomes and to cluster-specific propensity models, multi-source data with unknown subgroup labels can still get valid ATE inference if at least one cluster-based propensity model or the outcome model is right.
  • In the right heart catheterization analysis, the multiply robust estimate gives a small, statistically insignificant ATE, in contrast to two doubly robust baselines that report significantly negative effects—showing that the robustness property can change the substantive conclusion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the Neyman-orthogonal construction could plausibly be adapted to other low-dimensional targets, such as quantile treatment effects or weighted policy-relevant estimands, since the soft-calibration-plus-augmented-regression mechanism is not tied to the particular influence function of a mean.
  • Editorial inference: the outcome-only robustness branch depends on an unverifiable sparsity condition on the population calibration vector; in practice a sensitivity analysis varying the balancing tolerance and the set of working models would reveal how much of the guarantee is carried by that assumption.
  • Editorial inference: treating each clustering method or cluster count as one working propensity model turns clustering uncertainty into a robustness property, an implicit connection to model averaging that the paper does not develop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper develops a multiply robust inference procedure for the average treatment effect (ATE) under high-dimensional covariates (p >> n). The method combines high-dimensional empirical likelihood weighting with internal bias correction constraints and soft covariate balancing, together with a regularized augmented outcome regression. The central claim (Theorem 1) is that the resulting estimator has an asymptotically linear expansion with the same influence function under three scenarios: (i) at least one working propensity score model or their linear combination is correctly specified; (ii) the outcome regression model is correctly specified; or (iii) both are correctly specified, in which case the semiparametric efficiency bound is attained. Confidence intervals with nominal coverage follow (Proposition 2). The methodology is extended to generalized linear outcomes and to clustered data with unknown subgroup structure. Simulation studies compare the proposed estimator with five existing high-dimensional doubly robust methods, and the method is applied to the right heart catheterization dataset.

Significance. If the main theorems are correct, the paper makes a substantial contribution: it extends multiply robust inference from the fixed-dimensional setting to high-dimensional covariates, provides a new Neyman-orthogonal construction via soft calibration with augmented covariates and augmented outcome regression, and addresses a practically important clustered-data setting with heterogeneous treatment assignment mechanisms. The simulation evidence is encouraging and the application to the RHC dataset illustrates a situation where the proposed method changes the substantive conclusion. The main reservations are that the central proofs are relegated to a supplementary file that is not part of the preprint, and that the outcome-regression-only branch of the multiply robust claim relies on a nontrivial sparsity condition on an implicitly defined population quantity. These issues need to be resolved or explicitly qualified before the headline claim can be accepted.

major comments (4)
  1. [Section 4 and Appendix A] Theorem 1 and Proposition 2, as well as all GLM results, are proved only in the supplementary material; Appendix A is a sketch. In particular, the terms (A.4) and (A.5) are declared negligible 'due to the construction of the AOR' without a quantitative bound. These terms are load-bearing for the multiply robust expansion. I am unable to verify the central claim from the manuscript alone. Please either include the full proofs in a posted supplement or, at minimum, state the explicit rates for (A.4) and (A.5) and show how they are controlled by Conditions 5 and 6.
  2. [Condition 5(iii) and Theorem 1(ii)] The outcome-regression-only-correct branch of the multiply robust claim requires Condition 5(iii): when all working PS models are misspecified, the probability limit lambda_{2,*} of the soft-calibration multipliers must be s_lambda-sparse with s_1^{1/2} s_lambda log(m) = o(n^{1/2}). This is an assumption on an implicit, data-dependent population quantity. The l1 penalty in (3.6) produces sparse finite-sample estimates, but it does not imply that the population limit is sparse. Thus the abstract's statement that valid inference follows if the outcome regression model is correctly specified is stronger than what Theorem 1 actually establishes. This condition should be stated prominently in the abstract and introduction, or the OR-only branch needs a proof that avoids the sparsity of lambda_{2,*}.
  3. [Condition 6] Condition 6, (s_0 s_1 s_2)^{1/2} log(m) = o(n^{1/2}), is stronger than the condition max{s_1, s_2} log(m)/n^{1/2} = o(1) used in Ning et al. (2020) and Tan (2020a). The authors acknowledge this, but its implications for the practical high-dimensional regime are not discussed. The coupling of the PS and OR sparsity levels is a real restriction on the multiply robust claim. I ask for a discussion of when Condition 6 holds (e.g., concrete sparsity regimes) and whether the condition can be relaxed, or whether it is an artifact of the proof technique.
  4. [Proposition 1(ii)] Proposition 1(ii) only establishes feasibility of the high-dimensional EL problem when all PS models are misspecified; it does not provide a rate for (e_lambda_2 - lambda_{2,*}). The root-n expansion in Theorem 1(ii) requires such a rate, and Condition 5(iii) is exactly what supplies it. This should be made explicit: the reader should see that the OR-only branch is not a consequence of the EL construction alone, but of the additional sparsity assumption on lambda_{2,*}.
minor comments (5)
  1. [Notation] Proposition 1 uses the symbol 'wps' while the methodology and Theorem 1 use 'omega_ps'. Please make the notation consistent.
  2. [Equation (3.7)] The expression 'Di / e_pi_i^2' is rendered ambiguously. It should be clarified that the regression is weighted by e_pi_i^{-2} = (n e_p_i)^2.
  3. [Condition 5(i)] The Lipschitz condition '|pi_k(X_i, gamma_k) - pi_k(X_i, gamma_k')| <= C |X_i^T(gamma_{k,*} - gamma_k')|' has a scalar left-hand side and a vector-absolute-value right-hand side. Please use |X_i^T(gamma_{k,*} - gamma_k')| or a norm, as appropriate.
  4. [Section 5] The definition of e_b_i^(g) is hard to parse because 'b^(g)' and 'bb(X_i)' are used without explicitly stating dimensions. Please define all components clearly.
  5. [Section 8] The interpretation of the RHC result as 'RHC treatment had an insignificant impact' is based on one confidence interval. A more cautious wording would acknowledge that the conclusion depends on the correctness of the identification assumptions and the estimated clustering.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the multiply robust estimator is defined by an explicit EL/AOR construction, and Theorem 1 is a genuine asymptotic result under stated conditions; no load-bearing step reduces to its own inputs by construction, and the paper contains no self-citations.

full rationale

The paper's central claim is a theorem, not a fitted or renamed input. The estimator is explicitly defined via the constrained EL problem (3.2)-(3.4) and the regularized augmented outcome regression (3.7). The influence-function expansions in Theorem 1 are derived in Appendix A using Taylor expansions and standard high-dimensional results; the multiply robust property is a consequence of Neyman orthogonality engineered through the soft calibration and AOR, and the proof is deferred to the supplement rather than assumed. The conditions (1-6) are regularity and sparsity assumptions; in particular, Condition 5(iii) requiring sparsity of the probability limit lambda_{2,*} when all PS models are misspecified is a strong assumption that restricts the OR-only-correct branch, and Condition 6 is admittedly stronger than in Ning et al. (2020) and Tan (2020a). This is an honest limitation and an overclaim in the abstract relative to the theorem, but it is not circular: the theorem does not assume its conclusion. All external citations (Chang et al. 2018/2021 for high-dimensional EL feasibility, van de Geer 2008 for Lasso rates, Jin et al. 2017 and Abbe et al. 2022 for spectral clustering) are independent, non-self-citations and provide standard mathematical support. No equation is defined in terms of the target parameter, and no fitted quantity is relabeled as a prediction. Therefore the derivation chain is self-contained and no circular step is present.

Assumptions & free parameters 3 free parameters · 8 assumptions · 1 invented entities

The central claim rests on standard causal assumptions (unconfoundedness, overlap), standard high-dimensional regularity conditions, and two assumptions particular to this construction: sparsity of the limit soft-calibration multiplier λ2,* when all PS models are misspecified (Condition 5(iii)) and coupled sparsity growth (Condition 6). These are the price of the multiply robust property, and the headline claim in the abstract does not state them. The only 'entity' is the standard latent cluster indicator borrowed from the clustering literature.

free parameters (3)
  • soft balancing tolerance ωps = ωps ≍ {s1 log(m)/n}^{1/2} in theory; selected by 5-fold CV in simulations
    Controls the slack in the soft covariate balancing constraint (3.4); its rate enters the oracle expansions in Theorem 1, and in practice it is a data-dependent tuning parameter.
  • AOR penalty ωor = ωor ≍ {log(m)/n}^{1/2} in theory; selected by 5-fold CV
    Penalty in the regularized weighted regression (3.7); it interacts with the sparsity of the AOR coefficients β1 and β2.
  • number of clusters v = v = 2 in Section 7.2 and Section 8
    The clustered-data procedure requires choosing the number of latent subgroups; the analysis fixes v = 2 for both the simulation and the RHC data.
assumptions (8)
  • domain assumption Unconfoundedness: Y(d) ⊥ D | X (Condition 1)
    The identification of µd and τ from observed data rests entirely on this standard causal assumption; stated in Section 2.
  • domain assumption Overlap: c0 ≤ π(X) ≤ 1 - c0 (Condition 2)
    Needed for the EL weights and influence functions to be well defined; stated in Section 4.
  • domain assumption Sub-Gaussianity of augmented covariates b*(X), OR errors, and bounded errors under a misspecified OR model (Condition 3)
    Basis for the concentration inequalities used in the proofs; stated in Section 4.
  • domain assumption Bounded eigenvalues of Σ = Cov((b*(X), π*)) (Condition 4)
    Compatibility condition for the high-dimensional regularized regressions; stated in Section 4.
  • standard math Convergence rates of PS estimates γ̂ (Condition 5(i)-(ii))
    L1/L2 rates such as max_k ||γ̂k - γk,*||1 ≤ C s1 sqrt(log p / n) are satisfied by Lasso for GLMs (van de Geer 2008; Tan 2020b), as cited in Section 4.
  • ad hoc to paper Sparsity of the limit soft-calibration multiplier λ2,* when all PS models are misspecified (Condition 5(iii))
    Requires the probability limit of the EL dual solution to be sparse in the all-misspecified case; unverifiable from data and not implied by standard estimation theory. The paper itself calls it 'necessary when all the PS models are misspecified' (Section 4).
  • standard math Coupled sparsity growth (s0 s1 s2)^{1/2} log m = o(n^{1/2}) (Condition 6)
    Adapted from high-dimensional EL feasibility (Chang et al. 2021); the authors note it is stronger than the max{s1,s2} log m / n^{1/2} condition of Ning et al. (2020) and Tan (2020a) (Section 4).
  • domain assumption Clustering error rate O(s1 log p / n) (Condition 7)
    Required for the clustered-data extension (Section 6); satisfied by spectral clustering only under sufficient separation of cluster means, which must hold in the data.
invented entities (1)
  • Latent subgroup indicator U_i = (U_i1,...,U_iv) in model (6.1)
    purpose: Models unobserved population heterogeneity so that cluster-specific propensity models can be fitted; enables the clustered-data extension in Sections 6-8.
    U is a standard latent mixture variable borrowed from the high-dimensional clustering literature (Jin et al. 2017; Abbe et al. 2022); it has no falsifiable handle outside the paper, but it is not a novel entity introduced to make the derivation work. Its identifiability is delegated to the clustering algorithm and Condition 7.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multiply Robust Inference of Average Treatment Effects by High-dimensional Empirical Likelihood." pith.science (2026). https://pith.science/paper/5U2HRI25

@misc{pith2026250900312,
  author       = {Pith},
  title        = {Pith review of: Multiply Robust Inference of Average Treatment Effects by High-dimensional Empirical Likelihood},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5U2HRI25}},
  note         = {Machine review of arXiv:2509.00312}
}
read the original abstract

In this paper, we develop a multiply robust inference procedure of the average treatment effect (ATE) for data with high-dimensional covariates. We consider the case where it is difficult to correctly specify a single parametric model for the propensity scores (PS). For example, the target population is formed from heterogeneous sources with different treatment assignment mechanisms. We propose a novel high-dimensional empirical likelihood weighting method under soft covariate balancing constraints to combine multiple working PS models. An extended set of calibration functions is used, and a regularized augmented outcome regression is developed to correct the bias due to non-exact covariate balancing. Those two approaches provide a new way to construct the Neyman orthogonal score of the ATE. The proposed confidence interval for the ATE achieves asymptotically valid nominal coverage under high-dimensional covariates if any of the PS models, their linear combination, or the outcome regression model is correctly specified. The proposed method is extended to generalized linear models for the outcome variable. Specifically, we consider estimating the ATE for data with unknown clusters, where multiple working PS models can be fitted based on the estimated clusters. Our proposed approach enables robust inference of the ATE for clustered data. We demonstrate the advantages of the proposed approach over the existing doubly robust inference methods under high-dimensional covariates via simulation studies. We analyzed the right heart catheterization dataset, initially collected from five medical centers and two different phases of studies, to demonstrate the effectiveness of the proposed method in practice.

Figures

Figures reproduced from arXiv: 2509.00312 by the authors.

Figure 1
Figure 1. Scatter plot of the first two principal components for the observations under [PITH_FULL_IMAGE:figures/full_fig_p029_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 28 canonical work pages

  1. [1]

    Abbe, E., Fan, J., and Wang, K. (2022). An _p theory of PCA and spectral clustering. Ann. Statist. , 50(4):2359--2385

  2. [2]

    W., and Wager, S

    Athey, S., Imbens, G. W., and Wager, S. (2018). Approximate residual balancing: debiased inference of average treatment effects in high dimensions. Journal of the Royal Statistical Society Series B: Statistical Methodology , 80(4):597--623

  3. [3]

    and Robins, J

    Bang, H. and Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics , 61(4):962--973

  4. [4]

    Belloni, A., Chernozhukov, V., and Hansen, C. (2014). Inference on treatment effects after selection among high-dimensional controls. Review of Economic Studies , 81(2):608--650

  5. [5]

    T., Ma, J., and Zhang, L

    Cai, T. T., Ma, J., and Zhang, L. (2019). Chime: Clustering of high-dimensional gaussian mixtures with em algorithm and its optimality. Annals of Statistics , 47(3):1234--1267

  6. [6]

    A., and Davidian, M

    Cao, W., Tsiatis, A. A., and Davidian, M. (2009). Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data. Biometrika , 96(3):723--734

  7. [7]

    Chan, K. C. G. and Yam, S. C. P. (2014). Oracle, Multiple Robust and Multipurpose Calibration in a Missing Response Problem . Statistical Science , 29(3):380 -- 396

  8. [8]

    X., Tang, C

    Chang, J., Chen, S. X., Tang, C. Y., and Wu, T. T. (2021). High-dimensional empirical likelihood inference. Biometrika , 108(1):127--147

Show all 32 references
  1. [9]

    Y., and Wu, T

    Chang, J., Tang, C. Y., and Wu, T. T. (2018). A new scope of penalized empirical likelihood with high-dimensional estimating equations. The Annals of Statistics , 46(6B):3185--3216

  2. [10]

    and Haziza, D

    Chen, S. and Haziza, D. (2017). Multiply robust imputation procedures for the treatment of item nonresponse in surveys. Biometrika , 104(2):439--453

  3. [11]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters: Double/debiased machine learning. The Econometrics Journal , 21(1)

  4. [12]

    F., Speroff, T., Dawson, N

    Connors, A. F., Speroff, T., Dawson, N. V., Thomas, C., Harrell, F. E., Wagner, D., Desbiens, N., Goldman, L., Wu, A. W., Califf, R. M., et al. (1996). The effectiveness of right heart catheterization in the initial care of critically III patients . JAMA: Journal of the Americ...

  5. [13]

    Farrell, M. H. (2015). Robust inference on average treatment effects with possibly more covariates than observations. Journal of Econometrics , 189(1):1--23

  6. [14]

    A., Fleischmann, K

    Fleisher, L. A., Fleischmann, K. E., et al. (2014). 2014 ACC/AHA guideline on perioperative cardiovascular evaluation and management of patients undergoing noncardiac surgery: a report of the American College of Cardiology/American Heart Association Task Force on practice guid...

  7. [15]

    R., Kanwar, M., et al

    Garan, A. R., Kanwar, M., et al. (2020). Complete hemodynamic profiling with pulmonary artery catheters in cardiogenic shock is associated with lower in-hospital mortality. JACC: Heart Failure , 8(11):903--913

  8. [16]

    Hahn, J. (1998). On the role of the propensity score in efficient semiparametric estimation of average treatment effects. Econometrica , pages 315--331

  9. [17]

    Han, P. (2014). Multiply robust estimation in regression analysis with missing data. Journal of the American Statistical Association , 109(507):1159--1173

  10. [18]

    and Imbens, G

    Hirano, K. and Imbens, G. W. (2001). Estimation of causal effects using propensity score weighting: An application to data on right heart catheterization. Health Services and Outcomes research methodology , 2:259--278

  11. [19]

    and Ratkovic, M

    Imai, K. and Ratkovic, M. (2014). Covariate balancing propensity score. Journal of the Royal Statistical Society Series B: Statistical Methodology , 76(1):243--263

  12. [20]

    T., and Wang, W

    Jin, J., Ke, Z. T., and Wang, W. (2017). Phase transitions for high dimensional clustering and related problems . Ann. Statist. , 45(5):2151--2189

  13. [21]

    Kang, J. D. and Schafer, J. L. (2007). Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data. Statistical Science , 22(4):523--539

  14. [22]

    Li, W., Gu, Y., and Liu, L. (2020). Demystifying a class of multiply robust estimators. Biometrika , 107(4):919--933

  15. [23]

    Ning, Y., Sida, P., and Imai, K. (2020). Robust estimation of causal effects via a high-dimensional covariate balancing propensity score. Biometrika , 107(3):533--554

  16. [24]

    Rosenbaum, P. R. and Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika , 70(1):41--55

  17. [25]

    Stevenson, L. W. (2005). Evaluation study of congestive heart failure and pulmonary artery catheterization effectiveness: the escape trial. JAMA: Journal of the American Medical Association , 294(13)

  18. [26]

    Tan, Z. (2020a). Model-assisted inference for treatment effects using regularized calibrated estimation with high-dimensional data. The Annals of Statistics , 48(2):811--837

  19. [27]

    Tan, Z. (2020b). Regularized calibrated estimation of propensity scores with model misspecification and high-dimensional data. Biometrika , 107(1):137--158

  20. [28]

    B., Ritov, Y., and Dezeure, R

    van de Geer, Sara, P. B., Ritov, Y., and Dezeure, R. (2014). On asymptotically optimal confidence regions and tests for high-dimensional models. The Annals of Statistics , 42(3):1166--1202

  21. [29]

    van de Geer, S. A. (2008). High-dimensional generalized linear models and the lasso. The Annals of Statistics , 36(2):614

  22. [30]

    and Vansteelandt, S

    Vermeulen, K. and Vansteelandt, S. (2015). Bias-reduced doubly robust estimation. Journal of the American Statistical Association , 110(511):1024--1036

  23. [31]

    Wang, Z., B \"u hlmann, P., and Guo, Z. (2023). Distributionally robust machine learning with multi-source data. arXiv preprint arXiv:2309.02211

  24. [32]

    and Zhang, S

    Zhang, C.-H. and Zhang, S. S. (2014). Confidence intervals for low dimensional parameters in high dimensional linear models. Journal of the Royal Statistical Society Series B: Statistical Methodology , 76(1):217--242

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.