Pith. sign in

REVIEW 5 minor 38 references

Regression-adjusted average treatment effect estimates in stratified randomized experiments

T0 review · 0 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read In stratified randomized experiments, regression-adjusted estimates of the average treatment effect are consistent, asymptotically normal, and asymptotically no more variable than unadjusted estimates.

desk verdict Solid many-small-strata extension of design-based regression adjustment, but the variance-reduction guarantee requires common treatment proportions and the abstract oversells it. read the letter →

arxiv 1908.01628 v2 pith:FOUFZOTR submitted 2019-08-05 math.ST stat.TH

classification math.STstat.TH MSC 62D0562F1262K10
keywords BlockingRandomized-blockdesignRandomizedexperimentsRandomization-basedinferenceStratifiedsamplingRegressionadjustmentAveragetreatmenteffectFinitepopulationcentrallimittheorem
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves that in stratified randomized experiments, linear-regression adjustment for baseline covariates is a safe efficiency tool: the adjusted average-treatment-effect estimate $\hat{\tau}_{\mathrm{ols}}$ is consistent and asymptotically normal, and, when treatment proportions converge to a common value across strata, its asymptotic variance is no larger than that of the stratified difference in means $\hat{\tau}_{\mathrm{unadj}}$. The theory is randomization-based: potential outcomes and covariates are fixed, and all randomness comes from treatment assignment. It covers the regime in which the number of strata grows with the sample size, including many small strata, and also handles a few large strata with a stratum-specific estimator. If the theorems are right, applied researchers can regress outcomes on covariates and treatment-by-covariate interactions after stratifying, then use the provided conservative variance estimator to get large-sample confidence intervals that are at least as short as the unadjusted ones.

What carries the argument

The load-bearing mechanism is a finite-population central limit theorem for stratified random samples whose main condition is that the largest squared deviation of an outcome from its stratum mean, scaled by $N$, goes to zero; this replaces the usual triangular-array condition with a more interpretable bound. On top of it, the paper studies a weighted linear regression of the outcome on treatment, stratum indicators, centered covariates, and treatment-by-covariate interactions. The variance comparison is carried by the population projections $y_{ij}(z)=y_{i\cdot}(z)+(X_{ij}-X_{i\cdot})^{\top}\beta_z+\varepsilon_{ij}(z)$ for $z=0,1$; because the projection errors are orthogonal to the covariates, the cross terms in $N(\sigma^2_{\mathrm{unadj}}-\sigma^2_{\mathrm{ols}})$ vanish when treatment proportions converge to a common $p$, leaving the nonnegative quadratic form $\sum_i c_i \bar{\beta}_i^{\top}S_{iXX}\bar{\beta}_i/\{p_i(1-p_i)\}$ as the efficiency gain.

What would settle it

Consider a two-stratum design with $p_1\to 0.2$, $p_2\to 0.8$, a strong covariate in both strata, and bounded potential outcomes, and compute the limit of $N(\sigma^2_{\mathrm{ols}}-\sigma^2_{\mathrm{unadj}})$ from the paper's formulas; if any such sequence makes that limit positive, the 'never worse' claim fails. A simulation under Conditions 1--6 should also show $\hat{\sigma}^2_{\mathrm{ols}}\le \hat{\sigma}^2_{\mathrm{unadj}}$ in probability; a configuration where the adjusted interval is systematically wider would delimit the theorem.

Watch

Extended reading notes

Core claim

The paper's central claim is that covariate adjustment after stratified randomization is asymptotically no worse than stratification alone. Its main result, Theorem 3, states that under Conditions 1--6, with at least two treated and two control units per stratum, $(\hat{\tau}_{\mathrm{ols}}-\tau)/\sigma_{\mathrm{ols}}$ converges in distribution to $N(0,1)$, and if $p_i$ converges uniformly to a common $p$, the difference between the asymptotic variances of $\sqrt{N}\hat{\tau}_{\mathrm{ols}}$ and $\sqrt{N}\hat{\tau}_{\mathrm{unadj}}$ is the limit of $-\sum_i c_i \bar{\beta}_i^{\top}S_{iXX}\bar{\beta}_i/\{p_i(1-p_i)\}\le 0$, where $\bar{\beta}_i=(1-p_i)\beta_1+p_i\beta_0$ combines the population projection coefficients for treatment and control. Theorem 2 supplies the supporting finite-population central limit theorem for the unadjusted estimator, allowing the number of strata to tend to infinity. The paper also proves an analogous result, Theorem 4, for a stratum-interacted estimator in designs with a few large strata, and provides conservative variance estimators that yield large-sample confidence intervals with at least nominal coverage.

Load-bearing premise

The headline guarantee that regression adjustment never increases variance depends on every stratum's treatment fraction converging to the same number $p$; if treatment fractions stay different across strata, extra terms appear in the variance comparison that the proof does not control.

Editorial extensions

If this is right

  • In experiments with many small strata, researchers can use $\hat{\tau}_{\mathrm{ols}}$ with the conservative variance estimator $\hat{\sigma}^2_{\mathrm{ols}}$; the resulting intervals have asymptotic coverage at least nominal and are asymptotically no wider than intervals from $\hat{\tau}_{\mathrm{unadj}}$.
  • Even when treatment proportions differ across strata, the stratified difference-in-means estimator is consistent and asymptotically normal, so the paper's central limit theorem applies to designs with many small strata; only the variance-reduction claim needs the common-$p$ condition.
  • With a few large strata and heterogeneous covariate-outcome relationships, the stratum-interacted estimator $\hat{\tau}_{\mathrm{ols,int}}$ is asymptotically at least as efficient as the common-coefficient estimator $\hat{\tau}_{\mathrm{ols}}$.
  • The conservative variance estimators mean confidence intervals based on the adjusted estimators are asymptotically at least as short as those based on the unadjusted estimator, in line with the simulations showing 4--19% shorter intervals.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's equal-$p$ condition, a direct numerical study of the cross terms in $N(\sigma^2_{\mathrm{ols}}-\sigma^2_{\mathrm{unadj}})$ could show whether the no-worse guarantee survives when treatment shares differ across strata; the paper's proof does not cover that case.
  • The variance-difference formula identifies exactly where efficiency gains come from, so before fitting the adjusted regression one could compute a sample analogue of $\sum_i c_i \bar{\beta}_i^{\top}S_{iXX}\bar{\beta}_i/\{p_i(1-p_i)\}$ as a planning diagnostic for how much a given covariate is worth.
  • An immediate extension the paper leaves open is high-dimensional regression adjustment; its projection-and-residual argument would need a concentration inequality for stratified sampling, and the paper names that as the main technical obstacle.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper develops randomization-based inference for average treatment effects in stratified randomized experiments, allowing the number of strata to grow with the sample size. It re-establishes a finite-population central limit theorem for stratified samples (Theorem 1), proves asymptotic normality and conservative variance estimation for the stratified difference-in-means estimator (Theorem 2), and analyzes regression adjustment with treatment-by-covariate interactions (Theorem 3). Under Conditions 1–6, the regression-adjusted estimator is consistent and asymptotically normal; when the stratum-specific treatment proportions converge uniformly to a common limit, its asymptotic variance is no larger than that of the unadjusted estimator, and a conservative variance estimator is provided. A second estimator with stratum-specific regression coefficients is treated for the few-large-strata regime (Theorem 4 and Corollary 1). Simulations and an application to an iron-deficiency schooling trial illustrate the methods.

Significance. If the results hold, this is a valuable extension of the completely randomized regression-adjustment theory of Lin and of Li and Ding to stratified designs with many small strata. The paper gives precise conditions, a transparent variance decomposition, and conservative variance estimators that support large-sample confidence intervals, and it carefully distinguishes the regimes where common regression coefficients versus stratum-specific coefficients are appropriate. The proof appendix is detailed and largely self-contained, and the simulation study covers several design regimes with results that match the theoretical predictions. The main caveat—that the advertised variance reduction for the first estimator requires the asymptotic treatment proportions to be common across strata—is explicitly stated in Condition 1 and Remark 3, although it is underemphasized in the abstract and in the empirical discussion.

minor comments (5)
  1. [Abstract and Section 6] The abstract states that the asymptotic variance of the regression-adjusted estimator is 'no greater' than that of the difference-in-means estimator without mentioning the condition pi,infinity = p. This is a genuine scope restriction: Theorem 3 establishes the variance-reduction claim only under this condition, and the application in Section 6 uses strata with pi ranging from 0.636 to 0.688, so the reported shorter confidence intervals for tau_ols are not formally covered by the theorem. Please qualify the abstract and add a sentence in Section 6 noting that the efficiency gain there is empirical rather than guaranteed by Theorem 3.
  2. [Remark 3] Remark 3 correctly acknowledges that the efficiency improvement requires pi to tend uniformly to a common limit, but the discussion would be strengthened by stating explicitly that, when treatment proportions differ across strata, the cross terms 2*sum_i c_i p_i^{-1} S_iXepsilon(1)^T beta_i and 2*sum_i c_i (1-p_i)^{-1} S_iXepsilon(0)^T beta_i need not vanish and are not sign-definite, so no no-worse claim is made in that case.
  3. [Section 4.2, equation for tau_i,ols_int] In the definition of tau_i,ols_int, the second term appears to contain a typo: it reads {yhat_i.(0) - X_i.}^T betahat_0i, which should presumably be {Xhat_i.(0) - X_i.}^T betahat_0i, matching the treatment-group expression and the preceding text.
  4. [Appendix, Proof of Theorem 3] There is a minor typo: 'Thereom 2' should be 'Theorem 2'. Also, in the proof of Lemma 1, the bound in equation (18) is clear, but the sentence following it could state explicitly which quantities are bounded by Condition 1 and Condition 9, since the current wording is slightly compressed.
  5. [Theorem 1] The condition for the consistency of the variance estimator is stated as 2 <= n1i <= ni - 2 for every stratum, which requires at least two treated and two control units. This is noted in Remark 2, but it would be helpful to repeat the restriction in Theorem 1's statement so that the reader immediately sees the difference from the Bickel-Freedman result cited.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central theorems are proved from independent external benchmarks and from algebraic properties of population projections, not by fitting the conclusions into the inputs.

full rationale

The load-bearing results are self-contained arguments built on external benchmarks. Theorem 1 is proved by showing that the paper's condition m1N/N -> 0 implies the Lindeberg-Feller condition of Bickel and Freedman [4], an independent finite-population central limit theorem; the paper explicitly says 'Our proof relies on the following finite population central limit theorem proved by Bickel and Freedman [4]'. Theorem 2 applies Theorem 1 to the constructed population Pi_a, with Conditions 1-3 ensuring that the needed moment and distance conditions hold; this is a direct derivation, not a circular one. Theorem 3 defines tau_ols as the ordinary least squares estimator from a weighted regression with treatment-by-covariate interactions, and the population projection coefficients beta_1, beta_0 are fixed minimizers of deterministic weighted least squares criteria. The variance comparison is obtained by decomposing N*sigma2_unadj - N*sigma2_ols into a nonnegative quadratic term plus cross terms that vanish because sum_i c_i S_iXepsilon(1) = 0 and because Condition 1 with p_i,infty = p gives uniform convergence beta_i/p_i - beta_infty/p -> 0. That the cross terms vanish in this way is shown by the paper's own algebra, not by defining sigma2_ols to match sigma2_unadj. The consistency of the variance estimators is proved from Lemmas 1 and 2, which are established from sampling theory, and from the Bickel-Freedman CLT, rather than from the conclusion being estimated. Theorem 4 explicitly relies on Proposition 3 of Li and Ding [20], an independent published result, and is a straightforward extension to independent large strata. The only self-citation, reference [22], appears in the Discussion as a pointer to high-dimensional extensions and is not load-bearing for any theorem. The stated restriction that the no-worse-variance claim requires asymptotically common treatment proportions p_i,infty = p is a genuine scope condition, explicitly acknowledged in Condition 1, Theorem 3, and Remark 3; it limits the applicability of the efficiency claim but does not make the derivation circular. No fitted parameter is renamed as a prediction, and no claimed result reduces by construction to its own inputs.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper introduces no free parameters fitted to data and no new postulated entities. Its assumptions are standard domain conditions for randomization-based causal inference, plus regularity conditions on moments and covariate covariance matrices. All are stated explicitly as Conditions 1-7; the central results rest on these rather than on any hidden calibrated constants.

assumptions (6)
  • domain assumption Potential outcomes and covariates are fixed finite-population quantities; randomness arises only from treatment assignment (Neyman-Rubin model).
    Adopted at the start of Section 2; this is the randomization-based inference framework the paper builds on.
  • domain assumption Stable unit treatment value assumption (SUTVA), with no interference and no hidden treatment versions.
    Invoked via citation to Rubin in Section 1; it justifies defining unit-level potential outcomes and the average treatment effect.
  • domain assumption Condition 1: treatment proportions pi are bounded away from 0 and 1 and converge to limits; for Theorem 3, they converge to a common p.
    Needed for uniform control of stratum weights throughout the CLT and variance comparisons.
  • domain assumption Condition 2: maximum within-stratum squared deviations of potential outcomes are o(N).
    This replaces the Lindeberg-Feller condition and is used to verify the Bickel-Freedman CLT in Theorem 1.
  • domain assumption Conditions 4-6: covariates have vanishing maximum squared deviation relative to N, and the weighted covariance matrix converges to a finite invertible limit.
    Required for consistency of the estimated regression coefficients and for the asymptotic normality of the adjusted estimator.
  • standard math Bickel and Freedman's finite-population central limit theorem is accepted as an external theorem.
    The paper restates it as Theorem 5 in the appendix and proves its condition is implied by the paper's own m1N/N → 0 condition.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Regression-adjusted average treatment effect estimates in stratified randomized experiments." pith.science (2026). https://pith.science/paper/FOUFZOTR

@misc{pith2026190801628,
  author       = {Pith},
  title        = {Pith review of: Regression-adjusted average treatment effect estimates in stratified randomized experiments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FOUFZOTR}},
  note         = {Machine review of arXiv:1908.01628}
}
read the original abstract

Researchers often use linear regression to analyse randomized experiments to improve treatment effect estimation by adjusting for imbalances of covariates in the treatment and control groups. Our work offers a randomization-based inference framework for regression adjustment in stratified randomized experiments. Under mild conditions, we re-establish the finite population central limit theorem for a stratified experiment. We prove that both the stratified difference-in-means and the regression-adjusted average treatment effect estimators are consistent and asymptotically normal. The asymptotic variance of the latter is no greater and is typically lesser than that of the former. We also provide conservative variance estimators to construct large-sample confidence intervals for the average treatment effect.

Figures

Figures reproduced from arXiv: 1908.01628 by the authors.

Figure 1
Figure 1. Box plot of average treatment effect estimator minus the true average treatment effect, τˆ − τ . In each sub￾slot of each scenario, the box plots correspond to the methods, from left to right, τˆunadj, τˆols, and τˆols int. In Scenario 1, we do not compute and present τˆols int because both the number of treated and control units are smaller than the number of covariates. 12 [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. Average treatment effect estimates and 95% confidence intervals for three outcomes: number of iron pills taken, average grade score, and average score of Wii games. The circle dots are average treatment effect estimators and the bars are 95% confidence intervals, with the lengths shown on top. control. The proportion of treated units in stratum i, pi , is approximately two third, 0.688, 0.672, 0.652, 0.636, and 0.66… view at source ↗
Figure 3
Figure 3. Changes of average treatment effect estimates (circle dots) and 95% confidence intervals (bars) for three outcomes: number of iron pills taken, average grade score, and average score of Wii games, when the number of covariates changes from five to nine. The confidence interval lengths are shown on top of the bars. The sub-figures in the first line are the results of τˆols, and the sub-figures in the second line are … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 36 canonical work pages

  1. [1]

    and Imbens, G

    Abadie, A. and Imbens, G. W. (2008). Estimation of the con ditional variance in paired experiments. Ann. Econ. Statist., 91/92:175–87

  2. [2]

    Angrist, J. D. and Imbens, G. W. (1995). Two-stage least s quares estimation of average causal effects in models with variable treatment intensity. Journal of the American Statistical Association , 90(430):431–442

  3. [3]

    D., Imbens, G

    Angrist, J. D., Imbens, G. W., and Rubin, D. B. (1996). Ide ntification of causal effects using instrumental variables. Journal of the American Statistical Association , 91(434):444–455

  4. [4]

    Bickel, P . J. and Freedman, D. A. (1984). Asymptotic norm ality and the bootstrap in stratified sampling. Ann. Statist., 12:470–82

  5. [5]

    H., Sekhon, J., and Y u, B

    Bloniarz, A., Liu, H., Zhang, C. H., Sekhon, J., and Y u, B. (2016). Lasso adjustments of treatment effect estimates in randomized experiments. Proc. Natl. Acad. Sci. U.S.A. , 113:7383–90. 18

  6. [6]

    Chong, A., Cohen, I., Field, E., Nakasone, E., and Torero , M. (2016). Iron deficiency and schooling attainment in peru. Am. Econ. J. Appl. Econ , 8:222–55

  7. [7]

    Cochran, W. G. (1977). Sampling Techniques. New Y ork: Wiley, 3rd edition

  8. [8]

    Fisher, R. A. (1926). The arrangement of field experiment s. J. Min. Agric. Gt Br ., 33:503–13

Show all 38 references
  1. [9]

    Fogarty, C. B. (2018a). On mitigating the analytical lim itations of finely stratified experiments. J. R. Statist. Soc. B, 80:1035–56

  2. [10]

    Fogarty, C. B. (2018b). Regression-assisted inferenc e for the average treatment effect in paired experiments. Biometrika, 105:994–1000

  3. [11]

    Freedman, D. A. (2008a). On regression adjustments in e xperiments with several treatments. The Annals of Applied Statistics, 2:176–196

  4. [12]

    Freedman, D. A. (2008b). Randomization does not justif y logistic regression. Statistical Science, 23:237–249

  5. [13]

    Gerber, A. S. and Green, D. P . (2012). Field Experiments: Design, Analysis and Interpretation . New Y ork: Norton

  6. [14]

    J., S¨ avje, F., and Sekhon, J

    Higgins, M. J., S¨ avje, F., and Sekhon, J. S. (2015). Blo cking estimators and inference under the neyman–rubin model. arXiv: 1510.01103

  7. [15]

    Imai, K. (2008). V ariance identification and efficiency analysis in randomized experiments under the matched- pair design. Statist. Med., 27:4857–73

  8. [16]

    Imai, K., King, G., and Stuart, E. A. (2008). Misunderst andings between experimentalists and observationalists about causal inference. J. R. Statist. Soc. A , 171:481–502

  9. [17]

    Imbens, G. W. and Angrist, J. D. (1994). Identification a nd estimation of local average treatment effects. Econometrica, 62(2):467–475

  10. [18]

    Imbens, G. W. and Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sc iences: An Introduction. New Y ork: Cambridge University Press

  11. [19]

    Kempthorne, O. (1955). The randomization theory of exp erimental inference. J. Am. Statist. Assoc., 50:946–67

  12. [20]

    and Ding, P

    Li, X. and Ding, P . (2017). General forms of finite popula tion central limit theorems with applications to causal inference. J. Am. Statist. Assoc. , 112:1759–69

  13. [21]

    Lin, W. (2013). Agnostic notes on regression adjustmen ts to experimental data: Reexamining Freedman’s critique. Ann. Appl. Statist., 7:295–318

  14. [22]

    and Y ang, Y

    Liu, H. and Y ang, Y . (2018). Penalized regression adjusted causal effect estimates in high dimensional random- ized experiments. arXiv preprint arXiv:1809.08732

  15. [23]

    W., Sekhon, J

    Miratrix, L. W., Sekhon, J. S., and Y u, B. (2013). Adjust ing treatment effect estimates by post-stratification in randomized experiments. J. R. Statist. Soc. B , 75:369–96. 19

  16. [24]

    Moore, K. L. and van der Lann, M. J. (2009). Covariate adj ustment in randomized trials with binary outcomes: targeted maximum likelihood estimation. Statistics in Medicine, 28:39–64

  17. [25]

    Morgan, K. L. and Rubin, D. B. (2012). Rerandomization t o improve covariate balance in experiments. Ann. Statist., 40:1263–82

  18. [26]

    Neyman, J. (1935). Statistical problems in agricultur al experimentation (with discussion). Supplement to J. R. Statist. Soc., 2:107–80

  19. [27]

    M., and Speed, T

    Neyman, J., Dabrowska, D. M., and Speed, T. (1990). On th e application of probability theory to agricultural experiments. Statist. Sci., 5:465–72

  20. [28]

    Pashley, N. E. and Miratrix, L. W. (2017). Insights on va riance estimation for blocked and matched pairs designs. arXiv: 1710.10342

  21. [29]

    Robinson, J. (1978). An asymptotic expansion for sampl es from a finite population. Ann. Statist., 6:1005–11

  22. [30]

    Rubin, D. B. (1974). Estimating causal effects of treat ments in randomized and nonrandomized studies. J. Educ. Psychol., 66:688–701

  23. [31]

    Rubin, D. B. (1980). Randomization analysis of experim ental data: the fisher randomization test comment. J. Am. Statist. Assoc., 75:591–93

  24. [32]

    Schochet, P . Z. (2016). Statistical theory for the rct- yes software: Design-based causal inference for the rcts, second edition. Technical report, Washington, DC: U.S. Department of Education, Institute of Education Sciences, National Center for Education Evaluation and Regi...

  25. [33]

    Sen, P . K. (1995). The H´ ajek asymptotics for finite popu lation sampling and their ramifications. Kybernetika, 31:251–68

  26. [34]

    Y ue, L., Li, G., Lian, H., and Wan, X. (2019). Regression adjustment for treatment effect with multicollinearity in high dimensions. Computational Statistics & Data Analysis , 134:17–35

  27. [35]

    A., and Davidian, M

    Zhand, M., Tsiatis, A. A., and Davidian, M. (2008). Impr oving efficiency of inferences in randomized clinical trials using auxiliary covariates. Biometrics, 55:707–715

  28. [36]

    variance weight

    Zhao, A., Ding, P ., Mukerjee, R., and Dasgupta, T. (2018 ). Randomization-based causal inference from split- plot designs. Ann. Statist., 46(5):1876–903. A. Proof of main theorems Proof of Theorem 1 Proof 1 Our proof relies on the following finite population central l imit the...

  29. [37]

    holds, which further implies that the Lindeberg–Feller condition holds. Proof of Theorem 2 Proof 2 We write ˆτi· − τi· as the average of a stratified random sample: ˆτi· − τi· = ni∑ j=1 { Zijyij(1) n1i − (1 − Zij)yij(0) n0i } − {yi·(1) − yi·(0)} = ni∑ j=1 [ Zij{yij(1) − yi·(1)}...

  30. [38]

    Proof of lemmas Proof of lemma 1 Proof 6 Applying standard sampling theory (see for example Cochran [7]), it is easy to show that the mean of sibd is Sibd, i.e., E(sibd) = Sibd

    and (14), we have B∑ i=1 ci { S2 iε(1)/p + S2 iε(0)/(1 − p) } − B∑ i=1 ci { S2 iη(1)/p + S2 iη(0)/(1 − p) } = 1 p B∑ i=1 ci ( β1i − β1 ) T SiXX ( β1i − β1 ) + 1 1 − p B∑ i=1 ci ( β0i − β0 ) T SiXX ( β0i − β0 ) ≥ 0, since the covariance matrix SiXX is positive-definite. Proof of...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.