{"id":"ffe96308-fe2d-4975-9bec-f58b62bceb90","arxiv_id":"1908.01628","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Regression-adjusted average treatment effect estimators are consistent, asymptotically normal, and asymptotically no less efficient than the unadjusted stratified difference-in-means estimator.","lead":"This paper proves that linear-regression adjustment in stratified randomized experiments is safe and efficient: the adjusted estimator is asymptotically normal and never has larger asymptotic variance than the unadjusted difference-in-means, under stated conditions. It also supplies conservative confidence intervals for the average treatment effect.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The no-worse variance claim hinges on pi,∞ = p; without common treatment proportions the cross terms in the variance comparison need not vanish, so the advertised efficiency gain is narrower than the abstract suggests.","rationale":"The reader's weakest assumption identifies exactly the condition I would flag. The proof of the variance comparison in Theorem 3 is clean up to the point where the cross terms are discarded, and discarding them requires pi -> p uniformly; without it the algebra leaves terms that are not sign-definite. I checked the supporting steps (Lemma 1, Lemma 2, Lemma 3, and the application of Theorem 2 to the residual outcomes) and found no additional error that would threaten the theorem as stated: the variance bound in Lemma 1 is valid under Conditions 1 and 9, the decomposition in equations (6)-(8) is correct, and the conservativeness argument for sigma_ols^2 is internally consistent. The common-p condition is explicit in the theorem and Remark 3, so the concern is a scope limitation rather than a defect. It does weaken the abstract's unqualified claim of variance no greater, and it means the empirical application with unequal p_i is not formally covered for the variance-reduction claim. Since the formal result is correctly conditioned, the reader's ACCEPT verdict stands and no change is needed.","tokens_in":27317,"tokens_out":13355,"duration_ms":137605,"concrete_test":"Simulate a many-small-strata sequence satisfying Conditions 1-6 but with unequal limiting treatment proportions, e.g. B=200 strata split into two equal-weight types with pi,infty equal to 0.25 and 0.75, one covariate, and stratum-specific S_iXepsilon(1) chosen so that the unweighted sum c_i S_iXepsilon(1) is zero but the 1/p_i-weighted sum is nonzero. Compare Monte Carlo variances of tau_ols and tau_unadj over many randomizations. If tau_ols has larger variance than tau_unadj, the common-p condition is genuinely load-bearing and the abstract's unqualified wording should be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 3's variance-reduction statement is the paper's central efficiency claim. In the proof, N(sigma2_unadj - sigma2_ols) is decomposed into a nonnegative term sum_i c_i Delta_i^2 plus two cross terms 2 sum_i c_i p_i^{-1} S_iXepsilon(1)^T beta_i and 2 sum_i c_i (1-p_i)^{-1} S_iXepsilon(0)^T beta_i. These cross terms are shown to vanish only because beta_i/p_i - beta_infty/p -> 0 uniformly, which uses Condition 1 together with pi,infty = p. Without the common limit, only sum_i c_i S_iXepsilon(1)=0 and sum_i c_i S_iXepsilon(0)=0 follow from the projection definitions; the 1/p_i-weighted versions need not vanish and are not sign-definite. Thus, for the many-small-strata regime with heterogeneous pi,infty, the paper establishes consistency and asymptotic normality but does not establish that regression adjustment never hurts. This is a genuine scope limitation rather than an internal inconsistency: it is explicit in Condition 1 and Remark 3. It is also practically relevant: in the iron-deficiency application the p_i range from 0.636 to 0.688, so the shorter confidence intervals reported there are not formally covered by the theorem's variance-reduction claim. The theorem as stated remains correct under its conditions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops randomization-based inference for average treatment effects in stratified randomized experiments, allowing the number of strata to grow with the sample size. It re-establishes a finite-population central limit theorem for stratified samples (Theorem 1), proves asymptotic normality and conservative variance estimation for the stratified difference-in-means estimator (Theorem 2), and analyzes regression adjustment with treatment-by-covariate interactions (Theorem 3). Under Conditions 1–6, the regression-adjusted estimator is consistent and asymptotically normal; when the stratum-specific treatment proportions converge uniformly to a common limit, its asymptotic variance is no larger than that of the unadjusted estimator, and a conservative variance estimator is provided. A second estimator with stratum-specific regression coefficients is treated for the few-large-strata regime (Theorem 4 and Corollary 1). Simulations and an application to an iron-deficiency schooling trial illustrate the methods.","tokens_in":27573,"tokens_out":5872,"duration_ms":68299,"significance":"If the results hold, this is a valuable extension of the completely randomized regression-adjustment theory of Lin and of Li and Ding to stratified designs with many small strata. The paper gives precise conditions, a transparent variance decomposition, and conservative variance estimators that support large-sample confidence intervals, and it carefully distinguishes the regimes where common regression coefficients versus stratum-specific coefficients are appropriate. The proof appendix is detailed and largely self-contained, and the simulation study covers several design regimes with results that match the theoretical predictions. The main caveat—that the advertised variance reduction for the first estimator requires the asymptotic treatment proportions to be common across strata—is explicitly stated in Condition 1 and Remark 3, although it is underemphasized in the abstract and in the empirical discussion.","major_comments":[],"minor_comments":[{"comment":"The abstract states that the asymptotic variance of the regression-adjusted estimator is 'no greater' than that of the difference-in-means estimator without mentioning the condition pi,infinity = p. This is a genuine scope restriction: Theorem 3 establishes the variance-reduction claim only under this condition, and the application in Section 6 uses strata with pi ranging from 0.636 to 0.688, so the reported shorter confidence intervals for tau_ols are not formally covered by the theorem. Please qualify the abstract and add a sentence in Section 6 noting that the efficiency gain there is empirical rather than guaranteed by Theorem 3.","section":"Abstract and Section 6"},{"comment":"Remark 3 correctly acknowledges that the efficiency improvement requires pi to tend uniformly to a common limit, but the discussion would be strengthened by stating explicitly that, when treatment proportions differ across strata, the cross terms 2*sum_i c_i p_i^{-1} S_iXepsilon(1)^T beta_i and 2*sum_i c_i (1-p_i)^{-1} S_iXepsilon(0)^T beta_i need not vanish and are not sign-definite, so no no-worse claim is made in that case.","section":"Remark 3"},{"comment":"In the definition of tau_i,ols_int, the second term appears to contain a typo: it reads {yhat_i.(0) - X_i.}^T betahat_0i, which should presumably be {Xhat_i.(0) - X_i.}^T betahat_0i, matching the treatment-group expression and the preceding text.","section":"Section 4.2, equation for tau_i,ols_int"},{"comment":"There is a minor typo: 'Thereom 2' should be 'Theorem 2'. Also, in the proof of Lemma 1, the bound in equation (18) is clear, but the sentence following it could state explicitly which quantities are bounded by Condition 1 and Condition 9, since the current wording is slightly compressed.","section":"Appendix, Proof of Theorem 3"},{"comment":"The condition for the consistency of the variance estimator is stated as 2 <= n1i <= ni - 2 for every stratum, which requires at least two treated and two control units. This is noted in Remark 2, but it would be helpful to repeat the restriction in Theorem 1's statement so that the reader immediately sees the difference from the Bickel-Freedman result cited.","section":"Theorem 1"}],"recommendation":"minor_revision","confidential_remarks":"The mathematical content appears sound. The reader's strongest claim about Theorem 3 is accurate under its stated conditions, and the stress-test concern about heterogeneous pi,infinity is real but already disclosed in Condition 1 and Remark 3. The main needed change is to align the abstract and the empirical discussion with the theorem's scope, which should be a quick revision. I do not see a correctness problem that would require a new proof or a change in the theoretical conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this paper closes a real gap: it gives randomization-based CLTs for stratified experiments when the number of strata grows, and it shows regression adjustment is asymptotically no worse than the unadjusted estimator under a common-treatment-proportion condition. Second, that condition is doing more work than the abstract acknowledges. If treatment proportions differ across strata, the no-worse variance result is not proven; the cross terms in the variance decomposition do not vanish without the common limit.\n\nThe core is Theorem 3. Under Conditions 1–6, tau_ols is consistent and asymptotically normal; when pi,infty = p, its asymptotic variance is no greater than tau_unadj. The proof is detailed, and the variance-estimator consistency is handled carefully. Theorem 2, the finite-population CLT, is a nice rederivation of Bickel–Freedman with a more interpretable condition. Theorem 4 is a straightforward extension of Li–Ding, which the authors say themselves.\n\nThe soft spots are real but not fatal. The common-p assumption is the main one. The abstract says 'no greater' without the qualifier, and the iron-deficiency application has p_i from .636 to .688, so the shorter confidence intervals there are not formally guaranteed by Theorem 3. The theorem remains correct as stated, but the empirical claim goes beyond it. Minor: the variance estimator needs at least two treated and two control per stratum, which excludes paired designs; the authors note this. I was not troubled by the proofs; the appendix is thorough.\n\nBottom line: this is a paper for people working in design-based causal inference. It deserves a serious referee. I would engage with it and likely cite it. The authors should be asked to sharpen the abstract and add a remark about heterogeneous-p, but don't send it back for a fundamental flaw.","headline":"Solid many-small-strata extension of design-based regression adjustment, but the variance-reduction guarantee requires common treatment proportions and the abstract oversells it.","tokens_in":28079,"tokens_out":6645,"would_cite":true,"duration_ms":61778,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62F12","62K10"],"pacs":[],"model":"deepseek-v4-flash","headline":"In stratified randomized experiments, regression-adjusted estimates of the average treatment effect are consistent, asymptotically normal, and asymptotically no more variable than unadjusted estimates.","keywords":["Blocking","Randomized-block design","Randomized experiments","Randomization-based inference","Stratified sampling","Regression adjustment","Average treatment effect","Finite population central limit theorem"],"falsifier":"Consider a two-stratum design with $p_1\\to 0.2$, $p_2\\to 0.8$, a strong covariate in both strata, and bounded potential outcomes, and compute the limit of $N(\\sigma^2_{\\mathrm{ols}}-\\sigma^2_{\\mathrm{unadj}})$ from the paper's formulas; if any such sequence makes that limit positive, the 'never worse' claim fails. A simulation under Conditions 1--6 should also show $\\hat{\\sigma}^2_{\\mathrm{ols}}\\le \\hat{\\sigma}^2_{\\mathrm{unadj}}$ in probability; a configuration where the adjusted interval is systematically wider would delimit the theorem.","tokens_in":27120,"feed_emoji":"📊","tokens_out":12250,"duration_ms":112223,"temperature":0.7,"pith_summary":"This paper proves that in stratified randomized experiments, linear-regression adjustment for baseline covariates is a safe efficiency tool: the adjusted average-treatment-effect estimate $\\hat{\\tau}_{\\mathrm{ols}}$ is consistent and asymptotically normal, and, when treatment proportions converge to a common value across strata, its asymptotic variance is no larger than that of the stratified difference in means $\\hat{\\tau}_{\\mathrm{unadj}}$. The theory is randomization-based: potential outcomes and covariates are fixed, and all randomness comes from treatment assignment. It covers the regime in which the number of strata grows with the sample size, including many small strata, and also handles a few large strata with a stratum-specific estimator. If the theorems are right, applied researchers can regress outcomes on covariates and treatment-by-covariate interactions after stratifying, then use the provided conservative variance estimator to get large-sample confidence intervals that are at least as short as the unadjusted ones.","feed_headline":"Regression adjustment never adds variance in stratified trials","feed_subtitle":"New asymptotic proof covers many small strata and yields conservative confidence intervals.","key_machinery":"The load-bearing mechanism is a finite-population central limit theorem for stratified random samples whose main condition is that the largest squared deviation of an outcome from its stratum mean, scaled by $N$, goes to zero; this replaces the usual triangular-array condition with a more interpretable bound. On top of it, the paper studies a weighted linear regression of the outcome on treatment, stratum indicators, centered covariates, and treatment-by-covariate interactions. The variance comparison is carried by the population projections $y_{ij}(z)=y_{i\\cdot}(z)+(X_{ij}-X_{i\\cdot})^{\\top}\\beta_z+\\varepsilon_{ij}(z)$ for $z=0,1$; because the projection errors are orthogonal to the covariates, the cross terms in $N(\\sigma^2_{\\mathrm{unadj}}-\\sigma^2_{\\mathrm{ols}})$ vanish when treatment proportions converge to a common $p$, leaving the nonnegative quadratic form $\\sum_i c_i \\bar{\\beta}_i^{\\top}S_{iXX}\\bar{\\beta}_i/\\{p_i(1-p_i)\\}$ as the efficiency gain.","core_discovery":"The paper's central claim is that covariate adjustment after stratified randomization is asymptotically no worse than stratification alone. Its main result, Theorem 3, states that under Conditions 1--6, with at least two treated and two control units per stratum, $(\\hat{\\tau}_{\\mathrm{ols}}-\\tau)/\\sigma_{\\mathrm{ols}}$ converges in distribution to $N(0,1)$, and if $p_i$ converges uniformly to a common $p$, the difference between the asymptotic variances of $\\sqrt{N}\\hat{\\tau}_{\\mathrm{ols}}$ and $\\sqrt{N}\\hat{\\tau}_{\\mathrm{unadj}}$ is the limit of $-\\sum_i c_i \\bar{\\beta}_i^{\\top}S_{iXX}\\bar{\\beta}_i/\\{p_i(1-p_i)\\}\\le 0$, where $\\bar{\\beta}_i=(1-p_i)\\beta_1+p_i\\beta_0$ combines the population projection coefficients for treatment and control. Theorem 2 supplies the supporting finite-population central limit theorem for the unadjusted estimator, allowing the number of strata to tend to infinity. The paper also proves an analogous result, Theorem 4, for a stratum-interacted estimator in designs with a few large strata, and provides conservative variance estimators that yield large-sample confidence intervals with at least nominal coverage.","pith_inferences":["Beyond the paper's equal-$p$ condition, a direct numerical study of the cross terms in $N(\\sigma^2_{\\mathrm{ols}}-\\sigma^2_{\\mathrm{unadj}})$ could show whether the no-worse guarantee survives when treatment shares differ across strata; the paper's proof does not cover that case.","The variance-difference formula identifies exactly where efficiency gains come from, so before fitting the adjusted regression one could compute a sample analogue of $\\sum_i c_i \\bar{\\beta}_i^{\\top}S_{iXX}\\bar{\\beta}_i/\\{p_i(1-p_i)\\}$ as a planning diagnostic for how much a given covariate is worth.","An immediate extension the paper leaves open is high-dimensional regression adjustment; its projection-and-residual argument would need a concentration inequality for stratified sampling, and the paper names that as the main technical obstacle."],"forward_implications":["In experiments with many small strata, researchers can use $\\hat{\\tau}_{\\mathrm{ols}}$ with the conservative variance estimator $\\hat{\\sigma}^2_{\\mathrm{ols}}$; the resulting intervals have asymptotic coverage at least nominal and are asymptotically no wider than intervals from $\\hat{\\tau}_{\\mathrm{unadj}}$.","Even when treatment proportions differ across strata, the stratified difference-in-means estimator is consistent and asymptotically normal, so the paper's central limit theorem applies to designs with many small strata; only the variance-reduction claim needs the common-$p$ condition.","With a few large strata and heterogeneous covariate-outcome relationships, the stratum-interacted estimator $\\hat{\\tau}_{\\mathrm{ols,int}}$ is asymptotically at least as efficient as the common-coefficient estimator $\\hat{\\tau}_{\\mathrm{ols}}$.","The conservative variance estimators mean confidence intervals based on the adjusted estimators are asymptotically at least as short as those based on the unadjusted estimator, in line with the simulations showing 4--19% shorter intervals."],"supporting_citations":[{"why":"Supplies the finite-population central limit theorem for stratified sampling that Theorem 1 adapts.","marker":"[4]"},{"why":"Provides the standard stratified-sampling variance facts behind the variance formulas and the common-p comparison.","marker":"[7]"},{"why":"Establishes the asymptotic benchmark for finely stratified experiments that Theorem 2 generalizes.","marker":"[9]"},{"why":"Establishes the paired-experiment benchmark that Theorem 2 also covers.","marker":"[10]"},{"why":"Sets out the stratified randomized experiment and difference-in-means framework the paper starts from.","marker":"[18]"},{"why":"Provides the finite-population central limit theorems and regression-adjustment asymptotics for completely randomized experiments that Theorems 2--4 extend.","marker":"[20]"},{"why":"Gives the treatment-by-covariate interaction regression whose weighted version defines the main adjusted estimator.","marker":"[21]"}],"fun_headline_variants":["Stratified RCTs: adjust without variance penalty","Covariate adjustment: no asymptotic downside in stratified designs","Stratified experiments: regression adjustment matches or improves","Adjustment after stratification: never worse, usually better","Stratified trials: regression adjustment is a safe bet"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline guarantee that regression adjustment never increases variance depends on every stratum's treatment fraction converging to the same number $p$; if treatment fractions stay different across strata, extra terms appear in the variance comparison that the proof does not control.","fun_headline_variants_meta":{"raw":{"variants":["Stratified RCTs: adjust without variance penalty","Covariate adjustment: no asymptotic downside in stratified designs","Stratified experiments: regression adjustment matches or improves","Adjustment after stratification: never worse, usually better","Stratified trials: regression adjustment is a safe bet"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000266,"raw_usage":{"total_tokens":1582,"prompt_tokens":890,"completion_tokens":692,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":614}},"tokens_in":506,"tokens_out":692,"duration_ms":7621,"temperature":1.0,"reasoning_tokens":614,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:06:41.201907+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Consider a two-stratum design with $p_1\\to 0.2$, $p_2\\to 0.8$, a strong covariate in both strata, and bounded potential outcomes, and compute the limit of $N(\\sigma^2_{\\mathrm{ols}}-\\sigma^2_{\\mathrm{unadj}})$ from the paper's formulas; if any such sequence makes that limit positive, the 'never worse' claim fails. A simulation under Conditions 1--6 should also show $\\hat{\\sigma}^2_{\\mathrm{ols}}\\le \\hat{\\sigma}^2_{\\mathrm{unadj}}$ in probability; a configuration where the adjusted interval is systematically wider would delimit the theorem.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the finite-population central limit theorem for stratified sampling that Theorem 1 adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the standard stratified-sampling variance facts behind the variance formulas and the common-p comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the asymptotic benchmark for finely stratified experiments that Theorem 2 generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the paired-experiment benchmark that Theorem 2 also covers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sets out the stratified randomized experiment and difference-in-means framework the paper starts from."},{"cited_title":"and Ding, P","cited_arxiv_id":null,"evidence_quote":"Provides the finite-population central limit theorems and regression-adjustment asymptotics for completely randomized experiments that Theorems 2--4 extend."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the treatment-by-covariate interaction regression whose weighted version defines the main adjusted estimator."}],"review_version":1}