Pith. sign in

REVIEW 3 major objections 5 minor 85 references

Asymptotic Validity and Finite-Sample Properties of Approximate Randomization Tests

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Residual randomization tests are asymptotically valid when the transformation leaves the design's information structure nearly unchanged, and exactly valid in finite samples when each cluster's design is proportional to the full design.

desk verdict The paper's central validity theorem is overclaimed as written: similarity of the randomized Gram matrix cannot substitute for distributional invariance of the errors, and the proof quietly assumes an invariance condition the theorem omits. read the letter →

arxiv 1908.04218 v3 pith:4C7AFS3O submitted 2019-08-12 stat.ME stat.ML

classification stat.MEstat.ML MSC 62G1062J0562G2062G09
keywords residualrandomizationinferentialprimitiveinvarianceexactinferenceclusterederrorswildbootstrapfinite-samplevaliditytests
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Residual randomization tests work by comparing an observed regression statistic with the same statistic evaluated on residuals that have been randomly permuted, sign-flipped, or both. The paper establishes two validity results for these tests: asymptotically, the test controls the rejection probability whenever the random transformation leaves the design's information structure approximately intact, formalized as $(X'X)^{-1}X'GX \xrightarrow{p} bI$; and in finite samples, the test is exact whenever the data can be split into clusters whose design matrices are exact scalar multiples of the full design. From these two results the paper derives validity conditions for exchangeable, heteroskedastic, one-way and two-way clustered, and other error structures, including an exact test for an instance of the Behrens–Fisher problem. A sympathetic reader would care because the method reaches settings—few clusters, skewed or heavy-tailed errors, small samples—where bootstrap variants and cluster-robust asymptotics are known to misbehave.

What carries the argument

The load-bearing object is the inferential primitive $G$, an algebraic group of $n\times n$ matrices encoding the analyst's symmetry assumption about the errors—permutations for exchangeability, diagonal sign matrices for orthant symmetry, cluster-block versions of either, and row-and-column permutations for two-way clustering. The argument runs through the similarity condition $(X'X)^{-1}X'GX \xrightarrow{p} bI$, which states that the random transformation leaves the Fisher information of the design nearly unchanged; under it the gap between the feasible test statistic $t_n(G\hat\varepsilon_o)$ and the idealized $t_n(G\varepsilon)$ is asymptotically negligible, so Theorem 1's proof reduces feasible validity to the classical randomization-test validity result for exact tests. The finite-sample counterpart is the proportionality condition $X_c'X_c = b_c X'X$ of Theorem 2, under which the same gap is exactly zero and the test's size equals the nominal level at every $n$.

What would settle it

Simulate a regression with a small fixed number of clusters (for instance $J=10$), covariates engineered to violate homogeneity so that cluster design matrices $X_c'X_c$ are far from scalar multiples of $X'X$, and within-cluster errors drawn from a skewed distribution so that the sign-symmetry invariance $\varepsilon \overset{d}{=} g\varepsilon$ fails. Applying the cluster sign test of Theorem 6 at the 5% level over many replications, the rate at which the true null is rejected settles the matter: rates around 5% would show the theorem's conditions are not necessary, while rates clearly above 5% confirm that the invariance and homogeneity assumptions are the load-bearing parts.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that the feasible residual randomization test—which replaces unknown errors with restricted OLS residuals $\hat\varepsilon_o$ and compares $T_n = \sqrt{n}(a'\hat\beta - a_0)$ with reference values $t_n(G\hat\varepsilon_o)$ over random transformations $G$—inherits the validity of the idealized test that uses true errors. Theorem 1 proves asymptotic validity, $\limsup_n E(\varphi(y;X)) \le \alpha$, from the similarity condition $(X'X)^{-1}X'GX \xrightarrow{p} bI$, because this makes the gap $t_n(G\hat\varepsilon_o) - t_n(G\varepsilon)$ vanish in probability under the null; the gap equals $-a'(X'X)^{-1}X'GX\sqrt{n}(\hat\beta_o - \beta)$, and the constraint $a'\hat\beta_o = a_0$ forces the limit to zero. Theorem 2 proves finite-sample exactness, $E(\varphi(y;X)) = \alpha$, when a fixed clustering satisfies $X_c'X_c = b_c X'X$ for every cluster and every allowed transformation acts on whole clusters, because then the same gap is identically zero. The paper then shows that various error structures satisfy these conditions: permutation tests are valid even when the similarity condition fails because the permutation bias is asymptotically negligible; cluster sign tests require homogeneity when the number of clusters is fixed, or a no-dominant-cluster condition when it grows; and exact tests are constructible whenever proportional clustering is possible.

Load-bearing premise

The load-bearing premise is that the error distribution is exactly unchanged by every transformation in the chosen group—errors symmetric about zero, or exchangeable within clusters—which must hold in finite samples when the number of clusters is fixed and cannot be verified from the data alone.

Editorial extensions

If this is right

  • For exchangeable errors, residual permutation tests are asymptotically valid even though the similarity condition fails, because the permutation bias is negligible in the limit; this recovers the classical residual-permutation test as a special case.
  • With a fixed number of clusters, cluster sign tests require the homogeneity condition $(X'X)^{-1}X'D_cX \to b_c I$; the result confirms that this requirement is intrinsic to few-cluster inference, not an artifact of the bootstrap.
  • When errors are exchangeable within clusters, centering the covariates removes the homogeneity requirement and permits valid inference with few clusters, at the cost of differencing out the intercept.
  • When the number of clusters grows and no cluster dominates in size ($\sum_c n_c^2/n^2 \to 0$), cluster sign tests and double-invariance tests are valid without homogeneity.
  • Exact finite-sample tests are available for any design admitting a proportional clustering; the paper demonstrates one for the Behrens–Fisher problem with only 30 observations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The similarity condition could be turned into a cheap diagnostic that the paper does not propose: average $(X'X)^{-1}X'GX$ over the sampled transformations and inspect its eigenvalues; if they do not concentrate around one value, the randomization null is likely miscalibrated.
  • Theorem 2's proportionality condition reads like an exact-balance requirement, so automated cluster construction guided by covariate balance could generate exact tests from any design rather than only from natural clusterings.
  • The reflection test for autocorrelated errors conditions on residuals that are only approximately zero, suggesting that Theorem 1's logic could extend to approximate invariances and bring autoregressive settings within reach of the same style of theory.
  • The paper's power findings are empirical, so a natural next step is a formal efficiency comparison against wild-bootstrap refinements to determine whether the flexibility of the transformation group buys power as well as calibration.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript develops a residual randomization method for inference in linear regression. Under a chosen group G of transformations (permutations, sign flips, or clustered variants), the test compares an observed statistic t_n(ε) with statistics computed on transformed restricted residuals t_n(G ε̂_o). The paper claims a general asymptotic validity theorem (Theorem 1) under a similarity condition on the randomized Gram matrix, a finite-sample exactness construction based on cluster proportionality (Theorem 2), and then derives specialized conditions for exchangeable, heteroskedastic, clustered, and two-way clustered errors. It also reports extensive simulations, an exact test for a Behrens–Fisher setup, and empirical extensions to high-dimensional regression and autocorrelated errors.

Significance. If the theorems were correct as stated, the paper would provide a genuinely unified invariance-based alternative to the bootstrap, with the unusual advantage of finite-sample exactness in specially constructed designs. The exact Behrens–Fisher construction and the two-way clustering analysis are original and potentially useful, and the simulation studies are extensive and described in enough detail to be reproducible in spirit. However, the central theorem as stated is false, and several later theorems inherit or repeat the same omission. The reported empirical findings are not enough to compensate for the missing invariance assumptions in the main theoretical claims; the contribution can be meaningful only after substantial correction and reframing.

major comments (3)
  1. [Theorem 1, §2.2] Theorem 1 omits the distributional invariance that the proof actually uses. The proof asserts 'By the invariance assumption, ε d= Grε' at Eq. (13), but no such assumption appears in the theorem, and the similarity condition (10) does not imply it. A concrete failure: take the model yi = β0 + εi with Gs the individual sign-flip group (Eq. 13) and εi iid with mean zero, unit variance, and positive skew. Then (1/n)X'GX = (1/n)Σ Si →p 0, so (10) holds with b = 0, yet t_n(ε) = √n ε̄ converges to a skewed law while t_n(Gε), conditional on ε, converges to N(0,σ²) by the Rademacher CLT. A nominal 5% residual-randomization test therefore over-rejects asymptotically, contradicting the theorem. Theorem 1 must be restated with finite-sample invariance (7), or the asymptotic analogue (11), as an explicit hypothesis, and Remark 2.5 must be revised accordingly. Because later theorems invoke Theorem 1 as a black box, this omission propagates even where those theorems do state a finite-sample invariance.
  2. [Theorem 3, §3.1, Eq. (20)] The proof of Theorem 3 only cancels the bias term for contrasts with a1 = 0, or when the intercept is excluded and covariates are centered. The final step '−a1√n(ε̂o − ε) + oP(1) →p 0' is not justified: with an intercept in X and a1 ≠ 0, the constrained residuals do not have zero sample mean in general; the Lagrange first-order condition gives 1'ε̂o = O_P(√n), so the displayed quantity need not vanish. Since the paper advertises tests of H0,j for all j = 1,...,p, including the intercept, Theorem 3 as stated overclaims. The theorem should explicitly restrict to contrasts with a1 = 0, or provide a different argument that covers intercept-involved hypotheses.
  3. [Theorem 2, §2.3 and §5.2.2] The exactness conclusion of Theorem 2 requires the error law to be invariant under G; the algebraic cluster condition Xc'Xc = bc X'X only guarantees that tn(Grε̂o) = tn(Grε). Without ε d= gε for all g ∈ G, the reference values are not exchangeable with the observed statistic, so the p-value is not uniform under the null. As stated, Theorem 2 is false; for example, with Gs and asymmetric errors, a cluster sign test is not exact even when the cluster-proportionality condition holds. Consequently, the Behrens–Fisher 'exact test' in §5.2.2 must state the sign-symmetry assumption on the errors explicitly. The simulation error models used there are symmetric, but the problem statement only assumes zero mean and variance heterogeneity, so the text as written does not establish the claimed exactness.
minor comments (5)
  1. [Theorem 8 proof, p. 44-45] The proof of Theorem 8 contains the placeholder text 'verylongeqs.' inside two displayed equations, leaving the derivation incomplete as printed; full algebra should be supplied.
  2. [Equation numbering, §2.2 and Appendix B] Equation numbering conflicts between the main text and the proofs: the proof of Theorem 1 labels the display for tn(Grε̂o) − tn(Grε) as Eq. (11), duplicating Remark 2.5's Eq. (11), and the appendix restarts equation numbers at (14). This makes it hard to verify cross-references.
  3. [Section 6] The extensions in Section 6 are explicitly not supported by the theorems: §6.1 proceeds 'without further theoretical investigation' and §6.2 says 'we leave the full proof for future work.' These sections should be labeled as empirical or heuristic; the current text could be read as claiming theoretical validity.
  4. [Theorems 5 and 7] The phrase 'clusters are centered' cannot include the intercept column, whose cluster means are identically 1. The statements should specify that all non-intercept covariates are centered, or that the intercept is excluded from the hypothesis.
  5. [Remark 3.2] Remark 3.2 says the proof of Theorem 4 does not require finite-sample sign symmetry, but the theorem statement explicitly assumes ε d= gε for all g ∈ Gs. The statement and the remark should be aligned, either by weakening the theorem or by qualifying the remark.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the validity theorems reduce to explicit invariance and design assumptions, not to fitted outputs or self-citations.

full rationale

The derivation chain is self-contained. Theorem 1 states asymptotic validity under Assumptions 1(a)-(c) plus the similarity condition (10), and the proof uses the invariance assumption (7) (or its asymptotic version (11)) to assert that the randomization-reference law matches the null law: 'By the invariance assumption, ε d= Grε, variable T(r) also has a limit law...'. That invariance is an explicit input, not a renamed version of the conclusion. The similarity condition (10) is a design condition on X and G that is verified, not fitted; no parameter is estimated from the data being predicted. Theorem 2's exactness follows from the algebraic factorization X_c'X_c = b_c X'X, which makes tn(Gr^εo) − tn(Grε) exactly zero; the cluster construction is a valid recipe, not a fitted calibration. Simulations are Monte Carlo evaluations, not calibrations. The one self-citation (Basse, Feller, and Toulis 2019) appears in a background sentence about randomization tests under interference and is not load-bearing for any theorem. The skeptical critique that Theorem 1 may omit the distributional invariance needed to conclude validity is a correctness concern, not circularity: adding condition (11) as an explicit hypothesis would not make the conclusion equivalent to the premise. Overall, the paper's claims are derived from stated assumptions rather than assumed by construction.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central results are derivations from explicit invariance and moment assumptions; no parameters are fitted to data and no new entities are postulated. The main burden is the finite-sample invariance and the auxiliary cluster conditions.

assumptions (6)
  • domain assumption Linear regression model y = X beta + epsilon with fixed p and n > p (Equation 2).
    The entire method is defined for this model; the distribution of epsilon is otherwise unrestricted beyond invariance.
  • domain assumption Assumption 1(a)-(c): X is non-stochastic, X'X/n is bounded and invertible and converges to positive definite S, and (1/sqrt(n))X'epsilon converges in distribution to some Z.
    Used in every proof to ensure beta hat behaves well and the test statistic has a limit law.
  • domain assumption Finite-sample error invariance: epsilon has the same distribution as g epsilon for all g in the chosen group G (Equation 7).
    Load-bearing for the idealized exact test and for cluster sign tests with finite clusters; it is untestable from the data.
  • domain assumption For clustered results, errors are independent across clusters and Assumptions 1(a)-(b) hold within clusters.
    States the within-cluster moment conditions used in Theorems 5 through 8.
  • domain assumption For two-way clustering, the error structure is additive: epsilon_i = epsilon_r + epsilon_c + epsilon_rc (Equation 19).
    This additive structure supplies the two-way correlation and is assumed, not derived.
  • domain assumption Cluster homogeneity (X'X)^{-1}X'D_cX converges to b_c I (Equation 15), or centered clusters, when invoked by Theorems 5 through 7.
    These conditions are assumed, not derived, and are needed for validity with a fixed number of clusters.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Asymptotic Validity and Finite-Sample Properties of Approximate Randomization Tests." pith.science (2026). https://pith.science/paper/4C7AFS3O

@misc{pith2026190804218,
  author       = {Pith},
  title        = {Pith review of: Asymptotic Validity and Finite-Sample Properties of Approximate Randomization Tests},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4C7AFS3O}},
  note         = {Machine review of arXiv:1908.04218}
}
read the original abstract

Randomization tests rely on simple data transformations and possess an appealing robustness property. In addition to being finite-sample valid if the data distribution is invariant under the transformation, these tests can be asymptotically valid under a suitable studentization of the test statistic, even if the invariance does not hold. However, practical implementation often encounters noisy data, resulting in approximate randomization tests that may not be as robust. In this paper, our key theoretical contribution is a non-asymptotic bound on the discrepancy between the size of an approximate randomization test and the size of the original randomization test using noiseless data. This allows us to derive novel conditions for the validity of approximate randomization tests under data invariances, while being able to leverage existing results based on studentization if the invariance does not hold. We illustrate our theory through several examples, including tests of significance in linear regression. Our theory can explain certain aspects of how randomization tests perform in small samples, addressing limitations of prior theoretical results.

Figures

Figures reproduced from arXiv: 1908.04218 by the authors.

Figure 1
Figure 1. Histogram of p-values for a sequence of tests, [PITH_FULL_IMAGE:figures/full_fig_p017_1.png] view at source ↗
Figure 2
Figure 2. Power of de-sparsified LASSO method (Dezeure et al., 2015) and residual randomization in the high-dimensional simulation of Section 6.1. The different panels indicate how the active parameters are generated. The figure aggregates power data for both s0 = 3 and s0 = 15. for s0 = 3 and s0 = 15, and present the individual plots in Appendix E.1. In the simulations below we test all individual hypotheses H0,j : βj = 0, a… view at source ↗
Figure 2
Figure 2. Power of de-sparsified LASSO method (Dezeure et al., 2015) and residual randomization in the high-dimensional simulation of Section 6.1. The different panels indicate how the active parameters are generated. The figure aggregates power data for both s0 = 3 and s0 = 15. tive, but residual randomization reaches much closer to the nominal 5% level. Further improvements may be possible since the Bonferroni correction us… view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p055_3.png]
Figure 4
Figure 4. Figure 4: Power of de-sparsified LASSO method (Dezeure et al., 2015) and residual randomization in the high-dimensional simulation of Section 6.1 with normal errors. The different panels indicate how the active parameters are generated. The figure aggregates power data only for …
Figure 5
Figure 5. Figure 5: Power of de-sparsified LASSO method (Dezeure et al., 2015) and residual randomization in the high-dimensional simulation of Section 6.1 with normal errors. The different panels indicate how the active parameters are generated. The figure aggregates power data only for …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 72 canonical work pages

  1. [1]

    @esa ( ) , n @biblabelnum##1 ##1

    \@ifclassloaded aguplus natbib The aguplus class already includes natbib coding, so you should not add it explicitly Type <Return> for now, but then later remove the command natbib from the document \@ifclassloaded nlinproc natbib The nlinproc class already includes natbib coding, so you should not add it explicitly Type <Return> for now, but then later r...

  2. [2]

    @stdbsttrue NAT@ctr \@lbibitem[ NAT@ctr ] \@lbibitem[#1]#2 \@ifundefined b@#2\@extra@b@citeb @num @parse #2 [ @natanchorstart #2 \@biblabel @num @natanchorend] @ifcmd#1()()\@nil #2 @lbibitem\@undefined @lbibitem\@lbibitem \@lbibitem[#1]#2 @lbibitem[#1] #2 @ @@label #2 @stdbst @stdbstfalse @stdbst @filesw \@auxout @numberstrue [2] \@ifundefined b@#1\@extra...

  3. [4]

    W., and Wooldridge, J

    Abadie, A., Athey, S., Imbens, G. W., and Wooldridge, J. (2017). When should you adjust standard errors for clustering? Technical report, National Bureau of Economic Research

  4. [5]

    Anderson, M. J. and Robinson, J. (2001). Permutation tests for linear models. Australian & New Zealand Journal of Statistics , 43(1):75--88

  5. [6]

    Andrews, D. (1991). Heteroskedasticity and autocorrelation consistent covariant matrix estimation. Econometrica , 59(3):817--858

  6. [7]

    Angrist, J. D. and Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion . Princeton university press

  7. [8]

    Arellano, M. (1987). Practitioners corner: Computing robust standard errors for within-groups estimators. Oxford bulletin of Economics and Statistics , 49(4):431--434

  8. [9]

    Aronow, P., Cyrus, S., and Assenova, V. A. (2015). Cluster–robust variance estimation for dyadic data. Political Analysis , 23(4):564--577

Show all 85 references
  1. [10]

    Athey, S., Eckles, D., and Imbens, G. W. (2018). Exact p-values for network interference. Journal of the American Statistical Association , 113(521):230--240

  2. [11]

    Basse, G., Feller, A., and Toulis, P. (2019). Randomization tests of causal effects under interference. Biometrika , 106(2):487--494

  3. [12]

    Belloni, A., Chernozhukov, V., and Hansen, C. (2014). Inference on treatment effects after selection among high-dimensional controls. The Review of Economic Studies , 81(2):608--650

  4. [13]

    A., Conley, T

    Bester, C. A., Conley, T. G., and Hansen, C. B. (2011). Inference with dependent data using cluster covariance estimators. Journal of Econometrics , 165(2):137--151

  5. [14]

    Bickel, P. J. and Freedman, D. A. (1983). Bootstrapping regression models with many parameters. Festschrift for Erich L. Lehmann , pages 28--48

  6. [15]

    J., Freedman, D

    Bickel, P. J., Freedman, D. A., et al. (1981). Some asymptotic theory for the bootstrap. The annals of statistics , 9(6):1196--1217

  7. [16]

    C., Gelbach, J

    Cameron, A. C., Gelbach, J. B., and Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. The Review of Economics and Statistics , 90(3):414--427

  8. [17]

    C., Gelbach, J

    Cameron, A. C., Gelbach, J. B., and Miller, D. L. (2011). Robust inference with multiway clustering. Journal of Business & Economic Statistics , 29(2):238--249

  9. [18]

    Cameron, A. C. and Miller, D. L. (2015). A practitioner's guide to cluster-robust inference. Journal of Human Resources , 50(2):317--372

  10. [19]

    A., Romano, J

    Canay, I. A., Romano, J. P., and Shaikh, A. M. (2017). Randomization tests under an approximate symmetry assumption. Econometrica , 85(3):1013--1030

  11. [20]

    small" number of

    Canay, I. A., Santos, A., Shaikh, A. M., et al. (2018). The wild bootstrap with a" small" number of" large" clusters

  12. [21]

    V., Schnepel, K

    Carter, A. V., Schnepel, K. T., and Steigerwald, D. G. (2017). Asymptotic behavior of at-test robust to cluster heterogeneity. Review of Economics and Statistics , 99(4):698--709

  13. [22]

    Casini, A. (2018). Theory of evolutionary spectra for heteroskesdasticity and autocorrelation robust inference in possibly misspecified and nonstationary models

  14. [23]

    D., Frandsen, B

    Cattaneo, M. D., Frandsen, B. R., and Titiunik, R. (2015). Randomization inference in the regression discontinuity design: An application to party advantages in the us senate. Journal of Causal Inference , 3(1):1--24

  15. [24]

    D., Jansson, M., and Newey, W

    Cattaneo, M. D., Jansson, M., and Newey, W. K. (2018). Inference in linear regression models with many covariates and heteroscedasticity. Journal of the American Statistical Association , 113(523):1350--1361

  16. [25]

    Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters

  17. [26]

    Chernozhukov, V., Wuthrich, K., and Zhu, Y. (2017). An exact and robust conformal inference method for counterfactual and synthetic controls. arXiv preprint arXiv:1712.09089

  18. [27]

    and Jewitt, I

    Chesher, A. and Jewitt, I. (1987). The bias of a heteroskedasticity consistent covariance matrix estimator. Econometrica: Journal of the Econometric Society , pages 1217--1222

  19. [28]

    Christakis, N. A. and Fowler, J. H. (2013). Social contagion theory: examining dynamic social networks and human behavior. Statistics in medicine , 32(4):556--577

  20. [29]

    Conley, T. G. (1999). Gmm estimation with cross sectional dependence. Journal of econometrics , 92(1):1--45

  21. [30]

    Conley, T. G. and Taber, C. R. (2011). Inference with difference in differences with a small number of policy changes. The Review of Economics and Statistics , 93(1):113--125

  22. [31]

    D'Adamo, R. (2018). Cluster-robust standard errors for linear regression models with many controls. arXiv preprint arXiv:1806.07314

  23. [32]

    and Flachaire, E

    Davidson, R. and Flachaire, E. (2008). The wild bootstrap, tamed at last. Journal of Econometrics , 146(1):162--169

  24. [33]

    Dekker, D., Krackhardt, D., and Snijders, T. A. (2007). Sensitivity of mrqap tests to collinearity and autocorrelation conditions. Psychometrika , 72(4):563--581

  25. [34]

    Dezeure, R., B \"u hlmann, P., Meier, L., and Meinshausen, N. (2015). High-dimensional inference: Confidence intervals, p-values and r-software hdi. Statistical science , pages 533--558

  26. [35]

    DiCiccio, T. J. and Efron, B. (1996). Bootstrap confidence intervals. Statistical science , pages 189--212

  27. [36]

    DiCiccio, T. J. and Romano, J. P. (1988). A review of bootstrap confidence intervals. Journal of the Royal Statistical Society: Series B (Methodological) , 50(3):338--354

  28. [37]

    Ding, P. et al. (2017). A paradox from randomization-based causal inference. Statistical science , 32(3):331--345

  29. [38]

    Ding, P., Feller, A., and Miratrix, L. (2016). Randomization inference for treatment effect variation. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 78(3):655--671

  30. [39]

    Donald, S. G. and Lang, K. (2007). Inference with difference-in-differences and other panel data. The review of Economics and Statistics , 89(2):221--233

  31. [40]

    Efron, B. (1969). Student's t-test under symmetry conditions. Journal of the American Statistical Association , 64(328):1278--1302

  32. [41]

    Efron, B. (1979). Bootstrap methods: Another look at the jackknife. The Annals of Statistics , pages 1--26

  33. [42]

    Efron, B. (1981). Nonparametric estimates of standard error: the jackknife, the bootstrap and other methods. Biometrika , 68(3):589--599

  34. [43]

    and Tibshirani, R

    Efron, B. and Tibshirani, R. (1986). Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy. Statistical science , pages 54--75

  35. [44]

    Eicker, F. et al. (1963). Asymptotic normality and consistency of the least squares estimators for families of linear regressions. The Annals of Mathematical Statistics , 34(2):447--456

  36. [45]

    and Gubert, F

    Fafchamps, M. and Gubert, F. (2007). The formation of risk sharing networks. Journal of development Economics , 83(2):326--350

  37. [46]

    Fisher, R. A. (1935). The design of experiments . Oliver And Boyd; Edinburgh; London

  38. [47]

    Flachaire, E. (2005). Bootstrapping heteroskedastic regression models: wild bootstrap vs. pairs bootstrap. Computational Statistics & Data Analysis , 49(2):361--376

  39. [48]

    Fraser, D. A. (1968). The Structure of Inference . John Wiley & Sons

  40. [49]

    and Lane, D

    Freedman, D. and Lane, D. (1983a). A nonstochastic interpretation of reported significance levels. Journal of Business & Economic Statistics , 1(4):292--298

  41. [50]

    Freedman, D. A. et al. (1981). Bootstrapping regression models. The Annals of Statistics , 9(6):1218--1228

  42. [51]

    Freedman, D. A. and Lane, D. (1983b). Significance testing in a nonstochastic setting. A festschrift for Erich L. Lehmann , pages 185--208

  43. [52]

    Freedman, D. A. and Peters, S. C. (1984). Bootstrapping an econometric model: Some empirical results. Journal of Business & Economic Statistics , 2(2):150--158

  44. [53]

    Gerber, A. S. and Green, D. P. (2012). Field experiments: Design, analysis, and interpretation . WW Norton

  45. [54]

    Godfrey, L. G. and Orme, C. D. (2001). Significance levels of heteroskedasticity-robust tests for specification and misspecification: some results on the use of wild bootstraps. Unpublished Manuscript

  46. [55]

    Hansen, B. (2018). The exact distribution of the t-ratio with robust and clustered standard errors

  47. [56]

    Hansen, C. B. (2007). Asymptotic properties of a robust variance matrix estimator for panel data when t is large. Journal of Econometrics , 141(2):597--620

  48. [57]

    Horowitz, J. L. (2001). The bootstrap. In Handbook of econometrics , volume 5, pages 3159--3228. Elsevier

  49. [58]

    Hubert, L. (1986). Assignment methods in combinational data analysis , volume 73. CRC Press

  50. [59]

    and Schultz, J

    Hubert, L. and Schultz, J. (1976). Quadratic assignment as a general data analysis strategy. British journal of mathematical and statistical psychology , 29(2):190--241

  51. [60]

    and M \"u ller, U

    Ibragimov, R. and M \"u ller, U. K. (2010). t-statistic based correlation and heterogeneity robust inference. Journal of Business & Economic Statistics , 28(4):453--468

  52. [61]

    and M \"u ller, U

    Ibragimov, R. and M \"u ller, U. K. (2016). Inference with few heterogeneous clusters. Review of Economics and Statistics , 98(1):83--96

  53. [62]

    Imbens, G. W. and Kolesar, M. (2016). Robust standard errors in small samples: Some practical advice. Review of Economics and Statistics , 98(4):701--712

  54. [63]

    Imbens, G. W. and Rubin, D. B. (2015). Causal inference in statistics, social, and biomedical sciences . Cambridge University Press

  55. [64]

    and Montanari, A

    Javanmard, A. and Montanari, A. (2014). Confidence intervals and hypothesis testing for high-dimensional regression. The Journal of Machine Learning Research , 15(1):2869--2909

  56. [65]

    Kennedy, F. E. (1995). Randomization tests in econometrics. Journal of Business & Economic Statistics , 13(1):85--94

  57. [66]

    Krackhardt, D. (1988). Predicting with networks: Nonparametric multiple regression analysis of dyadic data. Social networks , 10(4):359--381

  58. [67]

    Lehmann, E. L. and Romano, J. P. (2005). Testing statistical hypotheses . New York: Springer

  59. [68]

    Liu, R. Y. et al. (1988). Bootstrap procedures under some non-iid models. The Annals of Statistics , 16(4):1696--1708

  60. [69]

    J., and Tibshirani, R

    Lockhart, R., Taylor, J., Tibshirani, R. J., and Tibshirani, R. (2014). A significance test for the lasso. Annals of statistics , 42(2):413

  61. [70]

    MacKinnon, J. G. (2002). Bootstrap inference in econometrics. Canadian Journal of Economics/Revue canadienne d' \'e conomique , 35(4):615--645

  62. [71]

    MacKinnon, J. G. and White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of econometrics , 29(3):305--325

  63. [72]

    Mammen, E. et al. (1993). Bootstrap and wild bootstrap for high dimensional linear models. The annals of statistics , 21(1):255--285

  64. [73]

    Manly, B. F. (2018). Randomization, bootstrap and Monte Carlo methods in biology . Chapman and Hall/CRC

  65. [74]

    McCaffrey, D. F. and Bell, R. M. (2002). Bias reduction in standard errors for linear and generalized linear models with multi-stage samples. In Proceedings of Statistics Canada Symposium , pages 1--10

  66. [75]

    Moulton, B. R. (1986). Random group effects and the precision of regression estimates. Journal of econometrics , 32(3):385--397

  67. [76]

    and Freedman, D

    Peters, S. and Freedman, D. (1984). Some notes on the bootstrap in regression problems. Journal of Business & Economic Statistics , 2(4):406--09

  68. [77]

    Petersen, M. A. (2009). Estimating standard errors in finance panel data sets: Comparing approaches. The Review of Financial Studies , 22(1):435--480

  69. [78]

    Rosenbaum, P. R. (2002). Observational studies. In Observational studies , pages 1--17. Springer

  70. [79]

    Scheff \'e , H. (1970). Practical solutions of the behrens-fisher problem. Journal of the American Statistical Association , 65(332):1501--1508

  71. [80]

    Ter Braak, C. J. (1992). Permutation versus bootstrap significance tests in multiple regression and anova. In Bootstrapping and related techniques , pages 79--85. Springer

  72. [81]

    Welch, B. L. (1951). On the comparison of several mean values: an alternative approach. Biometrika , 38(3-4):330--336

  73. [82]

    White, H. (1984). Asymptotic theory for econometricians . Academic press

  74. [83]

    White, H. et al. (1980). A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity. econometrica , 48(4):817--838

  75. [84]

    Wu, C.-F. J. et al. (1986). Jackknife, bootstrap and other resampling methods in regression analysis. the Annals of Statistics , 14(4):1261--1295

  76. [85]

    Zeileis, A. (2004). Econometric computing with hc and hac covariance matrix estimators

  77. [86]

    and Zhang, S

    Zhang, C.-H. and Zhang, S. S. (2014). Confidence intervals for low dimensional parameters in high dimensional linear models. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 76(1):217--242

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.