{"id":"c36ecac4-da70-4617-a64c-50b846033e2d","arxiv_id":"1908.04218","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Residual randomization tests are shown to be asymptotically valid under a general group-invariance condition and exact in finite samples under a cluster-similarity condition, extending permutation inference to clustered, heteroskedastic, and two-way clustered errors.","lead":"This paper builds a class of regression significance tests that work by randomly transforming the residuals of a fitted model, for example by shuffling them within groups or flipping their signs. It proves conditions under which these tests keep their promised error rate even when the usual assumptions fail, and shows they can give exact results in small samples.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 omits a necessary distributional invariance condition: the similarity condition (10) does not imply that the randomization reference law matches the null law, so asymptotic validity as stated is false for skewed errors under random signs.","rationale":"Good faith summary: the paper's contribution is a general asymptotic-validity theorem for residual randomization under a similarity condition, plus an exact finite-sample construction. The proof of Theorem 1 is valid only if the randomization values tn(Gε) share the null limit law of tn(ε). This is exactly the role of the invariance property, and the proof uses it. Since the theorem statement omits any such condition and Remark 2.5 claims only the weaker (11) is needed, the statement overreaches; (11) is not entailed by (10). The sign-flip/skewed-error example is a minimal concrete failure mode: the design satisfies Assumption 1 and similarity, yet the randomization reference is normal while the null statistic is skewed, so rejection probability exceeds α. I am not claiming the whole framework collapses: the specialized theorems (3)-(8) explicitly assume finite-sample invariance (7), and with that assumption the main argument is coherent. The same omission appears in Theorem 2, which derives exactness from algebraic design conditions without stating the distributional invariance of the errors under G; exactness requires Eq. (7). The Reader's weakest assumption identified Eq. (7) as fragile, which is related but not identical: the sharpest issue is that Theorem 1 claims to avoid it while silently relying on a variant of it. A concrete simulation with centered chi-square errors and individual sign flips would settle whether the theorem as written is false. If it over-rejects, the paper needs a corrected theorem statement and a proof under (11), or an explicit global adoption of (7); the advertised 'similarity alone' result is not established. Therefore the verdict should remain conditional rather than unchanged: acceptance requires correcting Theorem 1's hypotheses and reproducing the argument under a clearly stated invariance condition.","tokens_in":44417,"tokens_out":19937,"duration_ms":233976,"concrete_test":"Simulate the Section 2.1 test with G_s (Eq. 13): n=1000, balanced x_i=±1, β0=β1=0, errors ε_i=(χ²_1−1)/√2 (zero mean, unit variance, positive skew), R=10,000, α=0.05 two-sided, and 10,000 replications using the Appendix D.1 decision rule. The similarity condition (10) holds, but the null statistic is skewed while the randomization law is asymptotically normal. If the empirical rejection rate exceeds 0.05 by more than two simulation standard errors, Theorem 1 as stated is refuted; if it is within 0.05±0.004, the missing invariance is not the culprit and the theorem may be salvageable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorem 1, as stated, concludes asymptotic validity from Assumptions 1 plus the similarity condition (10). But the proof requires tn(Gε) to have the same limiting law as tn(ε): it invokes \"the invariance assumption\" and uses Eq. (13). Remark 2.5 correctly notes that finite-sample invariance (7) is stronger than needed and that asymptotic invariance (11) suffices, yet (11) is not an assumption of Theorem 1 and is not implied by (10). It is in fact necessary. With X including an intercept and x_i=±1, and G the individual sign-flip group G_s, Eq. (10) holds (the Theorem 4 calculation gives (1/n)X'GX→0). Take ε_i iid with mean 0, variance 1, and positive skew. Then tn(ε) converges to a skewed Z, while tn(Gε), conditional on ε, converges to N(0,1) by the Rademacher CLT. The feasible values tn(Gεhat_o) approximate tn(Gε), so the reference distribution is approximately normal and a 5% residual-randomization test over-rejects for right-skewed errors. Thus the central theorem is false exactly in the regime where the paper claims invariance is unnecessary. The theorem must be restated with (11) as an explicit hypothesis, or the paper must commit globally to (7) and drop the stronger claim in Remark 2.5. The same omission affects Theorem 2, whose exactness conclusion also requires the distributional invariance of errors under the chosen group G, not merely the algebraic design condition X_c'X_c∝X'X.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript develops a residual randomization method for inference in linear regression. Under a chosen group G of transformations (permutations, sign flips, or clustered variants), the test compares an observed statistic t_n(ε) with statistics computed on transformed restricted residuals t_n(G ε̂_o). The paper claims a general asymptotic validity theorem (Theorem 1) under a similarity condition on the randomized Gram matrix, a finite-sample exactness construction based on cluster proportionality (Theorem 2), and then derives specialized conditions for exchangeable, heteroskedastic, clustered, and two-way clustered errors. It also reports extensive simulations, an exact test for a Behrens–Fisher setup, and empirical extensions to high-dimensional regression and autocorrelated errors.","tokens_in":44758,"tokens_out":11652,"duration_ms":120986,"significance":"If the theorems were correct as stated, the paper would provide a genuinely unified invariance-based alternative to the bootstrap, with the unusual advantage of finite-sample exactness in specially constructed designs. The exact Behrens–Fisher construction and the two-way clustering analysis are original and potentially useful, and the simulation studies are extensive and described in enough detail to be reproducible in spirit. However, the central theorem as stated is false, and several later theorems inherit or repeat the same omission. The reported empirical findings are not enough to compensate for the missing invariance assumptions in the main theoretical claims; the contribution can be meaningful only after substantial correction and reframing.","major_comments":[{"comment":"Theorem 1 omits the distributional invariance that the proof actually uses. The proof asserts 'By the invariance assumption, ε d= Grε' at Eq. (13), but no such assumption appears in the theorem, and the similarity condition (10) does not imply it. A concrete failure: take the model yi = β0 + εi with Gs the individual sign-flip group (Eq. 13) and εi iid with mean zero, unit variance, and positive skew. Then (1/n)X'GX = (1/n)Σ Si →p 0, so (10) holds with b = 0, yet t_n(ε) = √n ε̄ converges to a skewed law while t_n(Gε), conditional on ε, converges to N(0,σ²) by the Rademacher CLT. A nominal 5% residual-randomization test therefore over-rejects asymptotically, contradicting the theorem. Theorem 1 must be restated with finite-sample invariance (7), or the asymptotic analogue (11), as an explicit hypothesis, and Remark 2.5 must be revised accordingly. Because later theorems invoke Theorem 1 as a black box, this omission propagates even where those theorems do state a finite-sample invariance.","section":"Theorem 1, §2.2"},{"comment":"The proof of Theorem 3 only cancels the bias term for contrasts with a1 = 0, or when the intercept is excluded and covariates are centered. The final step '−a1√n(ε̂o − ε) + oP(1) →p 0' is not justified: with an intercept in X and a1 ≠ 0, the constrained residuals do not have zero sample mean in general; the Lagrange first-order condition gives 1'ε̂o = O_P(√n), so the displayed quantity need not vanish. Since the paper advertises tests of H0,j for all j = 1,...,p, including the intercept, Theorem 3 as stated overclaims. The theorem should explicitly restrict to contrasts with a1 = 0, or provide a different argument that covers intercept-involved hypotheses.","section":"Theorem 3, §3.1, Eq. (20)"},{"comment":"The exactness conclusion of Theorem 2 requires the error law to be invariant under G; the algebraic cluster condition Xc'Xc = bc X'X only guarantees that tn(Grε̂o) = tn(Grε). Without ε d= gε for all g ∈ G, the reference values are not exchangeable with the observed statistic, so the p-value is not uniform under the null. As stated, Theorem 2 is false; for example, with Gs and asymmetric errors, a cluster sign test is not exact even when the cluster-proportionality condition holds. Consequently, the Behrens–Fisher 'exact test' in §5.2.2 must state the sign-symmetry assumption on the errors explicitly. The simulation error models used there are symmetric, but the problem statement only assumes zero mean and variance heterogeneity, so the text as written does not establish the claimed exactness.","section":"Theorem 2, §2.3 and §5.2.2"}],"minor_comments":[{"comment":"The proof of Theorem 8 contains the placeholder text 'verylongeqs.' inside two displayed equations, leaving the derivation incomplete as printed; full algebra should be supplied.","section":"Theorem 8 proof, p. 44-45"},{"comment":"Equation numbering conflicts between the main text and the proofs: the proof of Theorem 1 labels the display for tn(Grε̂o) − tn(Grε) as Eq. (11), duplicating Remark 2.5's Eq. (11), and the appendix restarts equation numbers at (14). This makes it hard to verify cross-references.","section":"Equation numbering, §2.2 and Appendix B"},{"comment":"The extensions in Section 6 are explicitly not supported by the theorems: §6.1 proceeds 'without further theoretical investigation' and §6.2 says 'we leave the full proof for future work.' These sections should be labeled as empirical or heuristic; the current text could be read as claiming theoretical validity.","section":"Section 6"},{"comment":"The phrase 'clusters are centered' cannot include the intercept column, whose cluster means are identically 1. The statements should specify that all non-intercept covariates are centered, or that the intercept is excluded from the hypothesis.","section":"Theorems 5 and 7"},{"comment":"Remark 3.2 says the proof of Theorem 4 does not require finite-sample sign symmetry, but the theorem statement explicitly assumes ε d= gε for all g ∈ Gs. The statement and the remark should be aligned, either by weakening the theorem or by qualifying the remark.","section":"Remark 3.2"}],"recommendation":"major_revision","confidential_remarks":"Major Comment 1 is not a technicality: Theorem 1's conclusion is false under its stated assumptions, and the same omission affects Theorem 2. The paper has a useful framework and credible simulations, so I am not recommending rejection, but the authors must restate the theorems with explicit invariance assumptions, restrict the claims that genuinely go beyond invariance, and correct the proof of Theorem 3 for intercept hypotheses. If a corrected Theorem 1 cannot be provided, the central claim should be withdrawn."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you spend time here. First, the framework is genuinely useful: it puts residual permutation tests and cluster wild bootstrap under one group-theoretic roof, and the exact-test construction in Section 2.3 is a nice idea with a clean algebraic proof. Second, the main theorem as stated is not true. Theorem 1 claims asymptotic validity from Assumption 1 plus the similarity condition (10), but the proof uses 'the invariance assumption' to conclude that tn(Gε) has the same limit law as tn(ε). That assumption is not in the statement, and it is not implied by (10). The stress-test note's specific counterexample with x_i = ±1 does not land—there tn(ε) is asymptotically normal even for skewed errors—but the underlying point survives: condition (10) only gets you tn(Gεhat_o) ≈ tn(Gε); it does nothing to align the reference law with the null law. Theorem 1 needs either asymptotic invariance (11) or finite-sample invariance (7) as an explicit hypothesis. As written, it overclaims.\n\nWhat is good: the paper is honest about prior art (Freedman-Lane, Canay et al.), the simulations are extensive and align with the theory where the theory is correctly stated, and the Behrens-Fisher exact test in Section 5.2.2 is a clever, reproducible construction. No code is shipped, which is a real weakness for a methods paper with this many simulations.\n\nSoft spots, in proportion. Theorem 3's proof of validity under permutation has a bias term that only vanishes when a_1 = 0 or covariates are centered; intercept hypotheses are not covered. Theorem 2 also omits the distributional invariance condition from its assumptions; the algebraic equality is exact, but the 'exact test' conclusion only holds under error invariance. This matters because the Behrens-Fisher example only has symmetry by construction, not by assumption. Theorem 8's second-case proof is sketchy as written. And there is no code.\n\nWho this is for: econometricians and biostatisticians who want an alternative to cluster wild bootstrap with small-sample guarantees, plus methodologists working on randomization inference. It deserves a serious referee: the framework is valuable, and the problems are fixable with a revision that restates Theorem 1 correctly, fixes the intercept case, and releases code. I would send it to review—but I would tell the authors up front that Theorem 1 needs a major rewrite.","headline":"The paper's central validity theorem is overclaimed as written: similarity of the randomized Gram matrix cannot substitute for distributional invariance of the errors, and the proof quietly assumes an invariance condition the theorem omits.","tokens_in":45237,"tokens_out":9136,"would_cite":false,"duration_ms":96363,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62J05","62G20","62G09"],"pacs":[],"model":"deepseek-v4-flash","headline":"Residual randomization tests are asymptotically valid when the transformation leaves the design's information structure nearly unchanged, and exactly valid in finite samples when each cluster's design is proportional to the full design.","keywords":["residual randomization","inferential primitive","invariance","exact inference","clustered errors","wild bootstrap","finite-sample validity","randomization tests"],"falsifier":"Simulate a regression with a small fixed number of clusters (for instance $J=10$), covariates engineered to violate homogeneity so that cluster design matrices $X_c'X_c$ are far from scalar multiples of $X'X$, and within-cluster errors drawn from a skewed distribution so that the sign-symmetry invariance $\\varepsilon \\overset{d}{=} g\\varepsilon$ fails. Applying the cluster sign test of Theorem 6 at the 5% level over many replications, the rate at which the true null is rejected settles the matter: rates around 5% would show the theorem's conditions are not necessary, while rates clearly above 5% confirm that the invariance and homogeneity assumptions are the load-bearing parts.","tokens_in":44239,"feed_emoji":"🎲","tokens_out":16174,"duration_ms":136671,"temperature":0.7,"pith_summary":"Residual randomization tests work by comparing an observed regression statistic with the same statistic evaluated on residuals that have been randomly permuted, sign-flipped, or both. The paper establishes two validity results for these tests: asymptotically, the test controls the rejection probability whenever the random transformation leaves the design's information structure approximately intact, formalized as $(X'X)^{-1}X'GX \\xrightarrow{p} bI$; and in finite samples, the test is exact whenever the data can be split into clusters whose design matrices are exact scalar multiples of the full design. From these two results the paper derives validity conditions for exchangeable, heteroskedastic, one-way and two-way clustered, and other error structures, including an exact test for an instance of the Behrens–Fisher problem. A sympathetic reader would care because the method reaches settings—few clusters, skewed or heavy-tailed errors, small samples—where bootstrap variants and cluster-robust asymptotics are known to misbehave.","feed_headline":"Regression tests by shuffled residuals stay valid in small samples","feed_subtitle":"No normality or bootstrap tuning needed, and some versions are exactly valid at any sample size.","key_machinery":"The load-bearing object is the inferential primitive $G$, an algebraic group of $n\\times n$ matrices encoding the analyst's symmetry assumption about the errors—permutations for exchangeability, diagonal sign matrices for orthant symmetry, cluster-block versions of either, and row-and-column permutations for two-way clustering. The argument runs through the similarity condition $(X'X)^{-1}X'GX \\xrightarrow{p} bI$, which states that the random transformation leaves the Fisher information of the design nearly unchanged; under it the gap between the feasible test statistic $t_n(G\\hat\\varepsilon_o)$ and the idealized $t_n(G\\varepsilon)$ is asymptotically negligible, so Theorem 1's proof reduces feasible validity to the classical randomization-test validity result for exact tests. The finite-sample counterpart is the proportionality condition $X_c'X_c = b_c X'X$ of Theorem 2, under which the same gap is exactly zero and the test's size equals the nominal level at every $n$.","core_discovery":"The central claim, stated on the paper's own terms, is that the feasible residual randomization test—which replaces unknown errors with restricted OLS residuals $\\hat\\varepsilon_o$ and compares $T_n = \\sqrt{n}(a'\\hat\\beta - a_0)$ with reference values $t_n(G\\hat\\varepsilon_o)$ over random transformations $G$—inherits the validity of the idealized test that uses true errors. Theorem 1 proves asymptotic validity, $\\limsup_n E(\\varphi(y;X)) \\le \\alpha$, from the similarity condition $(X'X)^{-1}X'GX \\xrightarrow{p} bI$, because this makes the gap $t_n(G\\hat\\varepsilon_o) - t_n(G\\varepsilon)$ vanish in probability under the null; the gap equals $-a'(X'X)^{-1}X'GX\\sqrt{n}(\\hat\\beta_o - \\beta)$, and the constraint $a'\\hat\\beta_o = a_0$ forces the limit to zero. Theorem 2 proves finite-sample exactness, $E(\\varphi(y;X)) = \\alpha$, when a fixed clustering satisfies $X_c'X_c = b_c X'X$ for every cluster and every allowed transformation acts on whole clusters, because then the same gap is identically zero. The paper then shows that various error structures satisfy these conditions: permutation tests are valid even when the similarity condition fails because the permutation bias is asymptotically negligible; cluster sign tests require homogeneity when the number of clusters is fixed, or a no-dominant-cluster condition when it grows; and exact tests are constructible whenever proportional clustering is possible.","pith_inferences":["The similarity condition could be turned into a cheap diagnostic that the paper does not propose: average $(X'X)^{-1}X'GX$ over the sampled transformations and inspect its eigenvalues; if they do not concentrate around one value, the randomization null is likely miscalibrated.","Theorem 2's proportionality condition reads like an exact-balance requirement, so automated cluster construction guided by covariate balance could generate exact tests from any design rather than only from natural clusterings.","The reflection test for autocorrelated errors conditions on residuals that are only approximately zero, suggesting that Theorem 1's logic could extend to approximate invariances and bring autoregressive settings within reach of the same style of theory.","The paper's power findings are empirical, so a natural next step is a formal efficiency comparison against wild-bootstrap refinements to determine whether the flexibility of the transformation group buys power as well as calibration."],"forward_implications":["For exchangeable errors, residual permutation tests are asymptotically valid even though the similarity condition fails, because the permutation bias is negligible in the limit; this recovers the classical residual-permutation test as a special case.","With a fixed number of clusters, cluster sign tests require the homogeneity condition $(X'X)^{-1}X'D_cX \\to b_c I$; the result confirms that this requirement is intrinsic to few-cluster inference, not an artifact of the bootstrap.","When errors are exchangeable within clusters, centering the covariates removes the homogeneity requirement and permits valid inference with few clusters, at the cost of differencing out the intercept.","When the number of clusters grows and no cluster dominates in size ($\\sum_c n_c^2/n^2 \\to 0$), cluster sign tests and double-invariance tests are valid without homogeneity.","Exact finite-sample tests are available for any design admitting a proportional clustering; the paper demonstrates one for the Behrens–Fisher problem with only 30 observations."],"supporting_citations":[{"why":"Launches the residual-permutation idea that this paper extends from permutations to general transformation groups.","marker":"Freedman and Lane (1983a)"},{"why":"Supplies the exact randomization-test validity theorems that both the asymptotic proof of Theorem 1 and the exactness of Theorem 2 reduce to.","marker":"Lehmann and Romano (2005)"},{"why":"Proves asymptotic validity of the cluster wild bootstrap under the homogeneity condition that Theorem 6 adopts and contrasts with its own conditions.","marker":"Canay et al. (2018)"},{"why":"Introduces cluster wild bootstrap and provides the simulation design used to evaluate cluster sign and permutation tests.","marker":"Cameron et al. (2008)"},{"why":"Defines the wild bootstrap that the sign-symmetry residual randomization test (Theorem 4) is compared with and related to.","marker":"Wu et al. (1986)"},{"why":"Provides the small-sample bias-corrected robust standard errors used as the baseline in the Behrens–Fisher exact-test study.","marker":"Imbens and Kolesar (2016)"},{"why":"Supplies the hormone dataset and bootstrap standard-error results that the Section 5.1 illustration reproduces with residual randomization.","marker":"Efron and Tibshirani (1986)"},{"why":"Provides the evidence for Rademacher (random-sign) weights in the wild bootstrap, matching the sign-symmetry primitive behind Theorems 4 and 6.","marker":"Davidson and Flachaire (2008)"}],"fun_headline_variants":["Shuffled residuals yield exact tests in small samples","Noisy data no longer breaks randomization tests","Finite-sample proof for approximate randomization tests","Residual shuffles keep regression tests valid","Randomization tests stay valid even with noisy errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the error distribution is exactly unchanged by every transformation in the chosen group—errors symmetric about zero, or exchangeable within clusters—which must hold in finite samples when the number of clusters is fixed and cannot be verified from the data alone.","fun_headline_variants_meta":{"raw":{"variants":["Shuffled residuals yield exact tests in small samples","Noisy data no longer breaks randomization tests","Finite-sample proof for approximate randomization tests","Residual shuffles keep regression tests valid","Randomization tests stay valid even with noisy errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1573,"prompt_tokens":1031,"completion_tokens":542,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":473}},"tokens_in":647,"tokens_out":542,"duration_ms":7542,"temperature":1.0,"reasoning_tokens":473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:49:55.252560+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a regression with a small fixed number of clusters (for instance $J=10$), covariates engineered to violate homogeneity so that cluster design matrices $X_c'X_c$ are far from scalar multiples of $X'X$, and within-cluster errors drawn from a skewed distribution so that the sign-symmetry invariance $\\varepsilon \\overset{d}{=} g\\varepsilon$ fails. Applying the cluster sign test of Theorem 6 at the 5% level over many replications, the rate at which the true null is rejected settles the matter: rates around 5% would show the theorem's conditions are not necessary, while rates clearly above 5% confirm that the invariance and homogeneity assumptions are the load-bearing parts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the exact randomization-test validity theorems that both the asymptotic proof of Theorem 1 and the exactness of Theorem 2 reduce to."},{"cited_title":"small\" number of","cited_arxiv_id":null,"evidence_quote":"Proves asymptotic validity of the cluster wild bootstrap under the homogeneity condition that Theorem 6 adopts and contrasts with its own conditions."},{"cited_title":"C., Gelbach, J","cited_arxiv_id":null,"evidence_quote":"Introduces cluster wild bootstrap and provides the simulation design used to evaluate cluster sign and permutation tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the small-sample bias-corrected robust standard errors used as the baseline in the Behrens–Fisher exact-test study."},{"cited_title":"and Flachaire, E","cited_arxiv_id":null,"evidence_quote":"Provides the evidence for Rademacher (random-sign) weights in the wild bootstrap, matching the sign-symmetry primitive behind Theorems 4 and 6."}],"review_version":1}