{"id":"6dc95dc7-5769-4487-bcf6-ffabb917aae3","arxiv_id":"2506.03599","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A mosaic permutation test gives finite-sample valid tests and confidence intervals for panel regressions under local exchangeability, with asymptotic robustness under cluster independence.","lead":"This paper introduces a mosaic permutation test for panel data regressions, which can test whether clusters of observations are truly independent and can build confidence intervals that do not require full cluster independence. The method offers finite-sample guarantees under a local exchangeability condition, and the authors report that on three real datasets many standard methods produce standard errors several times too small.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The mosaic CI formula (4.5) divides by 1−eρ, which is 0 with probability 2^(−M) when all B_m=0, leaving the interval undefined; Theorem 4.1 as stated is therefore not well-defined for finite M.","rationale":"The paper's core results—finite-sample validity of the mosaic permutation test under MI/JI and asymptotic robustness under cluster independence—are supported by careful proofs. Theorems 3.1 and 4.1 are otherwise coherent, and the asymptotic arguments in Sections 3.2 and B are technically detailed. My primary load-bearing concern is the undefined CI formula on the all-zero randomization draw, which is a concrete formal gap for finite M: the display in Eq 4.5 is not an almost-surely well-defined expression, and the theorem as stated does not specify how to handle it. The reader's weakest assumption about MI is a known limitation that the paper explicitly acknowledges, rather than a hidden flaw. A secondary issue is the internally inconsistent 'strictly weaker' claim in Section 1.2, which contradicts the non-nested statement in Section 2; this should be corrected in revision but does not affect the mathematical validity of the theorems. The verdict remains CONDITIONAL because the central construction appears correct after a small amendment to the definition of CI_mosaic.","tokens_in":32159,"tokens_out":29985,"duration_ms":327519,"concrete_test":"Analytical check: for any fixed design with M=2 and ⟨D,D⟩>0, enumerate the four randomization outcomes. On B=(0,0), compute eρ and the ratio (eρ β̂_mosaic − \\tilde β)/(1−eρ); both numerator and denominator are zero, so the endpoint in Eq 4.5 is undefined. This event has probability 1/4, so the interval in Eq 4.5 is not almost-surely defined. The check settles the concern without simulation; it also indicates the needed fix: define CI_mosaic directly as {b: pval(b) ≥ α} and state Eq 4.5 as an equivalent expression valid when 1−eρ ≠ 0.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorem 4.1 claims P(β* ∈ CI_mosaic) ≥ 1−α for the interval defined in Eq 4.5. In that formula, eρ = ⟨D,\\tilde D⟩/√(⟨D,D⟩⟨\\tilde D,\\tilde D⟩). Under the randomization with B_1,...,B_M i.i.d. Bernoulli(1/2), the event {B_1=...=B_M=0} has probability 2^(−M) > 0 for every finite M. On this event, \\tilde D = D and \\tilde ε = ε̂, so eρ = 1 and \\tilde β = β̂_mosaic. The ratio (eρ β̂_mosaic − \\tilde β)/(1−eρ) is therefore 0/0, and Eq 4.5 has no well-defined endpoints. The proof of Theorem 4.1 in Appendix A.2 establishes that the CI equals the inversion of the mosaic permutation test; that inversion is well-defined, so the gap is in the statement of the formula, not the proof. But the theorem is stated for Eq 4.5, and an undefined randomization outcome at positive probability invalidates the formal statement for finite M. If the all-zero draws are discarded, the randomization distribution is no longer the full group and the exchangeability argument (Step 2 of Theorem 3.1's proof) may fail to yield exact coverage for small M. The paper does not specify a convention for this event.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a mosaic permutation test for panel-data regressions. The test is designed to assess the cluster-independence assumption and, by inversion, to yield confidence intervals for a regression coefficient. The main theoretical claims are: (i) finite-sample Type I error control of the test under a marginal invariance assumption (Assumption MI) plus cluster independence, equivalently under Assumption JI; (ii) asymptotic Type I error control under cluster independence without invariance for a quadratic test statistic; (iii) finite-sample confidence intervals under Assumption JI; and (iv) asymptotic confidence-interval validity under cluster independence and a Lyapunov condition. The paper also presents empirical diagnostics on three economic datasets, arguing that standard cluster-robust methods undercover while mosaic intervals are better calibrated.","tokens_in":32492,"tokens_out":30938,"duration_ms":358781,"significance":"The paper addresses an important problem: cluster-robust inference in panel data is known to undercover when the clustering or independence assumptions fail, and the proposed mosaic construction offers a conceptually new route by replacing or supplementing cluster independence with local exchangeability. The finite-sample test Theorem 3.1 is proved from explicit assumptions, the asymptotic robustness results are nontrivial, and the paper ships public code and detailed proofs. If the confidence-interval formulation is repaired, the paper would be a valuable contribution to the growing literature on randomization-based inference for panel data.","major_comments":[{"comment":"The confidence interval formula is undefined with positive probability. When B_1 = ... = B_M = 0, we have D̃ = D and ε̃ = ε̂, so eρ = 1 and both the numerator and denominator of (eρ β̂_mosaic − β̃)/(1 − eρ) vanish; this event has probability 2^{−M} for every finite M. The proof in Appendix A.2 establishes exact coverage for the interval obtained by inverting the mosaic permutation test, but the algebraic equivalence used there fails on this event. Simply conditioning on {B ≠ 0} or assigning an arbitrary value to the ratio does not obviously preserve validity: with M = 1 the conditional randomization distribution is a single point, giving a zero-length interval, and the exchangeability-based inequality P(S > Q_{1−α/2}(S̃)) ≤ α/2 is false for point-mass conditional distributions (e.g., X = 0/1 with Y = 1 − X). The theorem should be restated for the inversion-based interval, or an exact tie-breaking convention that keeps the full-group randomization distribution must be supplied.","section":"Section 4, Eq. (4.5), Theorem 4.1"},{"comment":"The p-value defined in Eq. (3.3) is one-sided, but Section 4 inverts it to obtain a two-sided confidence interval. No two-sided p-value is defined, and the proof of Theorem 4.1 bounds the two tail events using Q_{α/2} and Q_{1−α/2} without showing that these events are equivalent to {p_val(b) < α}. The text should either define the two-sided p-value explicitly (e.g., 2 min{p^+, p^−} suitably capped) and prove that CI_mosaic is its inversion, or present Theorem 4.1 as a direct statement about the two-sided randomization interval rather than about an inverted p-value.","section":"Sections 3.1 and 4"}],"minor_comments":[{"comment":"The second term in the displayed union uses Q_{1−α/2} where Q_{α/2} is required; the printed formula repeats the same quantile in both terms.","section":"Appendix A.2, Eq. (A.13)"},{"comment":"The statement that the method is valid under assumptions that are strictly weaker than Assumption S is too strong: Assumptions JI and S are non-nested, and the paper's guarantees are finite-sample under JI and asymptotic under S. A formulation such as valid under a different set of assumptions that are arguably weaker in practical panel settings would be more accurate.","section":"Section 1.2"},{"comment":"The diagnostic plots report averages over random data splits without error bars or confidence bands; since the split is random, the plots would be more informative with a measure of sampling variability.","section":"Section 5, Figures 1 and 2"},{"comment":"The lemmas assume that the augmented design X̃ has full column rank so that (X̃^T X̃)^{-1} exists; this rank condition should be stated explicitly in Section 3.1 when cluster-by-cluster residuals are defined.","section":"Appendix A, Lemmas A.1 and A.2"},{"comment":"Remark 7 defines σ̂_mosaic as the standard deviation of the same ratio that appears in CI_mosaic; this is subject to the same all-zero event issue as Eq. (4.5).","section":"Section 4, Remark 7"}],"recommendation":"major_revision","confidential_remarks":"The main issue is confined to the statement of Theorem 4.1 and its relationship to the p-value; the underlying idea of inverting the mosaic test is sound and the finite-sample test theorem appears correct. I would be willing to accept after a revision that restates the CI theorem in terms of the inversion-based interval and clarifies the two-sided construction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline is that this is a real contribution: a finite-sample valid permutation test for cluster independence in panel regressions, and a confidence interval construction that only assumes joint invariance (e.g. local exchangeability) plus an asymptotic fallback under the usual cluster-independence assumptions. The technical core is careful. The mosaic residual construction, where residuals are estimated cluster-by-cluster on an invariance-augmented design, is exactly what makes the joint invariance hold for estimated residuals. The method-of-moments result in Proposition 3.1—showing the randomization distribution matches the unconditional law of a degenerate U-statistic—is clever and, as far as I can tell, correct. The empirical diagnostics on three datasets are a good idea, and the finding that standard errors from standard methods can be several times too small is important.\n\nNow the soft spots, in order of size.\n\nThe formula for the mosaic CI, Eq (4.5), divides by 1−eρ. When all randomization draws B_m are 0—probability 2^{−M}—both numerator and denominator are 0. The theorem as stated is therefore not well-defined for finite M. The proof, however, shows the CI equals the inversion of the mosaic permutation test, and that inversion is well-defined. So this is a statement-level bug, not a broken result, but the authors need to state a convention (e.g. define the CI via inversion, or condition on eρ<1) and adjust Theorem 4.1.\n\nThe phrase “strictly weaker assumptions” is not quite right. Joint invariance and cluster independence are non-nested. Finite-sample validity requires joint invariance; asymptotic validity requires cluster independence plus regularity conditions (Assumptions 3.3, 3.4, 4.1) that are plausible but not verified. I’d soften that language.\n\nThe empirical diagnostics in Figures 1 and 2 are averaged point estimates with no error bars or confidence bands. The exclusion of three outliers in Figure 7 is disclosed, but it would be better to show them. These are minor but worth fixing.\n\nOverall, the central finite-sample claims are proven carefully, and the asymptotic results are a substantial technical achievement. The issues I found are addressable and do not undermine the main contribution. I’d send this to a serious referee. The stress-test note about the undefined CI formula should be forwarded to the authors; it’s exactly the kind of edge case that needs a one-line fix.","headline":"A solid, careful paper that delivers finite-sample valid tests and CIs for panel data under local exchangeability; the CI formula has a small undefined-event bug but the inversion proof fixes it.","tokens_in":33001,"tokens_out":3285,"would_cite":true,"duration_ms":33380,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G09","62G10","62G20","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"A permutation test on carefully built residuals gives exact finite-sample p-values and confidence intervals for panel regressions, with an asymptotic fallback to classical assumptions.","keywords":["panel data","permutation test","cluster independence","local exchangeability","confidence intervals","Type I error control","linear regression"],"falsifier":"Simulate a panel with 100 units, $T=10$ time points, and $M=20$ clusters, with errors drawn as independent Gaussians and a single covariate; run the mosaic permutation test 20,000 times and estimate the Type I error at $\\alpha=0.05$. If the empirical rejection rate exceeds 5 percent by more than Monte Carlo error, Theorem 3.1 would be refuted.","tokens_in":31972,"feed_emoji":"🔀","tokens_out":10475,"duration_ms":107517,"temperature":0.7,"pith_summary":"The paper sets out to weaken the cluster-independence assumption that most panel-data inference relies on. It introduces a mosaic permutation test that, under a mild invariance condition such as local exchangeability of adjacent time points, controls the Type I error rate exactly in finite samples for any test statistic; inverting the test yields confidence intervals with finite-sample coverage under a joint invariance condition. If the invariance condition is false, the same procedures are still asymptotically valid under the classical setting of independent clusters with a growing number of clusters. The practical stakes are shown on three published panel datasets, where standard cluster-robust methods produce variance estimates up to five times too small while the mosaic intervals behave more consistently.","feed_headline":"Swap time points to get valid panel-data confidence intervals","feed_subtitle":"Exact finite-sample inference under local exchangeability, and classical guarantees when it fails","key_machinery":"The carrying mechanism is the mosaic residual estimate: residuals are computed cluster-by-cluster from an invariance-augmented regression that adds transformed covariates $XP$ to the design, where $P$ is a symmetric idempotent matrix encoding the invariance (for local exchangeability, the permutation that swaps adjacent time points). Each cluster's residual matrix is then independently multiplied by $P$ with probability $1/2$. Because the augmented projection satisfies $PHP=H$, the estimated residuals obey the same joint invariance as the true errors, which makes the finite-sample p-value and confidence interval exact. For the asymptotic robustness results, the key technical input is a finite-sample moment bound showing that all moments of the randomization distribution track the unconditional moments of the normalized quadratic test statistic at rate $1/M$, so the test recovers the non-universal limiting law of a degenerate U-statistic without studentizing or estimating its variance.","core_discovery":"The central claim is that finite-sample valid inference for linear panel regressions can be built from residuals that inherit the invariances of the true errors, without assuming cluster independence. Under the null that clusters are independent and errors satisfy marginal invariance (Assumption MI), the mosaic p-value satisfies $\\mathbb{P}(\\mathrm{pval}\\le\\alpha)\\le\\alpha$ for all $\\alpha\\in(0,1)$ and every choice of test statistic (Theorem 3.1). Under joint invariance (Assumption JI), the inverted interval satisfies $\\mathbb{P}(\\beta^\\star\\in\\mathrm{CI}_{\\mathrm{mosaic}})\\ge 1-\\alpha$ in finite samples (Theorem 4.1). Under only mean-zero, cluster-independent errors with a growing number of clusters and regularity conditions, the test and the interval recover asymptotic validity even when the invariance assumptions fail (Theorems 3.2 and 4.2).","pith_inferences":["The non-nested relationship between the two assumptions suggests a practical workflow: run both mosaic and cluster-robust intervals and use the mosaic test as a diagnostic; disagreement would indicate which assumption is driving the conclusions.","The moment-matching argument appears transferable to other degenerate U-statistic settings with few clusters, since it does not require $T$ to grow and avoids consistent variance estimation; a natural test would be to apply it to other between-cluster quadratic forms.","Because local exchangeability allows arbitrary cross-sectional dependence within clusters and arbitrary unit heterogeneity, it may be especially plausible for spatial panels where the disturbance distribution drifts slowly across time; the fold-based overlap diagnostics could be used to test this directly.","The width penalty observed in experiments (typically 1.1 to 1.5 times wider, with rare larger outliers) suggests a concrete goal for follow-up work: choose alternative invariances or adaptive cluster merging to reduce variance inflation while preserving exact coverage."],"forward_implications":["Researchers can test the cluster-independence null with exact finite-sample Type I error control, using essentially any test statistic, so hidden cross-cluster dependence can be diagnosed before cluster-robust standard errors are trusted.","Confidence intervals for a regression coefficient can be reported with exact finite-sample coverage under local exchangeability, without requiring clusters to be independent or the number of clusters to be large.","In the classical regime of independent clusters and a growing number of clusters, the same procedures remain asymptotically valid even when the invariance assumption fails, so the new method keeps the old guarantee as a fallback.","The fold-splitting diagnostics give a concrete, method-agnostic check of whether a reported standard error is trustworthy on a given dataset; in the three datasets studied, classical and cluster-robust intervals undercover while mosaic intervals track the theoretical overlap probability.","Because local exchangeability and cluster independence are non-nested, mosaic inference is valid under a genuinely different assumption, not merely a relaxation of the standard one."],"supporting_citations":[{"why":"Supplies the original mosaic permutation test and the invariance-augmented residual construction that this paper adapts to panel regressions.","marker":"Spector et al. (2024)"},{"why":"Establishes finite-sample regression permutation tests under exchangeable errors, the line of work this method extends to milder invariances.","marker":"Lei and Bickel (2020)"},{"why":"Provides the augmented-regression permutation test for linear models whose design augmentation the mosaic residuals build on.","marker":"Guan (2024)"},{"why":"Supplies the quantile-convergence lemma used to turn moment matching into asymptotic Type I error control.","marker":"Romano and Shaikh (2012)"},{"why":"Shows the limiting law of the quadratic test statistic is non-universal, motivating the moment-based recovery of the randomization distribution.","marker":"Bhattacharya et al. (2022)"},{"why":"Provides the WCR bootstrap and cluster-robust guidance used as baselines in the empirical comparison.","marker":"MacKinnon et al. (2023)"},{"why":"Contributes the cluster-level estimation idea and the fold-split diagnostic that the empirical evaluation adapts.","marker":"Ibragimov and Müller (2016)"}],"fun_headline_variants":["Mosaic permutation test relaxes cluster-independence for panel data","Finite-sample panel inference without full cluster independence","Panel confidence intervals valid under local exchangeability","Test hidden cluster dependencies with mosaic permutation","Reliable panel inference despite nearby cluster dependencies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The finite-sample guarantees stand or fall on the condition that the errors within each cluster are distributionally unchanged when adjacent time periods are swapped (or another stated invariance holds); if errors trend or autocorrelate, those guarantees lapse and only the asymptotic, many-independent-clusters version remains.","fun_headline_variants_meta":{"raw":{"variants":["Mosaic permutation test relaxes cluster-independence for panel data","Finite-sample panel inference without full cluster independence","Panel confidence intervals valid under local exchangeability","Test hidden cluster dependencies with mosaic permutation","Reliable panel inference despite nearby cluster dependencies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000574,"raw_usage":{"total_tokens":2707,"prompt_tokens":939,"completion_tokens":1768,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":1698}},"tokens_in":555,"tokens_out":1768,"duration_ms":12326,"temperature":1.0,"reasoning_tokens":1698,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:58:39.458597+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a panel with 100 units, $T=10$ time points, and $M=20$ clusters, with errors drawn as independent Gaussians and a single covariate; run the mosaic permutation test 20,000 times and estimate the Type I error at $\\alpha=0.05$. If the empirical rejection rate exceeds 5 percent by more than Monte Carlo error, Theorem 3.1 would be refuted.","supporting_citations":[{"cited_title":"F., Hastie, T., Kahn, R","cited_arxiv_id":null,"evidence_quote":"Supplies the original mosaic permutation test and the invariance-augmented residual construction that this paper adapts to panel regressions."},{"cited_title":"and Bickel, P","cited_arxiv_id":null,"evidence_quote":"Establishes finite-sample regression permutation tests under exchangeable errors, the line of work this method extends to milder invariances."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the augmented-regression permutation test for linear models whose design augmentation the mosaic residuals build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the quantile-convergence lemma used to turn moment matching into asymptotic Type I error control."},{"cited_title":"B., Das, S., Mukherjee, S., and Mukherjee, S","cited_arxiv_id":null,"evidence_quote":"Shows the limiting law of the quadratic test statistic is non-universal, motivating the moment-based recovery of the randomization distribution."},{"cited_title":"G., Ørregaard Nielsen, M., and Webb, M","cited_arxiv_id":null,"evidence_quote":"Provides the WCR bootstrap and cluster-robust guidance used as baselines in the empirical comparison."},{"cited_title":"and Müller, U","cited_arxiv_id":null,"evidence_quote":"Contributes the cluster-level estimation idea and the fold-split diagnostic that the empirical evaluation adapts."}],"review_version":1}