REVIEW 2 major objections 5 minor 50 references
Mosaic inference on panel data
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A permutation test on carefully built residuals gives exact finite-sample p-values and confidence intervals for panel regressions, with an asymptotic fallback to classical assumptions.
desk verdict A solid, careful paper that delivers finite-sample valid tests and CIs for panel data under local exchangeability; the CI formula has a small undefined-event bug but the inversion proof fixes it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the mosaic residual estimate: residuals are computed cluster-by-cluster from an invariance-augmented regression that adds transformed covariates $XP$ to the design, where $P$ is a symmetric idempotent matrix encoding the invariance (for local exchangeability, the permutation that swaps adjacent time points). Each cluster's residual matrix is then independently multiplied by $P$ with probability $1/2$. Because the augmented projection satisfies $PHP=H$, the estimated residuals obey the same joint invariance as the true errors, which makes the finite-sample p-value and confidence interval exact. For the asymptotic robustness results, the key technical input is a finite-sample moment bound showing that all moments of the randomization distribution track the unconditional moments of the normalized quadratic test statistic at rate $1/M$, so the test recovers the non-universal limiting law of a degenerate U-statistic without studentizing or estimating its variance.
What would settle it
Simulate a panel with 100 units, $T=10$ time points, and $M=20$ clusters, with errors drawn as independent Gaussians and a single covariate; run the mosaic permutation test 20,000 times and estimate the Type I error at $\alpha=0.05$. If the empirical rejection rate exceeds 5 percent by more than Monte Carlo error, Theorem 3.1 would be refuted.
Extended reading notes
Core claim
The central claim is that finite-sample valid inference for linear panel regressions can be built from residuals that inherit the invariances of the true errors, without assuming cluster independence. Under the null that clusters are independent and errors satisfy marginal invariance (Assumption MI), the mosaic p-value satisfies $\mathbb{P}(\mathrm{pval}\le\alpha)\le\alpha$ for all $\alpha\in(0,1)$ and every choice of test statistic (Theorem 3.1). Under joint invariance (Assumption JI), the inverted interval satisfies $\mathbb{P}(\beta^\star\in\mathrm{CI}_{\mathrm{mosaic}})\ge 1-\alpha$ in finite samples (Theorem 4.1). Under only mean-zero, cluster-independent errors with a growing number of clusters and regularity conditions, the test and the interval recover asymptotic validity even when the invariance assumptions fail (Theorems 3.2 and 4.2).
Load-bearing premise
The finite-sample guarantees stand or fall on the condition that the errors within each cluster are distributionally unchanged when adjacent time periods are swapped (or another stated invariance holds); if errors trend or autocorrelate, those guarantees lapse and only the asymptotic, many-independent-clusters version remains.
Editorial extensions
If this is right
- Researchers can test the cluster-independence null with exact finite-sample Type I error control, using essentially any test statistic, so hidden cross-cluster dependence can be diagnosed before cluster-robust standard errors are trusted.
- Confidence intervals for a regression coefficient can be reported with exact finite-sample coverage under local exchangeability, without requiring clusters to be independent or the number of clusters to be large.
- In the classical regime of independent clusters and a growing number of clusters, the same procedures remain asymptotically valid even when the invariance assumption fails, so the new method keeps the old guarantee as a fallback.
- The fold-splitting diagnostics give a concrete, method-agnostic check of whether a reported standard error is trustworthy on a given dataset; in the three datasets studied, classical and cluster-robust intervals undercover while mosaic intervals track the theoretical overlap probability.
- Because local exchangeability and cluster independence are non-nested, mosaic inference is valid under a genuinely different assumption, not merely a relaxation of the standard one.
Reading between the lines
- The non-nested relationship between the two assumptions suggests a practical workflow: run both mosaic and cluster-robust intervals and use the mosaic test as a diagnostic; disagreement would indicate which assumption is driving the conclusions.
- The moment-matching argument appears transferable to other degenerate U-statistic settings with few clusters, since it does not require $T$ to grow and avoids consistent variance estimation; a natural test would be to apply it to other between-cluster quadratic forms.
- Because local exchangeability allows arbitrary cross-sectional dependence within clusters and arbitrary unit heterogeneity, it may be especially plausible for spatial panels where the disturbance distribution drifts slowly across time; the fold-based overlap diagnostics could be used to test this directly.
- The width penalty observed in experiments (typically 1.1 to 1.5 times wider, with rare larger outliers) suggests a concrete goal for follow-up work: choose alternative invariances or adaptive cluster merging to reduce variance inflation while preserving exact coverage.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a mosaic permutation test for panel-data regressions. The test is designed to assess the cluster-independence assumption and, by inversion, to yield confidence intervals for a regression coefficient. The main theoretical claims are: (i) finite-sample Type I error control of the test under a marginal invariance assumption (Assumption MI) plus cluster independence, equivalently under Assumption JI; (ii) asymptotic Type I error control under cluster independence without invariance for a quadratic test statistic; (iii) finite-sample confidence intervals under Assumption JI; and (iv) asymptotic confidence-interval validity under cluster independence and a Lyapunov condition. The paper also presents empirical diagnostics on three economic datasets, arguing that standard cluster-robust methods undercover while mosaic intervals are better calibrated.
Significance. The paper addresses an important problem: cluster-robust inference in panel data is known to undercover when the clustering or independence assumptions fail, and the proposed mosaic construction offers a conceptually new route by replacing or supplementing cluster independence with local exchangeability. The finite-sample test Theorem 3.1 is proved from explicit assumptions, the asymptotic robustness results are nontrivial, and the paper ships public code and detailed proofs. If the confidence-interval formulation is repaired, the paper would be a valuable contribution to the growing literature on randomization-based inference for panel data.
major comments (2)
- [Section 4, Eq. (4.5), Theorem 4.1] The confidence interval formula is undefined with positive probability. When B_1 = ... = B_M = 0, we have D̃ = D and ε̃ = ε̂, so eρ = 1 and both the numerator and denominator of (eρ β̂_mosaic − β̃)/(1 − eρ) vanish; this event has probability 2^{−M} for every finite M. The proof in Appendix A.2 establishes exact coverage for the interval obtained by inverting the mosaic permutation test, but the algebraic equivalence used there fails on this event. Simply conditioning on {B ≠ 0} or assigning an arbitrary value to the ratio does not obviously preserve validity: with M = 1 the conditional randomization distribution is a single point, giving a zero-length interval, and the exchangeability-based inequality P(S > Q_{1−α/2}(S̃)) ≤ α/2 is false for point-mass conditional distributions (e.g., X = 0/1 with Y = 1 − X). The theorem should be restated for the inversion-based interval, or an exact tie-breaking convention that keeps the full-group randomization distribution must be supplied.
- [Sections 3.1 and 4] The p-value defined in Eq. (3.3) is one-sided, but Section 4 inverts it to obtain a two-sided confidence interval. No two-sided p-value is defined, and the proof of Theorem 4.1 bounds the two tail events using Q_{α/2} and Q_{1−α/2} without showing that these events are equivalent to {p_val(b) < α}. The text should either define the two-sided p-value explicitly (e.g., 2 min{p^+, p^−} suitably capped) and prove that CI_mosaic is its inversion, or present Theorem 4.1 as a direct statement about the two-sided randomization interval rather than about an inverted p-value.
minor comments (5)
- [Appendix A.2, Eq. (A.13)] The second term in the displayed union uses Q_{1−α/2} where Q_{α/2} is required; the printed formula repeats the same quantile in both terms.
- [Section 1.2] The statement that the method is valid under assumptions that are strictly weaker than Assumption S is too strong: Assumptions JI and S are non-nested, and the paper's guarantees are finite-sample under JI and asymptotic under S. A formulation such as valid under a different set of assumptions that are arguably weaker in practical panel settings would be more accurate.
- [Section 5, Figures 1 and 2] The diagnostic plots report averages over random data splits without error bars or confidence bands; since the split is random, the plots would be more informative with a measure of sampling variability.
- [Appendix A, Lemmas A.1 and A.2] The lemmas assume that the augmented design X̃ has full column rank so that (X̃^T X̃)^{-1} exists; this rank condition should be stated explicitly in Section 3.1 when cluster-by-cluster residuals are defined.
- [Section 4, Remark 7] Remark 7 defines σ̂_mosaic as the standard deviation of the same ratio that appears in CI_mosaic; this is subject to the same all-zero event issue as Eq. (4.5).
Circularity Check
No circular derivation: finite-sample and asymptotic guarantees are proved from stated assumptions, not fitted to data.
full rationale
The central claims are proved from the stated null and invariance assumptions. Theorem 3.1 is established by Lemma A.2, which shows mosaic residuals inherit Assumption JI from the true errors, and by the standard randomization-test rank property that then yields P(pval <= alpha) <= alpha; the only delegation to Spector et al. (2024) is the subgroup-exchangeability argument, a parameter-free textbook fact and not a fitted input or an assumption of the panel conclusion. Theorem 4.1 follows algebraically from Lemmas A.4 and A.5 and the same randomization validity; no estimated parameter is inserted into the coverage statement. The asymptotic results are new proofs: Proposition 3.1 bounds the gap between randomization and unconditional moments, and Theorem 4.2 uses a Lyapunov CLT; neither reduces to a fitted quantity. The paper's own stated limitations in Section 6 (cluster-by-cluster estimation, power against joint-invariance alternatives, multi-way clustering, nonlinear models, and small-M asymptotics) are scope caveats, not circular steps. A separate non-circular correctness caveat is that Eq. (4.5) is undefined when all randomization bits B_m = 0 (probability 2^{-M}), since then eρ = 1 and the ratio is 0/0; this concerns well-definedness of the formula, not circularity of the derivation.
Assumptions & free parameters
free parameters (3)
- Cluster partition C_1,...,C_M
- Transformation matrix P
- Test statistic weights s_ij
assumptions (8)
- domain assumption Assumption MI: marginal invariance of errors within each cluster under a known symmetric idempotent transformation P.
- domain assumption Assumption JI: joint invariance of cluster error blocks under independent transformations.
- domain assumption Cluster independence of error blocks under H0 and for asymptotic robustness.
- domain assumption Mean-zero errors (Assumption 3.1).
- domain assumption Sub-Gaussian errors (Assumption 3.2).
- domain assumption sigma_delta^2 bounded away from zero (Assumption 3.3).
- domain assumption Anti-concentration of the test statistic (Assumption 3.4).
- domain assumption Lyapunov condition for cluster-level contributions (Assumption 4.1).
Cite this review
Pith. "Pith review of Mosaic inference on panel data." pith.science (2026). https://pith.science/paper/A6VE6D6E
@misc{pith2026250603599,
author = {Pith},
title = {Pith review of: Mosaic inference on panel data},
year = {2026},
howpublished = {\url{https://pith.science/paper/A6VE6D6E}},
note = {Machine review of arXiv:2506.03599}
}
read the original abstract
Analysis of panel data via linear regression is widespread across disciplines. To perform statistical inference, such analyses typically assume that clusters of observations are jointly independent. For example, one might assume that observations in New York are independent of observations in New Jersey. Are such assumptions plausible? Might there be hidden dependencies between nearby clusters? This paper introduces a mosaic permutation test that can (i) test the cluster-independence assumption and (ii) produce confidence intervals for linear models without assuming the full cluster-independence assumption. The key idea behind our method is to apply a permutation test to carefully constructed residual estimates that obey the same invariances as the true errors. As a result, our method yields finite-sample valid inferences under a mild "local exchangeability" condition. This condition differs from the typical cluster-independence assumption, as neither assumption implies the other. Furthermore, our method is asymptotically valid under cluster-independence (with no exchangeability assumptions). Together, these results show our method is valid under assumptions that are arguably weaker than the assumptions underlying many classical methods. In experiments on well-studied datasets from the literature, we find that many existing methods produce variance estimates that are up to five times too small, whereas mosaic methods produce reliable results. We implement our methods in the python package mosaicperm.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
A. Colin Cameron, J. B. G. and Miller, D. L. (2011). Robust inference with multiway clustering. Journal of Business & Economic Statistics , 29(2):238--249
work page 2011
-
[2]
Abadie, A., Athey, S., Imbens, G. W., and Wooldridge, J. (2020). Sampling-based vs.\ design-based uncertainty in regression analysis. Econometrica , 88:265--296
work page 2020
-
[3]
Abadie, A., Athey, S., Imbens, G. W., and Wooldridge, J. M. (2022). When Should You Adjust Standard Errors for Clustering?* . The Quarterly Journal of Economics , 138(1):1--35
work page 2022
-
[4]
Arkhangelsky, D. and Imbens, G. (2024). Causal models for longitudinal and panel data: a survey. The Econometrics Journal , 27(3):C1--C61
work page 2024
-
[5]
Bester, C. A., Conley, T. G., and Hansen, C. B. (2011). Inference with dependent data using cluster covariance estimators. Journal of Econometrics , 165:137--151
work page 2011
-
[6]
B., Das, S., Mukherjee, S., and Mukherjee, S
Bhattacharya, B. B., Das, S., Mukherjee, S., and Mukherjee, S. (2022). Asymptotic distribution of random quadratic forms
work page 2022
-
[7]
Billingsley, P. (1995). Probability and Measure . Wiley Series in Probability and Mathematical Statistics. John Wiley & Sons, New York, 3rd edition
work page 1995
-
[8]
Cai, Y. (2021). A modified randomization test for the level of clustering. Ar X iv e-prints 2105.01008, Northwestern University
work page Pith review arXiv 2021
Show all 50 references
-
[9]
A., Kim, D., and Shaikh, A
Cai, Y., Canay, I. A., Kim, D., and Shaikh, A. M. (2021). On the implementation of approximate randomization tests in linear models with a small number of clusters. Ar X iv e-prints 2102.09058v2, Northwestern University
2021 arXiv
-
[10]
C., Gelbach, J
Cameron, A. C., Gelbach, J. B., and Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. Review of Economics and Statistics , 90:414--427
2008
-
[11]
Cameron, A. C. and Miller, D. L. (2015). A practitioner's guide to cluster-robust inference. Journal of Human Resources , 50:317--372
2015
-
[12]
Canay, I. A. and Kamat, V. (2017). Approximate permutation tests and induced order statistics in the regression discontinuity design. The Review of Economic Studies , 85(3):1577--1608
2017
-
[13]
A., Romano, J
Canay, I. A., Romano, J. P., and Shaikh, A. M. (2017). Randomization tests under an approximate symmetry assumption. Econometrica , 85:1013--1030
2017
-
[14]
A., Santos, A., and Shaikh, A
Canay, I. A., Santos, A., and Shaikh, A. (2021). The wild bootstrap with a `small' number of `large' clusters. Review of Economics and Statistics , 103:346--363
2021
-
[15]
and Chen, S
Cao, Y. and Chen, S. (2022). Rebel on the canal: Disrupted trade access and social conflict in china, 1650–1911. American Economic Review , 112(5):1555–90
2022
-
[16]
Chen, J. (2025). Potential weights and implicit causal designs in linear regression
2025
-
[17]
and Romano, J
Chung, E. and Romano, J. P. (2013). Exact and asymptotically robust permutation tests . The Annals of Statistics , 41(2):484 -- 507
2013
-
[18]
and Romano, J
Chung, E. and Romano, J. P. (2016a). Asymptotically valid and exact permutation tests based on two-sample u-statistics. Journal of Statistical Planning and Inference , 168:97--105
2016
-
[19]
and Romano, J
Chung, E. and Romano, J. P. (2016b). Multivariate and multiple permutation tests. Journal of Econometrics , 193(1):76--91
2016
-
[20]
and D'Haultfœuille, X
de Chaisemartin, C. and D'Haultfœuille, X. (2020). Two-way fixed effects estimators with heterogeneous treatment effects. American Economic Review , 110(9):2964–96
2020
-
[21]
and Tuvaandorj, P
D'Haultfœuille, X. and Tuvaandorj, P. (2024). A robust permutation test for subvector inference in linear regressions. Quantitative Economics , 15(1):27--87
2024
-
[22]
A., Mac\-Kinnon, J
Djogbenou, A. A., Mac\-Kinnon, J. G., and Nielsen, M. . (2019). Asymptotic theory and wild bootstrap inference with clustered errors. Journal of Econometrics , 212:393--412
2019
-
[23]
M., and Sinkinson, M
Gentzkow, M., Shapiro, J. M., and Sinkinson, M. (2011). The effect of newspaper entry and exit on electoral politics. American Economic Review , 101(7):2980–3018
2011
-
[24]
Guan, L. (2024). A conformal test of linear models via permutation-augmented regressions . The Annals of Statistics , 52(5):2059 -- 2080
2024
-
[25]
Hagemann, A. (2019). Permutation inference with a finite number of heterogeneous clusters. Ar X iv e-prints 1907.01049, University of Michigan
2019 arXiv
-
[26]
Hansen, B. E. and Lee, S. (2019). Asymptotic theory for clustered samples. Journal of Econometrics , 210:268--290
2019
-
[27]
Hansen, C. B. (2007). Asymptotic properties of a robust variance matrix estimator for panel data when T is large. Journal of Econometrics , 141:597--620
2007
-
[28]
and Spamann, H
Hu, A. and Spamann, H. (2020). Inference with cluster imbalance: T he case of state corporate laws. Discussion paper, Harvard Law School
2020
-
[29]
and M\"uller, U
Ibragimov, R. and M\"uller, U. K. (2010). t -statistic based correlation and heterogeneity robust inference. Journal of Business & Economic Statistics , 28:453--468
2010
-
[30]
and Müller, U
Ibragimov, R. and Müller, U. K. (2016). Inference with Few Heterogeneous Clusters . The Review of Economics and Statistics , 98(1):83--96
2016
-
[31]
and Pauls, T
Janssen, A. and Pauls, T. (2003). How do bootstrap and permutation tests work? The Annals of Statistics , 31(3):768 -- 806
2003
-
[32]
and Pauls, T
Janssen, A. and Pauls, T. (2005). A monte carlo comparison of studentized bootstrap and permutation tests for heteroscedastic two-sample problems. Computational Statistics , 20(3):369--383
2005
-
[33]
and Bickel, P
Lei, L. and Bickel, P. J. (2020). An assumption-free exact test for fixed-design linear models with exchangeable errors . Biometrika , 108(2):397--412
2020
-
[34]
and Zeger, S
Liang, K.-Y. and Zeger, S. L. (1986). Longitudinal data analysis using generalized linear models. Biometrika , 73:13--22
1986
-
[35]
G., Nielsen, M
Mac\-Kinnon, J. G., Nielsen, M. ., and Webb, M. D. (2020). Testing for the appropriate level of clustering in linear regression models. QED Working Paper 1428, Queen's University
2020
-
[36]
G., Nielsen, M
Mac\-Kinnon, J. G., Nielsen, M. ., and Webb, M. D. (2022). Fast and reliable jackknife and bootstrap methods for cluster-robust inference. QED Working Paper 1485, Queen's University
2022
-
[37]
Mac\-Kinnon, J. G. and Webb, M. D. (2018). The wild bootstrap for few (treated) clusters. Econometrics Journal , 21:114--135
2018
-
[38]
G., Ørregaard Nielsen, M., and Webb, M
MacKinnon, J. G., Ørregaard Nielsen, M., and Webb, M. D. (2023). Cluster-robust inference: A guide to empirical practice. Journal of Econometrics , 232(2):272--299
2023
-
[39]
Menzel, K. (2021). Bootstrap with cluster-dependence in two or more dimensions. Econometrica , 89(5):2143--2188
2021
-
[40]
Neuhaus, G. (1993). Conditional Rank Tests for the Two-Sample Problem Under Random Censorship . The Annals of Statistics , 21(4):1760 -- 1779
1993
-
[41]
Pouliot, G. A. (2025). An exact t-test
2025
-
[42]
and Roth, J
Rambachan, A. and Roth, J. (2024). Design-based uncertainty for quasi-experiments
2024
-
[43]
Romano, J. P. (1990). On the behavior of randomization tests without a group invariance assumption. Journal of the American Statistical Association , 85(411):686--692
1990
-
[44]
Romano, J. P. and Shaikh, A. M. (2012). On the uniform asymptotic validity of subsampling and the bootstrap . The Annals of Statistics , 40(6):2798 -- 2822
2012
-
[45]
F., Hastie, T., Kahn, R
Spector, A., Barber, R. F., Hastie, T., Kahn, R. N., and Candès, E. (2024). The mosaic permutation test: an exact and nonparametric goodness-of-fit test for factor models
2024
-
[46]
and Verbeek, M
Vella, F. and Verbeek, M. (1998). Whose wages do unions raise? a dynamic model of unionism and wage rate determination for young men. Journal of Applied Econometrics , 13(2):163--183
1998
-
[47]
Wainwright, M. J. (2019). High-Dimensional Statistics: A Non-Asymptotic Viewpoint . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press
2019
-
[48]
Wen, K., Wang, T., and Wang, Y. (2022). Residual permutation test for high-dimensional regression coefficient testing
2022
-
[49]
White, H. (1984). Asymptotic Theory for Econometricians . Academic Press, San Diego
1984
-
[50]
Zajkowski, K. (2020). Bounds on tail probabilities for quadratic forms in dependent sub-gaussian random variables. Statistics & Probability Letters , 167:108898
2020
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.