REVIEW 5 major objections 4 minor 1 cited by
Clustering with Potential Multidimensionality: Inference and Practice
T0 review · 5 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper proves when cluster standard errors are needed and offers a shrinkage estimator that is valid and tighter than the standard two-way correction.
desk verdict First design-based multiway cluster-robust inference for M-estimators with a genuinely useful shrinkage estimator, but the simulations that showcase the two-way results use a design the paper itself rules out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are the variance decomposition terms for clustered scores: the individual heteroskedasticity component, the within-cluster correlation component, and the finite-population terms formed from products of score expectations. These extra terms are what make the usual variance estimators conservative in one-way clustering or potentially anticonservative in two-way clustering, and they are not directly identifiable because each unit is observed under only one assignment. The proposed shrinkage estimator estimates a lower bound on those terms by linearly projecting within-cluster sums of scores onto within-cluster sums of fixed attributes; the bound $0 \le \Delta^Z \le \Delta_E + \Delta_{EC}$ is what guarantees conservativeness. Asymptotic normality is obtained by combining standard M-estimation arguments with a one-way clustered central limit theorem and, for two-way dependence, a multiway central limit theorem that requires the variance not to be concentrated in a few clusters.
What would settle it
Simulate the paper's difference-in-means design with constant treatment effects, cluster sampling on one dimension, cluster assignment on a non-nested other dimension, and check whether one-way clustering on the assignment dimension gives empirical coverage at the nominal rate; the paper predicts exact coverage, so persistent undercoverage would contradict the claim.
Extended reading notes
Core claim
The paper's central claim is that, for general M-estimators with finite populations, the variance of the estimator decomposes into an individual (heteroskedasticity) component, within-cluster correlation components, and extra finite-population components that appear because cluster sampling and cluster assignment make even the expectations of scores cross-correlated. The usual one-way cluster-robust estimator converges to the superpopulation variance, which is matrix-wise no smaller than the finite-population variance, so it is conservative. The usual two-way CGM estimator need not be: the paper constructs cases where it is smaller than the true variance and hence anticonservative. To fix this while avoiding the over-conservatism of CGM2, the paper proposes projecting cluster-level score sums onto cluster-level covariates, producing an estimator whose probability limit lies between the true finite-population variance and the CGM2 limit. The resulting adjusted variance estimators are asymptotically conservative and have a smaller upper bound than CGM2. Along the way the paper provides design-based justifications for clustering in difference-in-means, fixed-effects, triple-difference, and two-assignment regressions, and shows that cluster dependence can change the interpretation of fixed-effects estimands, not just their standard errors.
Load-bearing premise
The two-way results rely on a central limit theorem requiring that the variance of the score sum is not concentrated in a few clusters; degenerate dependence that factors as a product of cluster effects is ruled out, and without this condition the variance estimators are not justified.
Editorial extensions
If this is right
- Researchers can choose the clustering level from the sampling and assignment design: cluster only when sampling or assignment is clustered, and use two-way clustering only when the two sources act on different dimensions or one is multiway.
- The standard two-way CGM variance estimator can be anticonservative; practitioners should prefer a conservative estimator such as CGM2 or the paper's adjusted version.
- The proposed adjusted estimator produces standard errors that are often substantially smaller than CGM2 while maintaining correct coverage, so empirical conclusions can become sharper without sacrificing validity.
- In regressions with two assignment variables clustered on different dimensions, one-way clustering on each variable's own dimension can suffice, avoiding unnecessarily large two-way standard errors.
- In triple-differences designs, one-way versus two-way clustering should be chosen according to whether both grouping indicators are stochastic assignment variables or one is a fixed attribute.
Reading between the lines
- The size of the shrinkage gain depends on how well cluster-level covariates predict cluster-level average scores; with no useful covariates the method reduces to the conservative benchmark, and with rich covariates it approaches the finite-population variance.
- The regression-based bound could in principle be applied to three or more clustering dimensions, since it only uses additive one-way cluster objects, provided a suitable multiway central limit theorem exists.
- Under the paper's design-based view, many fixed-effects coefficients in applied work should be interpreted as weighted averages of treatment effects rather than the ATE, with weights depending on cluster and assignment structure.
- Finite-population shrinkage is complementary to cluster bootstrap methods; for the few-clusters case the paper notes that wild cluster bootstrap remains the preferred alternative.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops finite-population asymptotic theory for M-estimators under one- and two-way clustered sampling or assignment. It characterizes when one-way versus multiway clustering is needed, shows that the Cameron-Gelbach-Miller two-way estimator can be anticonservative while the CGM2 estimator is conservative, and proposes covariate-based shrinkage variance estimators that remain conservative but are smaller than CGM2. The framework is applied to difference-in-means, fixed-effects estimands, two assignment variables clustered on different dimensions, triple differences, and an empirical illustration on tenure-clock policies.
Significance. If the results hold, the paper gives practical guidance on when clustering is justified and offers a less conservative valid alternative to CGM2 in design-based settings. Strengths include the general M-estimator setup, an explicit finite-population variance decomposition, clear practical recommendations, and Monte Carlo and empirical evidence. The paper also makes falsifiable predictions about coverage. However, several central proofs are deferred or omitted, and the formal conditions behind the two-way results are hard to verify, so the practical guarantee is not as cleanly established as the text suggests.
major comments (5)
- [Section 2.2.2, Theorem 2.3] The proof of Theorem 2.3 is not provided; the text states that it is "largely analogous to Theorem 2.1" and applies the CLT from Yap (2023) instead of Hansen and Lee (2019). Since Theorem 2.3 underpins Propositions 2.1-2.2 and Theorem 3.3, and Yap (2023) is a working paper by one of the present authors, this is a load-bearing gap. The manuscript should include a full proof or a precise verification that the conditions of Yap (2023) are satisfied under Assumptions 1-5 and Assumption A.3.
- [Appendix C.2, Theorem 3.2] The proof of Theorem 3.2 is omitted with the statement that it is "almost the same as that for Theorem 3.3 with sampling indicators." Theorem 3.2 is needed for Table 1, Cases 2 and 3, and the presence of sampling indicators changes the projection and convergence argument in a non-trivial way. The authors should provide the proof or a detailed statement of the required modifications.
- [Theorem 3.3] The statement of Theorem 3.3 is incomplete: the final sentence asserts that "either ... or ..." two convergence results hold, without specifying the condition that determines which case applies. Moreover, condition (iv) refers to an unstated "variance order condition (C.124) in Appendix C." Since Theorem 3.3 is the formal basis for the adjusted CGM2 estimator in Table 2, these conditions should be stated in the main text, or a simpler sufficient condition should be provided.
- [Assumption 5 and Sections 4, 5.4-5.5] The discussion after Assumption 5 says that a stronger way of stating the assumption rules out balanced two-way designs with one unit per intersection, which invites the reading that Assumption 5 itself rules them out. In fact, Assumption 5 may hold for such designs when two-way dependence makes the variance scale λ_M grow faster than M, but the paper never verifies this for the designs used in the simulations. Since Tables 3-5 are the main evidence for the proposed estimators, the authors should either verify Assumptions 5 and 6 for those DGPs or acknowledge that the simulations fall outside the stated formal domain and provide a separate justification.
- [Proposition 2.2, proof] The key inequality in equation (C.104), namely Δ_ehw,M + ρ_uM Δ_(G∩H),M ≥ ρ_uM ρ_gM ρ_hM (Δ_E,M + Δ_E(G∩H),M), is asserted without proof. It can be justified by writing the left side as ρ_uM times the expected outer product of intersection-level sums plus a nonnegative remainder, but this argument should be stated explicitly, as the conservativeness of CGM2 rests on it.
minor comments (4)
- [Appendix C, proof of Theorem 2.3] There is a typo: "Thereom 2.1" should be "Theorem 2.1."
- [Table 4] The table note misidentifies columns: the first and third data columns report results for X_g, and the second and fourth report results for X_h, not "second and fourth" and "third and fifth."
- [Theorem 3.3] The statement uses the same symbol Δ^Z_CE,M for both dimensions even though equations (42)-(43) define distinct objects Δ^Z_GE,M and Δ^Z_HE,M; this should be clarified.
- [Section 1] The survey claim that 70% of 133 AER articles reported cluster-robust standard errors is stated without describing the article selection or coding procedure; a brief description or reference would be useful.
Circularity Check
No circular reduction found; the shrinkage estimator is derived algebraically, and the only self-citation is an external general CLT, though the simulations sit outside Assumption 5.
full rationale
The paper's central proposal, the covariate-shrunk variance estimator in Table 2, is not fitted to the target variance or imported by definition. Its conservatism is proven by the projection inequality in (C.121)/(C.140), namely that the residual projection is positive semidefinite, and by the CGM2 dominance calculation in (C.103)-(C.105). The claim that adjusted standard errors remain conservative is therefore a theorem rather than a renamed fit or a prediction forced by a fitted parameter. The only load-bearing external theorem is Yap (2023)'s two-way CLT, invoked in the proof of Theorem 2.3 with the sentence: 'we apply the central limit theorem (CLT) from Yap (2023) instead of Hansen and Lee (2019)'. Yap is a coauthor, so this is a self-citation, but it is a parameter-free general CLT with stated assumptions (Assumption 5 and conditions in Yap 2023) and it does not itself contain the present paper's shrinkage estimator or design-based interpretations; under the rubric this is real supporting evidence, not circularity. One internal-validity limitation should be flagged, though it is not a circular step: the paper explicitly states that Assumption 5 'rules out two-way balanced clusters where there is one unit in every intersection: if there are G clusters on both the G and H dimensions, then M = G^2 so 1/M sum_g (M^G_g)^2 = G^3/G^2 = G -> infinity'. Yet the Section 4 simulation uses '50 clusters each on the two dimensions G and H with one unit for every (g,h) cluster pair', and Sections 5.4 and 5.5 use 100x100 one-unit-per-pair designs. Consequently Tables 3-5 cover a setting outside the formal domain of Theorems 2.3 and 3.3. This is a validity and coverage concern, not a circular derivation, and it is the reason I do not assign score 0.
Assumptions & free parameters
free parameters (1)
- choice of covariates z for shrinkage =
university characteristics in the empirical illustration
assumptions (7)
- domain assumption Bernoulli two-step cluster sampling at G and then H dimensions with constant probabilities, plus independent unit sampling (Assumption 1)
- domain assumption Assignment variables are independent across units sharing no cluster on either dimension and may be arbitrarily correlated within clusters (Assumption 2)
- domain assumption Assignment vector is independent of the sampling indicators (Assumption 3)
- domain assumption Cluster sizes satisfy Assumption 4/4' (and Assumption 5 for two-way): no dominant cluster, variance grows properly, eigenvalue conditions
- standard math Regularity conditions in Assumption A.1/A.3: compact parameter space, moment bounds, Lipschitz scores, nonsingular Hessian
- standard math Yap (2023) CLT for triangular arrays with two-way cluster dependence
- domain assumption Potential outcomes and attributes are non-stochastic; randomness comes only from assignment and sampling (footnote 2)
Cite this review
Pith. "Pith review of Clustering with Potential Multidimensionality: Inference and Practice." pith.science (2026). https://pith.science/paper/MHYA7ALE
@misc{pith2026241113372,
author = {Pith},
title = {Pith review of: Clustering with Potential Multidimensionality: Inference and Practice},
year = {2026},
howpublished = {\url{https://pith.science/paper/MHYA7ALE}},
note = {Machine review of arXiv:2411.13372}
}
read the original abstract
We show how clustering standard errors in one or more dimensions can be justified in M-estimation when there is sampling or assignment uncertainty. Since existing procedures for variance estimation are either conservative or invalid, we propose a variance estimator that refines a conservative procedure and remains valid. We then interpret environments where clustering is frequently employed in empirical work from our design-based perspective and provide insights on their estimands and inference procedures.
Forward citations
Cited by 1 Pith paper
-
Variance Estimation with Dependence and Heterogeneous Means
A variance estimator that adds an extra square term to cluster- and time-robust estimators is proposed to prevent over-rejection when means are heterogeneous and clusters are serially correlated.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archive author booktitle chapter collaboration doi edition editor eid eprint howpublished institution journal key month note number numpages organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state...
-
[2]
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
(2017), When should you adjust standard errors for clustering? Tech
Abadie, A., Athey, S., Imbens, G.W., and Wooldridge, J.M. (2017), When should you adjust standard errors for clustering? Tech. rep., NBER Working Paper No. 24003
work page 2017
-
[4]
(2020), Sampling-based versus design-based uncertainty in regression analysis
Abadie, A., Athey, S., Imbens, G.W., and Wooldridge, J.M. (2020), Sampling-based versus design-based uncertainty in regression analysis. Econometrica 88, 265--296
work page 2020
-
[5]
Abadie, A., Athey, S., Imbens, G.W., and Wooldridge, J.M. (2023), When should you adjust standard errors for clustering? The Quarterly Journal of Economics 138(1), 1--35
work page 2023
-
[6]
Antecol, H., Bedard, K., and Stearns, J. (2018), Equal but inequitable: Who benefits from gender-neutral tenure clock stopping policies? American Economic Review 108, 2420--2441
work page 2018
-
[7]
Arkhangelsky, D. and Imbens, G. (2024), Causal models for longitudinal and panel data: A survey. The Econometrics Journal 27(3), C1--C61
work page 2024
-
[8]
Athey, S. and Imbens, G.W. (2022), Design-based analysis in difference-in-differences settings with staggered adoption. Journal of Econometrics 226(1), 62--79
work page 2022
Show all 45 references
-
[9]
(2008), Econometric analysis of panel data
Baltagi, B.H. (2008), Econometric analysis of panel data
2008
-
[10]
(2021), Can policy change culture? government pension plans and traditional kinship practices
Bau, N. (2021), Can policy change culture? government pension plans and traditional kinship practices. American Economic Review 111(6), 1880--1917
2021
-
[11]
(2004), How much should we trust differences-in-differences estimates? Quarterly Journal of Economics 119, 249--275
Bertrand, M., Duflo, E., and Mullainathan, S. (2004), How much should we trust differences-in-differences estimates? Quarterly Journal of Economics 119, 249--275
2004
-
[12]
(2024), Revisiting event study designs: Robust and efficient estimation
Borusyak, K., Jaravel, X., and Spiess, J. (2024), Revisiting event study designs: Robust and efficient estimation. Review of Economic Studies
2024
-
[13]
and Sant’Anna, P.H
Callaway, B. and Sant’Anna, P.H. (2021), Difference-in-differences with multiple time periods. Journal of Econometrics 225(2), 200--230
2021
-
[14]
(2008), Bootstrap-based improvements for inference with clustered errors
Cameron, A.C., Gelbach, J.B., and Miller, D.L. (2008), Bootstrap-based improvements for inference with clustered errors. Review of Economics and Statistics 90, 414--427
2008
-
[15]
(2011), Robust inference with multiway clustering
Cameron, A.C., Gelbach, J.B., and Miller, D.L. (2011), Robust inference with multiway clustering. Journal of Business & Economic Statistics 29(2), 238--249
2011
-
[16]
(2024), Standard errors for two-way clustering with serially correlated time effects
Chiang, H.D., Hansen, B.E., and Sasaki, Y. (2024), Standard errors for two-way clustering with serially correlated time effects. Review of Economics and Statistics pp. 1--40
2024
-
[17]
and Sasaki, Y
Chiang, H.D. and Sasaki, Y. (2023), On using the two-way cluster-robust standard errors. arXiv preprint arXiv:2301.13775
2023 arXiv
-
[18]
(2018), Asymptotic results under multiway clustering
Davezies, L., D'Haultfoeuille, X., and Guyonvarch, Y. (2018), Asymptotic results under multiway clustering. Tech. rep., arXiv preprint arXiv:1807.07925
2018 arXiv
-
[19]
and d'Haultfoeuille, X
de Chaisemartin, C. and d'Haultfoeuille, X. (2024), Difference-in-differences estimators of intertemporal treatment effects. Review of Economics and Statistics pp. 1--45
2024
-
[20]
and d’Haultfoeuille, X
De Chaisemartin, C. and d’Haultfoeuille, X. (2020), Two-way fixed effects estimators with heterogeneous treatment effects. American economic review 110(9), 2964--2996
2020
-
[21]
and Ramirez-Cuellar, J
de Chaisemartin, C. and Ramirez-Cuellar, J. (2024), At what level should one cluster standard errors in paired and small-strata experiments? American Economic Journal: Applied Economics 16(1), 193--212
2024
-
[22]
and White, H
Gallant, A.R. and White, H. (1988), A unified theory of estimation and inference for nonlinear dynamic models. Blackwell
1988
-
[23]
(2024), Two-stage differences in differences
Gardner, J., Thakral, N., T \^o , L.T., and Yap, L. (2024), Two-stage differences in differences
2024
-
[24]
and Lee, S
Hansen, B.E. and Lee, S. (2019), Asymptotic theory for clustered samples. Journal of Econometrics 210, 268--290
2019
-
[25]
(2007), Asymptotic properties of a robust variance matrix estimator for panel data when T is large
Hansen, C.B. (2007), Asymptotic properties of a robust variance matrix estimator for panel data when T is large. Journal of Econometrics 141, 597--620
2007
-
[26]
(1998), Compensating differentials for gender-specific job injury risks
Hersch, J. (1998), Compensating differentials for gender-specific job injury risks. The American Economic Review 88(3), 598--607
1998
-
[27]
and Prucha, I.R
Jenish, N. and Prucha, I.R. (2009), Central limit theorems and uniform laws of large numbers for arrays of random fields. Journal of Econometrics 150(1), 86--98
2009
-
[28]
and Zeger, S.L
Liang, K. and Zeger, S.L. (1986), Longitudinal data analysis using generalized linear models. Biometrika 73, 13--22
1986
-
[29]
(2019), How cluster-robust inference is changing applied econometrics
MacKinnon, J.G. (2019), How cluster-robust inference is changing applied econometrics. Canadian Journal of Economics 52, 851--881
2019
-
[30]
, and Webb, M.D
MacKinnon, J.G., Nielsen, M. ., and Webb, M.D. (2021), Wild bootstrap and asymptotic inference with multiway clustering. Journal of Business & Economic Statistics 39(2), 505--519
2021
-
[31]
and Webb, M.D
MacKinnon, J.G. and Webb, M.D. (2017), Wild bootstrap inference for wildly different cluster sizes. Journal of Applied Econometrics 32, 233--254
2017
-
[32]
and Poyker, M
Marchingiglio, R. and Poyker, M. (2019), The employment effects of gender-specific minimum wage. Tech. rep., Working paper
2019
-
[33]
(2021), Bootstrap with cluster-dependence in two or more dimensions
Menzel, K. (2021), Bootstrap with cluster-dependence in two or more dimensions. Econometrica 89(5), 2143--2188
2021
-
[34]
(1991), Uniform convergence in probability and stochastic equicontinuity
Newey, W.K. (1991), Uniform convergence in probability and stochastic equicontinuity. Econometrica 59, 1161--1167
1991
-
[35]
and McFadden, D
Newey, W.K. and McFadden, D. (1994), Large sample estimation and hypothesis testing. In R.F. Engle and D.L. McFadden (eds.), Handbook of Econometrics, vol. 4, pp. 2111--2245, Elsevier
1994
-
[36]
and M en, J
Olden, A. and M en, J. (2022), The triple difference estimator. The Econometrics Journal 25(3), 531--553
2022
-
[37]
(2012), The treatment effect, the cross difference, and the interaction term in nonlinear ``difference-in-differences'' models
Puhani, P.A. (2012), The treatment effect, the cross difference, and the interaction term in nonlinear ``difference-in-differences'' models. Economics Letters 115, 85--87
2012
-
[38]
(2023), Decomposing triple-differences regression under staggered adoption
Strezhnev, A. (2023), Decomposing triple-differences regression under staggered adoption. Tech. rep., arXiv preprint arXiv:2307.02735
2023 arXiv
-
[39]
and Abraham, S
Sun, L. and Abraham, S. (2021), Estimating dynamic treatment effects in event studies with heterogeneous treatment effects. Journal of Econometrics 225(2), 175--199
2021
-
[40]
(2010), Econometric analysis of cross section and panel data (2nd ed.)
Wooldridge, J.M. (2010), Econometric analysis of cross section and panel data (2nd ed.). MIT press
2010
-
[41]
(2021), Two-way fixed effects, the two-way mundlak regression, and difference-in-differences estimators
Wooldridge, J.M. (2021), Two-way fixed effects, the two-way mundlak regression, and difference-in-differences estimators. Available at SSRN 3906345
2021
-
[42]
(2023), Simple approaches to nonlinear difference-in-differences with panel data
Wooldridge, J.M. (2023), Simple approaches to nonlinear difference-in-differences with panel data. The Econometrics Journal 26(3), C31--C66
2023
-
[43]
(2020), Potential outcomes and finite-population inference for M -estimators
Xu, R. (2020), Potential outcomes and finite-population inference for M -estimators. Econometrics Journal forthcoming
2020
-
[44]
and Wooldridge, J.M
Xu, R. and Wooldridge, J.M. (2022), A design-based approach to spatial correlation. Tech. rep., arXiv preprint arXiv:2211.14354
2022 arXiv
-
[45]
(2023), General conditions for valid inference in multi-way clustering
Yap, L. (2023), General conditions for valid inference in multi-way clustering. Tech. rep., arXiv preprint arXiv:2301.03805
2023 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.