REVIEW 3 major objections 5 minor 1 cited by
Uniform inference in linear mixed models
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proves that score-based chi-squared confidence regions for random-effect variances and covariances have uniformly correct coverage in linear mixed models, including on the singular boundary where variances are zero or…
desk verdict This paper makes a genuine advance in uniform inference for variance parameters in linear mixed models, including crossed random effects; the main risk is the deferred proof of the key lemma. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the ratio $a(v,\psi)$, defined as the spectral norm divided by the Frobenius norm of $A(v,\psi)=\Sigma(\psi)^{-1/2}\Sigma(v)\Sigma(\psi)^{-1/2}$; it measures how far a direction $v$ in parameter space is from being an identifiable, well-conditioned direction. Lemma 4, built on a normal approximation for quadratic forms, turns smallness of $a(v,\psi)$ into an explicit bound on $|g(t;v,\psi)-\phi(t)|$, the difference between the density of $v^T W^S(\psi)$ and the standard normal density. Lemma 5 converts the pointwise bound into a uniform one, and Theorem 2 converts that into asymptotic normality of every linear combination of the standardized score. For crossed random effects the analysis uses the fact that $\Sigma(\psi)$ is a linear combination of fixed projection matrices, so the eigenvectors of $\Sigma$ do not depend on $\psi$; this permits a spectral argument that yields the $\psi$-free bound involving $n_{\min}$ and $r$.
What would settle it
Simulate a crossed random-effects model with $r=3$ fixed factors but with one factor's size kept at $n_{\min}=1$ while the others grow, compute the restricted score confidence region at parameters where a random-effect variance is zero, and check whether the actual coverage error stays as small as the paper's uniform bound predicts; if it does, the paper's sufficient condition $n_{\min}\ge3$ is shown to be unnecessary, and if it does not, the condition is doing real work. A more direct check is to evaluate $a(v,\psi)$ numerically for such a design and verify whether the Theorem 4 bound is violated.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the distribution of any linear combination of the standardized score $W^S(\psi) = I(\psi)^{-1/2}S(\psi)$ is close to standard normal whenever the quantity $a(v,\psi)=\|A(v,\psi)\|/\|A(v,\psi)\|_F$ is small, where $A(v,\psi)=\Sigma(\psi)^{-1/2}\Sigma(v)\Sigma(\psi)^{-1/2}$ and $\Sigma(v)=\sum_j v_j \partial\Sigma/\partial\psi_j$. Theorem 2 formalizes this as asymptotic normality, uniformly over unit vectors $v_n$, and extends it to unknown fixed effects $\beta$ when $X$ has full column rank. In independent-cluster settings Theorem 3 gives the bound $a(v,\psi)\le c_3 m^{-1/2}(1+\|\psi_{-r}\|/\psi_r)$, so the approximation improves as the number of clusters grows even at the boundary. In crossed random effects settings Theorem 4 gives a bound depending only on the smallest factor size $n_{\min}$ and the number of factors $r$—not on $\psi$—so Corollary 4 yields $\sup_{\psi}|P_\psi\{\psi\in \tilde{C}^S_n(\alpha)\}-(1-\alpha)|\to0$ over the whole parameter set when $r\ge3$ is fixed and $n_{\min}\to\infty$. The paper thus claims that score-based chi-squared confidence regions solve the boundary problem for these models, and the experiments support that claim for the restricted likelihood version.
Load-bearing premise
The whole argument leans on the design matrices giving every random effect enough independent replication: in the clustered setting each cluster's within-cluster design matrix must be sufficiently well conditioned, and in the crossed settings every crossed factor must have at least two (or three for the restricted score) levels, with the smallest factor size growing for the asymptotic results; if that fails, the key ratio $a(v,\psi)$ need not be small and the uniform normality bound does not follow.
Editorial extensions
If this is right
- If the central claim is right, the chi-squared quantile gives confidence regions for random-effect variances and covariances with asymptotically correct uniform coverage, including on the boundary where a variance is zero or a correlation is ±1.
- For independent clusters, coverage improves at rate roughly $m^{-1/2}$ in the number of clusters, so the finite-sample bounds can be used to say how many clusters are enough before trusting a chi-squared region.
- For crossed random effects, the effective sample size for variance-component inference is governed by the smallest factor size $n_{\min}$, not the total number of observations, and uniform coverage holds over the entire parameter set when $r$ is fixed and $n_{\min}\to\infty$.
- The results also hold when the number of random-effect parameters grows with the sample size, and the restricted-likelihood version remains valid with a growing number of fixed-effect predictors.
- Wald and likelihood-ratio regions do not have this property, so score-based regions are the default choice when boundary parameters are a live possibility.
Reading between the lines
- We infer that the same $a(v,\psi)$ ratio could serve as a practical diagnostic: before reporting a score-based confidence region, one could compute its maximum over the boundary of the fitted region and flag cases where the normal approximation is likely poor.
- We infer that the mechanism is not fundamentally tied to normality: the score is a quadratic form, and the paper's remarks suggest the bound could be reworked for non-normal responses by replacing the normality-based moment conditions with estimates of third and fourth moments, perhaps via resampling.
- We infer that allocating experimental units across all crossed factors matters more than total sample size for variance-component inference, a point with direct implications for designing multi-factor studies.
- We infer that extending this style of bound to generalized linear mixed models would hit the same difficulty the paper identifies: without the linearity and joint normality, the spectral decomposition of $\Sigma$ is lost, so the crossed-random-effects proof would need a different route.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops finite-sample and asymptotic distributional approximations for score-based inference in Gaussian linear mixed models, with emphasis on uniform coverage near the boundary of the random-effects covariance parameter space, where variances may vanish or correlations may approach ±1. The main engine is Lemma 4, a finite-sample bound on the sup-norm distance between the density of a standardized linear combination of the score and the standard normal density, controlled by the ratio a(v, psi) of spectral to Frobenius norms of A(v, psi). Lemma 5 converts this into a uniform bound; Theorem 2 derives asymptotic normality of arbitrary linear combinations of the standardized score under sup_v a_n(v, psi_n) -> 0, including with unknown fixed effects. Theorems 3 and 4 give design-specific bounds for independent clusters and crossed random effects, and Corollaries 2-4 translate these into asymptotically correct uniform coverage for chi-square confidence regions, including at the boundary. Simulations illustrate near-nominal coverage for score regions and failures of Wald and likelihood-ratio regions in the same settings.
Significance. If the deferred proofs are correct, this is a substantial contribution: it provides a general treatment of uniform inference for covariance parameters at the singular boundary of the parameter set, with explicit finite-sample bounds and clear sufficient conditions. The design-specific bounds (m^{-1/2} for independent clusters, functions of n_min for crossed effects) are informative and lead to practically implementable confidence regions. The paper's strengths include explicit constants, a clean separation between the generic approximation and design-specific analyses, treatment of crossed random effects with diverging dimensions, reproducible simulation code, and falsifiable predictions about coverage. The main reservations are verification risks rather than demonstrated errors: the central normal approximation is imported from a co-authored prior paper whose proof is not in the reviewed material, and the extension to unknown beta requires a nontrivial argument that the main text does not supply.
major comments (3)
- [Section 4.1, Lemma 4] Lemma 4 is the sole finite-sample engine of the paper, but its proof is deferred to a supplement that was not included in the reviewed material, and the lemma is attributed to a normal approximation for quadratic forms due to Zhang et al. (2025), which was developed for a single variance component. In the present setting A(v, psi) is generally indefinite because the H_j contain off-diagonal ones, so the cited result must handle indefinite quadratic forms under the condition a(tilde v, psi)^2 < 1/8. Please state the precise conditions and either prove Lemma 4 or reproduce the relevant theorem from the cited paper; without this, the main results cannot be independently checked.
- [Section 4.1, Theorem 2, second assertion] The claim that u_n^T W_n^S(theta_n) converges to N(0,1) when beta is unknown is not immediate from Lemma 5. The score for beta is exactly normal, but it is uncorrelated with, not generally independent of, the quadratic score for psi; in a Gaussian vector a linear form and a quadratic form can have zero covariance yet be dependent. The proof must show how this dependence is controlled in the linear combination u_n^T W_n^S(theta_n). The remark that the result is 'suggested, but not implied' by the marginal normality of the beta-score confirms that a gap remains in the main text.
- [Sections 4.2 and 4.3, finite-sample applicability] The finite-sample bound (15) is valid only when a(tilde v, psi)^2 < 1/8. The design-specific bounds in Theorem 3 (18) and Theorem 4 are often larger than this threshold for small m or small n_min; for example, Theorem 4 with n_min = 3 gives tilde a^2 <= 1. The paper should state explicitly that the finite-sample statements are conditional on the right-hand side being below 1/8, and that the asymptotic corollaries rely on this threshold being eventually satisfied.
minor comments (5)
- [Section 2, Lemma 2] In the displayed definition of q_{1-alpha}(psi), the inequality appears reversed: it should be min{t : F(t) >= 1-alpha}, not min{t : F(1-alpha) >= t}.
- [Section 4.3] In the display defining the crossed-effects projections, P2 = Z^{(2)}Z^{(2)T}/n_2 should be divided by n_1, not n_2, to make P2 a projection matrix; the subsequent expression for Sigma confirms this.
- [Theorem 2] The notation P_n subset of R^{r_n x r_n} in the theorem statement should be P_n subset of R^{r_n}; the parameter vector psi_n is in R^{r_n}, not in a matrix space.
- [Section 4.3] The symbol P_j is used both for the n_j x n_j projection matrix 1_{n_j}1_{n_j}^T/n_j and for the n x n Kronecker-product projection; please use distinct symbols or otherwise clarify to avoid confusion.
- [Section 5] The phrase 'fine-sample coverage probabilities' should read 'finite-sample coverage probabilities'.
Circularity Check
No circularity: uniform-coverage claims follow from stated quantitative bounds; the co-authored Lemma 4 is an external, published inequality, not a fitted input or the target conclusion.
full rationale
The derivation chain is self-contained in structure. The score statistic is defined from the log-likelihood and Fisher information (Section 3). Lemma 4 supplies a finite-sample density bound for a linear combination of the standardized score, citing the normal-approximation result of Zhang et al. (2025); this is a published inequality whose stated assumptions do not include the paper's conclusions, and the bound is then stated explicitly in Eq. (15). Lemma 5 turns a uniform bound on a(v, psi) into a uniform density bound. Theorems 3 and 4 prove that this key quantity is small under explicit design-replication conditions (bounded eigenvalues of Z_i^T Z_i, n_min growing), rather than assuming it. Corollaries 2-4 then combine these bounds with Lemmas 1-2 to obtain uniform asymptotic coverage. No parameter is fitted to data and renamed as a prediction, no target result is assumed inside the proof of the smallness bound, and no uniqueness theorem is imported from the authors' prior work. The self-citations present are not load-bearing: Lemma 1 is compared to Lemma 2.5 of Ekvall and Bottai (2022) but is stated and proved independently, and Lemma 4 is attributed to Zhang et al. (2025), a co-authored paper, but is an external mathematical result with explicit assumptions and no dependence on the current paper's conclusions. The only caveat is that Lemma 4's proof is deferred to the supplement, which is a verification risk rather than circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The response follows the Gaussian linear mixed model (1) with correctly specified mean and covariance, and Psi is a linear function of psi whose element-level derivatives are 0/1 matrices.
- standard math Lemma 4, the normal approximation bound for quadratic forms borrowed from Zhang et al. (2025), is correct as stated, including the constants 0.14 and 0.29.
- domain assumption The parameter set P = {psi : Psi(psi) >= 0, psi_r > 0} with Psi positive semidefinite is the correct constraint set, and the Fisher information is positive definite under the stated identifiability condition (12).
Cite this review
Pith. "Pith review of Uniform inference in linear mixed models." pith.science (2026). https://pith.science/paper/H566U7MG
@misc{pith2026250719633,
author = {Pith},
title = {Pith review of: Uniform inference in linear mixed models},
year = {2026},
howpublished = {\url{https://pith.science/paper/H566U7MG}},
note = {Machine review of arXiv:2507.19633}
}
abstract
We provide finite-sample distribution approximations, that are uniform in the parameter, for inference in linear mixed models. Focus is on variances and covariances of random effects in cases where existing theory fails because their covariance matrix is nearly or exactly singular, and hence near or at the boundary of the parameter set. Quantitative bounds on the differences between the standard normal density and those of linear combinations of the score function enable, for example, the assessment of sufficient sample size. The bounds also lead to useful asymptotic theory in settings where both the number of parameters and the number of random effects grow with the sample size. We consider models with independent clusters and ones with a possibly diverging number of crossed random effects, which are notoriously complicated. Simulations indicate the theory leads to practically relevant methods. In particular, the studied confidence regions, which are straightforward to implement, have near-nominal coverage in finite samples even when some random effects have variances near or equal to zero, or correlations near or equal to $\pm 1$.
Figures
Forward citations
Cited by 1 Pith paper
-
Universal inference for variance components
A randomized split likelihood ratio test provides finite-sample valid inference for variance components at the boundary, including heritability near 1, with computational shortcuts for structured models.
Reference graph
Works this paper leans on
-
[1]
Bates, D. et al. (2005). Fitting linear mixed models in R . R news , 5(1):27--30
work page 2005
-
[2]
Battey, H. S. and McCullagh, P. (2024). An anomaly arising in the analysis of processes with more than one source of variability. Biometrika , 111(2):677--689
work page 2024
-
[3]
Bottai, M. (2003). Confidence regions when the Fisher information is zero. Biometrika , 90(1):73--84
work page 2003
-
[4]
Chesher, A. (1984). Testing for neglected heterogeneity. Econometrica , 52(4):865--872
work page 1984
-
[5]
Cox, D. R. and Hinkley, D. V. (2000). Theoretical Statistics . Chapman & Hall/CRC, Boca Raton
work page 2000
-
[6]
Crainiceanu, C. M. and Ruppert, D. (2004). Likelihood ratio tests in linear mixed models with one variance component. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 66(1):165--185
work page 2004
-
[7]
Ekvall, K. O. and Bottai, M. (2022). Confidence regions near singular information and boundary points with applications to mixed models. The Annals of Statistics , 50(3):1806--1832
work page 2022
-
[8]
Ekvall, K. O. and Jones, G. L. (2020). Consistent maximum likelihood estimation using subsets with applications to multivariate mixed models. The Annals of Statistics , 48(2):932--952
work page 2020
Show all 30 references
-
[9]
Geyer, C. J. (1994). On the Asymptotics of Constrained M-Estimation . The Annals of Statistics , 22(4):1993--2010
1994
-
[10]
M., K \"u chenhoff, H., and Peters, A
Greven, S., Crainiceanu, C. M., K \"u chenhoff, H., and Peters, A. (2008). Restricted likelihood ratio testing for zero variance components in linear mixed models. Journal of Computational and Graphical Statistics , 17(4):870--891
2008
-
[11]
Gu \'e don, T., Baey, C., and Kuhn, E. (2024). Bootstrap test procedure for variance components in nonlinear mixed effects models in the presence of nuisance parameters and a singular Fisher information matrix. Biometrika , 111(4):1331--1348
2024
-
[12]
Henderson, H. V. and Searle, S. R. (1981). On deriving the inverse of a sum of matrices. SIAM Review , 23(1):53--60
1981
-
[13]
Jiang, J. (1996). REML estimation: Asymptotic behavior and related topics. The Annals of Statistics , 24(1)
1996
-
[14]
Jiang, J. (2013). The subset argument and consistency of MLE in GLMM : Answer to an open problem and beyond. The Annals of Statistics , 41(1)
2013
-
[15]
Jiang, J. (2025). Asymptotic distribution of maximum likelihood estimator in generalized linear mixed models with crossed random effects. The Annals of Statistics , 53(3):1298--1318
2025
-
[16]
P., and Ghosh, S
Jiang, J., Wand, M. P., and Ghosh, S. (2024). Precise Asymptotics for Linear Mixed Models with Crossed Random Effects
2024
-
[17]
and Chesher, A
Lee, L.-F. and Chesher, A. (1986). Specification testing when score test statistics are identically zero. Journal of Econometrics , 31(2):121--149
1986
-
[18]
Lyu, Z., Sisson, S., and Welsh, A. (2024). Increasing dimension asymptotics for two-way crossed mixed effect models. The Annals of Statistics , 52(6):2956--2978
2024
-
[19]
Mikusheva, A. (2007). Uniform inference in autoregressive models. Econometrica , 75(5):1411--1452
2007
-
[20]
Moran, P. A. P. (1971). Maximum-likelihood estimation in non-standard conditions. Mathematical Proceedings of the Cambridge Philosophical Society , 70(3):441--450
1971
-
[21]
Patterson, H. D. and Thompson, R. (1971). Recovery of inter-block information when block sizes are unequal. Biometrika , 58(3):545--554
1971
-
[22]
Qu, L., Guennel, T., and Marshall, S. L. (2013). Linear score tests for variance components in linear mixed models and applications to genetic association studies: Linear score tests for variance components. Biometrics , 69(4):883--892
2013
-
[23]
Rothenberg, T. J. (1971). Identification in Parametric Models . Econometrica , 39(3):577
1971
-
[24]
R., Bottai, M., and Robins, J
Rotnitzky, A., Cox, D. R., Bottai, M., and Robins, J. (2000). Likelihood-based inference with singular information matrix. Bernoulli , 6(2):243--284
2000
-
[25]
Self, S. G. and Liang, K.-Y. (1987). Asymptotic properties of maximum likelihood estimators and likelihood ratio tests under nonstandard conditions. Journal of the American Statistical Association , 82(398):605--610
1987
-
[26]
Stern, S. E. and Welsh, A. H. (2000). Likelihood inference for small variance components. Canadian Journal of Statistics , 28(3):517--532
2000
-
[27]
Sung, Y. J. and Geyer, C. J. (2007). Monte Carlo likelihood inference for missing data models. The Annals of Statistics , 35(3):990--1011
2007
-
[28]
and Molenberghs, G
Verbeke, G. and Molenberghs, G. (2003). The use of score tests for inference on variance components. Biometrics , 59(2):254--262
2003
-
[29]
O., and Molstad, A
Zhang, Y., Ekvall, K. O., and Molstad, A. J. (2025). Fast and reliable confidence intervals for a variance component. Biometrika , 112(2):asaf010
2025
-
[30]
and Zhang, H
Zhu, H. and Zhang, H. (2006). Generalized score test of homogeneity for mixed effects models. The Annals of Statistics , 34(3):1545--1569
2006
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.