Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Uniform inference in linear mixed models

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper proves that score-based chi-squared confidence regions for random-effect variances and covariances have uniformly correct coverage in linear mixed models, including on the singular boundary where variances are zero or…

desk verdict This paper makes a genuine advance in uniform inference for variance parameters in linear mixed models, including crossed random effects; the main risk is the deferred proof of the key lemma. read the letter →

arxiv 2507.19633 v1 pith:H566U7MG submitted 2025-07-25 math.ST stat.MEstat.TH

classification math.STstat.MEstat.TH MSC 62F1262F2562J10
keywords linearmixedmodelsrandomeffectsscorestatisticuniforminferenceboundaryofparametersetcrossedfinite-sampleboundsconfidenceregions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proves that, for linear mixed models, the score statistic standardized by expected Fisher information is approximately normal with finite-sample error bounds that hold uniformly over the parameter set, including points where the random-effect covariance matrix is singular—variance zero or correlation ±1. That uniformity is the key advance: existing pointwise asymptotic results fail near the boundary because Wald and likelihood ratio statistics change their asymptotic distributions as the parameter approaches the boundary, while the score statistic does not. From the bounds the authors construct chi-squared confidence regions for variance and covariance parameters whose coverage error is controlled explicitly, so one can judge whether a sample is large enough and obtain asymptotically correct uniform coverage. The same mechanism covers crossed random effects, where the number of independent observations is not the sample size, and where no comparable boundary theory existed. Simulations show these regions stay near nominal coverage in finite samples in settings where Wald and likelihood ratio regions drift.

What carries the argument

The load-bearing object is the ratio $a(v,\psi)$, defined as the spectral norm divided by the Frobenius norm of $A(v,\psi)=\Sigma(\psi)^{-1/2}\Sigma(v)\Sigma(\psi)^{-1/2}$; it measures how far a direction $v$ in parameter space is from being an identifiable, well-conditioned direction. Lemma 4, built on a normal approximation for quadratic forms, turns smallness of $a(v,\psi)$ into an explicit bound on $|g(t;v,\psi)-\phi(t)|$, the difference between the density of $v^T W^S(\psi)$ and the standard normal density. Lemma 5 converts the pointwise bound into a uniform one, and Theorem 2 converts that into asymptotic normality of every linear combination of the standardized score. For crossed random effects the analysis uses the fact that $\Sigma(\psi)$ is a linear combination of fixed projection matrices, so the eigenvectors of $\Sigma$ do not depend on $\psi$; this permits a spectral argument that yields the $\psi$-free bound involving $n_{\min}$ and $r$.

What would settle it

Simulate a crossed random-effects model with $r=3$ fixed factors but with one factor's size kept at $n_{\min}=1$ while the others grow, compute the restricted score confidence region at parameters where a random-effect variance is zero, and check whether the actual coverage error stays as small as the paper's uniform bound predicts; if it does, the paper's sufficient condition $n_{\min}\ge3$ is shown to be unnecessary, and if it does not, the condition is doing real work. A more direct check is to evaluate $a(v,\psi)$ numerically for such a design and verify whether the Theorem 4 bound is violated.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the distribution of any linear combination of the standardized score $W^S(\psi) = I(\psi)^{-1/2}S(\psi)$ is close to standard normal whenever the quantity $a(v,\psi)=\|A(v,\psi)\|/\|A(v,\psi)\|_F$ is small, where $A(v,\psi)=\Sigma(\psi)^{-1/2}\Sigma(v)\Sigma(\psi)^{-1/2}$ and $\Sigma(v)=\sum_j v_j \partial\Sigma/\partial\psi_j$. Theorem 2 formalizes this as asymptotic normality, uniformly over unit vectors $v_n$, and extends it to unknown fixed effects $\beta$ when $X$ has full column rank. In independent-cluster settings Theorem 3 gives the bound $a(v,\psi)\le c_3 m^{-1/2}(1+\|\psi_{-r}\|/\psi_r)$, so the approximation improves as the number of clusters grows even at the boundary. In crossed random effects settings Theorem 4 gives a bound depending only on the smallest factor size $n_{\min}$ and the number of factors $r$—not on $\psi$—so Corollary 4 yields $\sup_{\psi}|P_\psi\{\psi\in \tilde{C}^S_n(\alpha)\}-(1-\alpha)|\to0$ over the whole parameter set when $r\ge3$ is fixed and $n_{\min}\to\infty$. The paper thus claims that score-based chi-squared confidence regions solve the boundary problem for these models, and the experiments support that claim for the restricted likelihood version.

Load-bearing premise

The whole argument leans on the design matrices giving every random effect enough independent replication: in the clustered setting each cluster's within-cluster design matrix must be sufficiently well conditioned, and in the crossed settings every crossed factor must have at least two (or three for the restricted score) levels, with the smallest factor size growing for the asymptotic results; if that fails, the key ratio $a(v,\psi)$ need not be small and the uniform normality bound does not follow.

Editorial extensions

If this is right

  • If the central claim is right, the chi-squared quantile gives confidence regions for random-effect variances and covariances with asymptotically correct uniform coverage, including on the boundary where a variance is zero or a correlation is ±1.
  • For independent clusters, coverage improves at rate roughly $m^{-1/2}$ in the number of clusters, so the finite-sample bounds can be used to say how many clusters are enough before trusting a chi-squared region.
  • For crossed random effects, the effective sample size for variance-component inference is governed by the smallest factor size $n_{\min}$, not the total number of observations, and uniform coverage holds over the entire parameter set when $r$ is fixed and $n_{\min}\to\infty$.
  • The results also hold when the number of random-effect parameters grows with the sample size, and the restricted-likelihood version remains valid with a growing number of fixed-effect predictors.
  • Wald and likelihood-ratio regions do not have this property, so score-based regions are the default choice when boundary parameters are a live possibility.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We infer that the same $a(v,\psi)$ ratio could serve as a practical diagnostic: before reporting a score-based confidence region, one could compute its maximum over the boundary of the fitted region and flag cases where the normal approximation is likely poor.
  • We infer that the mechanism is not fundamentally tied to normality: the score is a quadratic form, and the paper's remarks suggest the bound could be reworked for non-normal responses by replacing the normality-based moment conditions with estimates of third and fourth moments, perhaps via resampling.
  • We infer that allocating experimental units across all crossed factors matters more than total sample size for variance-component inference, a point with direct implications for designing multi-factor studies.
  • We infer that extending this style of bound to generalized linear mixed models would hit the same difficulty the paper identifies: without the linearity and joint normality, the spectral decomposition of $\Sigma$ is lost, so the crossed-random-effects proof would need a different route.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper develops finite-sample and asymptotic distributional approximations for score-based inference in Gaussian linear mixed models, with emphasis on uniform coverage near the boundary of the random-effects covariance parameter space, where variances may vanish or correlations may approach ±1. The main engine is Lemma 4, a finite-sample bound on the sup-norm distance between the density of a standardized linear combination of the score and the standard normal density, controlled by the ratio a(v, psi) of spectral to Frobenius norms of A(v, psi). Lemma 5 converts this into a uniform bound; Theorem 2 derives asymptotic normality of arbitrary linear combinations of the standardized score under sup_v a_n(v, psi_n) -> 0, including with unknown fixed effects. Theorems 3 and 4 give design-specific bounds for independent clusters and crossed random effects, and Corollaries 2-4 translate these into asymptotically correct uniform coverage for chi-square confidence regions, including at the boundary. Simulations illustrate near-nominal coverage for score regions and failures of Wald and likelihood-ratio regions in the same settings.

Significance. If the deferred proofs are correct, this is a substantial contribution: it provides a general treatment of uniform inference for covariance parameters at the singular boundary of the parameter set, with explicit finite-sample bounds and clear sufficient conditions. The design-specific bounds (m^{-1/2} for independent clusters, functions of n_min for crossed effects) are informative and lead to practically implementable confidence regions. The paper's strengths include explicit constants, a clean separation between the generic approximation and design-specific analyses, treatment of crossed random effects with diverging dimensions, reproducible simulation code, and falsifiable predictions about coverage. The main reservations are verification risks rather than demonstrated errors: the central normal approximation is imported from a co-authored prior paper whose proof is not in the reviewed material, and the extension to unknown beta requires a nontrivial argument that the main text does not supply.

major comments (3)
  1. [Section 4.1, Lemma 4] Lemma 4 is the sole finite-sample engine of the paper, but its proof is deferred to a supplement that was not included in the reviewed material, and the lemma is attributed to a normal approximation for quadratic forms due to Zhang et al. (2025), which was developed for a single variance component. In the present setting A(v, psi) is generally indefinite because the H_j contain off-diagonal ones, so the cited result must handle indefinite quadratic forms under the condition a(tilde v, psi)^2 < 1/8. Please state the precise conditions and either prove Lemma 4 or reproduce the relevant theorem from the cited paper; without this, the main results cannot be independently checked.
  2. [Section 4.1, Theorem 2, second assertion] The claim that u_n^T W_n^S(theta_n) converges to N(0,1) when beta is unknown is not immediate from Lemma 5. The score for beta is exactly normal, but it is uncorrelated with, not generally independent of, the quadratic score for psi; in a Gaussian vector a linear form and a quadratic form can have zero covariance yet be dependent. The proof must show how this dependence is controlled in the linear combination u_n^T W_n^S(theta_n). The remark that the result is 'suggested, but not implied' by the marginal normality of the beta-score confirms that a gap remains in the main text.
  3. [Sections 4.2 and 4.3, finite-sample applicability] The finite-sample bound (15) is valid only when a(tilde v, psi)^2 < 1/8. The design-specific bounds in Theorem 3 (18) and Theorem 4 are often larger than this threshold for small m or small n_min; for example, Theorem 4 with n_min = 3 gives tilde a^2 <= 1. The paper should state explicitly that the finite-sample statements are conditional on the right-hand side being below 1/8, and that the asymptotic corollaries rely on this threshold being eventually satisfied.
minor comments (5)
  1. [Section 2, Lemma 2] In the displayed definition of q_{1-alpha}(psi), the inequality appears reversed: it should be min{t : F(t) >= 1-alpha}, not min{t : F(1-alpha) >= t}.
  2. [Section 4.3] In the display defining the crossed-effects projections, P2 = Z^{(2)}Z^{(2)T}/n_2 should be divided by n_1, not n_2, to make P2 a projection matrix; the subsequent expression for Sigma confirms this.
  3. [Theorem 2] The notation P_n subset of R^{r_n x r_n} in the theorem statement should be P_n subset of R^{r_n}; the parameter vector psi_n is in R^{r_n}, not in a matrix space.
  4. [Section 4.3] The symbol P_j is used both for the n_j x n_j projection matrix 1_{n_j}1_{n_j}^T/n_j and for the n x n Kronecker-product projection; please use distinct symbols or otherwise clarify to avoid confusion.
  5. [Section 5] The phrase 'fine-sample coverage probabilities' should read 'finite-sample coverage probabilities'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: uniform-coverage claims follow from stated quantitative bounds; the co-authored Lemma 4 is an external, published inequality, not a fitted input or the target conclusion.

full rationale

The derivation chain is self-contained in structure. The score statistic is defined from the log-likelihood and Fisher information (Section 3). Lemma 4 supplies a finite-sample density bound for a linear combination of the standardized score, citing the normal-approximation result of Zhang et al. (2025); this is a published inequality whose stated assumptions do not include the paper's conclusions, and the bound is then stated explicitly in Eq. (15). Lemma 5 turns a uniform bound on a(v, psi) into a uniform density bound. Theorems 3 and 4 prove that this key quantity is small under explicit design-replication conditions (bounded eigenvalues of Z_i^T Z_i, n_min growing), rather than assuming it. Corollaries 2-4 then combine these bounds with Lemmas 1-2 to obtain uniform asymptotic coverage. No parameter is fitted to data and renamed as a prediction, no target result is assumed inside the proof of the smallness bound, and no uniqueness theorem is imported from the authors' prior work. The self-citations present are not load-bearing: Lemma 1 is compared to Lemma 2.5 of Ekvall and Bottai (2022) but is stated and proved independently, and Lemma 4 is attributed to Zhang et al. (2025), a co-authored paper, but is an external mathematical result with explicit assumptions and no dependence on the current paper's conclusions. The only caveat is that Lemma 4's proof is deferred to the supplement, which is a verification risk rather than circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No fitted constants enter the theory; the only external input is the published quadratic-form approximation. The Gaussian model, the linear parameterization of Psi, and the identifiability condition are standard assumptions rather than invented objects.

assumptions (3)
  • domain assumption The response follows the Gaussian linear mixed model (1) with correctly specified mean and covariance, and Psi is a linear function of psi whose element-level derivatives are 0/1 matrices.
    Section 3 sets up the model; all finite-sample and asymptotic results are derived under this model. Misspecification is mentioned only as a possible extension.
  • standard math Lemma 4, the normal approximation bound for quadratic forms borrowed from Zhang et al. (2025), is correct as stated, including the constants 0.14 and 0.29.
    Equation (15) is the basis of the main finite-sample bounds; it is cited, not re-proven in the main text, and the supplement was not available for review.
  • domain assumption The parameter set P = {psi : Psi(psi) >= 0, psi_r > 0} with Psi positive semidefinite is the correct constraint set, and the Fisher information is positive definite under the stated identifiability condition (12).
    Theorem 1 and the surrounding discussion; if the model is unidentifiable, the standardized score is not well defined.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Uniform inference in linear mixed models." pith.science (2026). https://pith.science/paper/H566U7MG

@misc{pith2026250719633,
  author       = {Pith},
  title        = {Pith review of: Uniform inference in linear mixed models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H566U7MG}},
  note         = {Machine review of arXiv:2507.19633}
}
abstract

We provide finite-sample distribution approximations, that are uniform in the parameter, for inference in linear mixed models. Focus is on variances and covariances of random effects in cases where existing theory fails because their covariance matrix is nearly or exactly singular, and hence near or at the boundary of the parameter set. Quantitative bounds on the differences between the standard normal density and those of linear combinations of the score function enable, for example, the assessment of sufficient sample size. The bounds also lead to useful asymptotic theory in settings where both the number of parameters and the number of random effects grow with the sample size. We consider models with independent clusters and ones with a possibly diverging number of crossed random effects, which are notoriously complicated. Simulations indicate the theory leads to practically relevant methods. In particular, the studied confidence regions, which are straightforward to implement, have near-nominal coverage in finite samples even when some random effects have variances near or equal to zero, or correlations near or equal to $\pm 1$.

Figures

Figures reproduced from arXiv: 2507.19633 by the authors.

Figure 1
Figure 1. Quantiles of test-statistics evaluated at a true [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Coverage probabilities with independent clusters and correlated random effects, for [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. Coverage probabilities with crossed, independent random effects, for different [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Universal inference for variance components

    stat.ME 2025-08 conditional novelty 6.0 of 10

    A randomized split likelihood ratio test provides finite-sample valid inference for variance components at the boundary, including heritability near 1, with computational shortcuts for structured models.

Reference graph

Works this paper leans on

30 extracted references · 29 canonical work pages · cited by 1 Pith paper

  1. [1]

    Bates, D. et al. (2005). Fitting linear mixed models in R . R news , 5(1):27--30

  2. [2]

    Battey, H. S. and McCullagh, P. (2024). An anomaly arising in the analysis of processes with more than one source of variability. Biometrika , 111(2):677--689

  3. [3]

    Bottai, M. (2003). Confidence regions when the Fisher information is zero. Biometrika , 90(1):73--84

  4. [4]

    Chesher, A. (1984). Testing for neglected heterogeneity. Econometrica , 52(4):865--872

  5. [5]

    Cox, D. R. and Hinkley, D. V. (2000). Theoretical Statistics . Chapman & Hall/CRC, Boca Raton

  6. [6]

    Crainiceanu, C. M. and Ruppert, D. (2004). Likelihood ratio tests in linear mixed models with one variance component. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 66(1):165--185

  7. [7]

    Ekvall, K. O. and Bottai, M. (2022). Confidence regions near singular information and boundary points with applications to mixed models. The Annals of Statistics , 50(3):1806--1832

  8. [8]

    Ekvall, K. O. and Jones, G. L. (2020). Consistent maximum likelihood estimation using subsets with applications to multivariate mixed models. The Annals of Statistics , 48(2):932--952

Show all 30 references
  1. [9]

    Geyer, C. J. (1994). On the Asymptotics of Constrained M-Estimation . The Annals of Statistics , 22(4):1993--2010

  2. [10]

    M., K \"u chenhoff, H., and Peters, A

    Greven, S., Crainiceanu, C. M., K \"u chenhoff, H., and Peters, A. (2008). Restricted likelihood ratio testing for zero variance components in linear mixed models. Journal of Computational and Graphical Statistics , 17(4):870--891

  3. [11]

    Gu \'e don, T., Baey, C., and Kuhn, E. (2024). Bootstrap test procedure for variance components in nonlinear mixed effects models in the presence of nuisance parameters and a singular Fisher information matrix. Biometrika , 111(4):1331--1348

  4. [12]

    Henderson, H. V. and Searle, S. R. (1981). On deriving the inverse of a sum of matrices. SIAM Review , 23(1):53--60

  5. [13]

    Jiang, J. (1996). REML estimation: Asymptotic behavior and related topics. The Annals of Statistics , 24(1)

  6. [14]

    Jiang, J. (2013). The subset argument and consistency of MLE in GLMM : Answer to an open problem and beyond. The Annals of Statistics , 41(1)

  7. [15]

    Jiang, J. (2025). Asymptotic distribution of maximum likelihood estimator in generalized linear mixed models with crossed random effects. The Annals of Statistics , 53(3):1298--1318

  8. [16]

    P., and Ghosh, S

    Jiang, J., Wand, M. P., and Ghosh, S. (2024). Precise Asymptotics for Linear Mixed Models with Crossed Random Effects

  9. [17]

    and Chesher, A

    Lee, L.-F. and Chesher, A. (1986). Specification testing when score test statistics are identically zero. Journal of Econometrics , 31(2):121--149

  10. [18]

    Lyu, Z., Sisson, S., and Welsh, A. (2024). Increasing dimension asymptotics for two-way crossed mixed effect models. The Annals of Statistics , 52(6):2956--2978

  11. [19]

    Mikusheva, A. (2007). Uniform inference in autoregressive models. Econometrica , 75(5):1411--1452

  12. [20]

    Moran, P. A. P. (1971). Maximum-likelihood estimation in non-standard conditions. Mathematical Proceedings of the Cambridge Philosophical Society , 70(3):441--450

  13. [21]

    Patterson, H. D. and Thompson, R. (1971). Recovery of inter-block information when block sizes are unequal. Biometrika , 58(3):545--554

  14. [22]

    Qu, L., Guennel, T., and Marshall, S. L. (2013). Linear score tests for variance components in linear mixed models and applications to genetic association studies: Linear score tests for variance components. Biometrics , 69(4):883--892

  15. [23]

    Rothenberg, T. J. (1971). Identification in Parametric Models . Econometrica , 39(3):577

  16. [24]

    R., Bottai, M., and Robins, J

    Rotnitzky, A., Cox, D. R., Bottai, M., and Robins, J. (2000). Likelihood-based inference with singular information matrix. Bernoulli , 6(2):243--284

  17. [25]

    Self, S. G. and Liang, K.-Y. (1987). Asymptotic properties of maximum likelihood estimators and likelihood ratio tests under nonstandard conditions. Journal of the American Statistical Association , 82(398):605--610

  18. [26]

    Stern, S. E. and Welsh, A. H. (2000). Likelihood inference for small variance components. Canadian Journal of Statistics , 28(3):517--532

  19. [27]

    Sung, Y. J. and Geyer, C. J. (2007). Monte Carlo likelihood inference for missing data models. The Annals of Statistics , 35(3):990--1011

  20. [28]

    and Molenberghs, G

    Verbeke, G. and Molenberghs, G. (2003). The use of score tests for inference on variance components. Biometrics , 59(2):254--262

  21. [29]

    O., and Molstad, A

    Zhang, Y., Ekvall, K. O., and Molstad, A. J. (2025). Fast and reliable confidence intervals for a variance component. Biometrika , 112(2):asaf010

  22. [30]

    and Zhang, H

    Zhu, H. and Zhang, H. (2006). Generalized score test of homogeneity for mixed effects models. The Annals of Statistics , 34(3):1545--1569

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.