Pith. sign in

REVIEW 2 major objections 5 minor 23 references

$L_2$-norm posterior contraction in Gaussian models with unknown variance

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper proves that a four-case likelihood-ratio test yields exponentially small testing errors for Gaussian hypotheses separated in the L2 metric even when variance is unknown, and that this restores optimal posterior contraction rates.

desk verdict A genuinely useful local test for unknown variance, with a fixable gap in the high-dimensional application. read the letter →

arxiv 2506.00401 v2 pith:K72CKLBQ submitted 2025-05-31 math.ST stat.TH

classification math.STstat.TH MSC 62C1062G2062G0862J07
keywords BayesiannonparametricsHigh-dimensionalregressionNonparametricTesting-basedposteriorcontractionL2-normUnknownvarianceGaussianmodelLocaltest
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proves that the testing-based route to posterior contraction works in Gaussian models even when the error variance is unknown. Its key device is a local test, built from likelihood-ratio and residual-thresholding ideas, that separates the true pair (mean vector, variance) from alternatives at least $\epsilon$ away in an $L_2$-type metric, with both error probabilities bounded by $e^{-K n \epsilon^2/\sigma_0^2}$ for a universal constant $K$. This is enough to build a global test over a sieve and to give sufficient conditions under which the posterior contracts at rate $\epsilon_n$ around the truth. The result is relevant because previous unknown-variance tests either required truncating the variance prior or lost the variance-normalized decay rate, blocking optimal rates when $\sigma_0$ varies with $n$. The paper demonstrates the machinery on high-dimensional sparse regression and adaptive nonparametric regression.

What carries the argument

The load-bearing object is the local test of Theorem 1, a four-case decision rule on the observation $y$ that separates $(\mu_0,\sigma_0)$ from alternatives at distance at least $\epsilon$ in the metric $d((\mu,\sigma),(\mu',\sigma')) = \sqrt{n^{-1}\|\mu-\mu'\|_2^2+|\sigma-\sigma'|^2}$. Case 1 thresholds a centered chi-square statistic; Case 2 uses the inner product $(\mu_1-\mu_0)^\top(y-\mu_0)$; Cases 3 and 4 threshold the mean absolute deviation $n^{-1}\|y-\mu_0\|_1$ depending on whether the alternative variance is above or below $\sigma_0$. Its key property is that both error types decay as $e^{-K n\epsilon^2/\sigma_0^2}$ uniformly over the $\epsilon/6$-neighborhood, a rate unavailable from the prior unknown-variance tests. Theorem 2 converts this local property into a global test via covering numbers, and Theorems 3 and 4 convert the global test into posterior contraction through prior concentration, entropy, and sieve-mass conditions.

What would settle it

Recalculate the Case 3 and Case 4 inequalities at the boundary $\epsilon=\sigma_0$ that the theorem excludes: the displayed bound $[-(7/20)\sqrt{7/8}+\sqrt{2}/6]\epsilon<-\epsilon/11$ still holds, but the step $\sigma\le 2M_1\sigma_0$ no longer turns the exponent into a multiple of $n\epsilon^2/\sigma_0^2$, so checking whether the type-II probability remains at most $e^{-K n\epsilon^2/\sigma_0^2}$ at that boundary would settle whether the $\epsilon<\sigma_0$ restriction is necessary.

Watch

Extended reading notes

Core claim

The paper's central claim is that a test exists for $y\sim N_n(\mu_0,\sigma_0^2 I_n)$ that separates $(\mu_0,\sigma_0)$ from any $(\mu_1,\sigma_1)$ with $d((\mu_1,\sigma_1),(\mu_0,\sigma_0))\ge\epsilon$, provided $\epsilon<\sigma_0$, where $d^2=n^{-1}\|\mu_1-\mu_0\|_2^2+|\sigma_1-\sigma_0|^2$. The test is a four-case construction: a chi-square threshold when $\sigma_1$ is much larger than $\sigma_0$; an inner-product test in the direction $\mu_1-\mu_0$ when the mean difference dominates; and a mean-absolute-deviation threshold when the variance difference dominates. In all four cases the type-I error and the worst-case type-II error over the $\epsilon/6$-ball around $(\mu_1,\sigma_1)$ are at most $e^{-K n\epsilon^2/\sigma_0^2}$. The author then assembles these local tests into a global test using metric entropy (Theorem 2) and plugs it into the standard testing framework for posterior contraction (Theorems 3 and 4), with applications to high-dimensional regression and nonparametric regression. The rate $n\epsilon^2/\sigma_0^2$ is what gives the result its scope: it matches the known-variance benchmark and remains meaningful when $\sigma_0$ changes with $n$.

Load-bearing premise

The whole construction assumes the target separation $\epsilon$ is smaller than the true noise level $\sigma_0$, so the results cannot certify posterior contraction when the noise variance shrinks faster than the desired accuracy.

Editorial extensions

If this is right

  • Any Gaussian model $y\sim N_n(\mu_0,\sigma_0^2 I_n)$ whose prior satisfies the three conditions in (2) contracts at rate $\epsilon_n$ in the metric $d$, under the scale assumptions $\epsilon_n/\sigma_0\to0$ and $n\epsilon_n^2/\sigma_0^2\to\infty$.
  • For high-dimensional sparse regression, the framework delivers the rate $\sigma_0\sqrt{|S_0|\log p/n}$ for $n^{-1/2}\|X(\beta-\beta_0)\|_2+|\sigma-\sigma_0|$, matching the known-variance rate up to constants.
  • For adaptive nonparametric regression with B-spline priors, it delivers $((\sigma_0^2\log n)/n)^{\alpha/(2\alpha+1)}$ in the empirical $L_2$ norm plus $|\sigma-\sigma_0|$ for $\alpha$-smooth regression functions.
  • Because the error rate is normalized by $\sigma_0^2$, the results stay valid when $\sigma_0$ varies with $n$, provided $\epsilon_n/\sigma_0\to0$.
  • Inverse gamma and half-Cauchy priors for the variance are admissible under mild Lipschitz and polynomial-tail conditions, so priors supported on the full positive line are covered.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The $\epsilon<\sigma_0$ condition is likely not a proof artifact: at the boundary the Case 3 and Case 4 bounds lose the $\sigma_0$-normalized exponent, so extending the theorem to all $\epsilon$ would need a genuinely scale-free test statistic, perhaps a Studentized version of the mean-absolute-deviation test.
  • The same four-case logic could transfer to dependent Gaussian settings such as autoregressive or white-noise-in-time models; the author names these as possible extensions, and a variance-normalized decaying error rate would be the property to check.
  • In the high-dimensional application, the assumptions $\|\beta_0\|_\infty=O(\log p)$ and $\|X\|_\infty=O(\log p)$ are used to pass from $\ell_\infty$ control to $\ell_2$ control; a compatibility-condition-only version would show whether these are removable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper develops a testing-based framework for L2 posterior contraction in Gaussian models N_n(µ0, σ0²I) with σ0² unknown. Theorem 1 constructs a local test with exponentially small type-I and type-II errors at rate nε²/σ0² for separations ε<σ0. Theorem 2 combines these local tests through metric entropy into a global test, and Theorems 3 and 4 give posterior contraction sufficient conditions, with Theorem 4 isolating convenient conditions on the prior for σ². The framework is applied to high-dimensional sparse regression and to adaptive B-spline nonparametric regression.

Significance. The local-test construction is the main contribution and is genuinely useful: unlike Lemma 8.27 of Ghosal and van der Vaart (2017) it does not require truncating the σ² prior, and unlike Proposition S2 of Naulet and Barat (2018) it has the rate nε²/σ0² needed when σ0² varies with n. The proofs are self-contained, use standard tail bounds, and do not fit parameters to data; the restriction ε<σ0 is explicitly acknowledged. If the application gap identified below is repaired, Theorems 1 and 2 provide a clean testing route for L2 contraction in unknown-variance Gaussian models.

major comments (2)
  1. [Section 3.1, verification of condition (5)]
  2. [Section 3.1, same verification]
minor comments (5)
  1. [Section 3.1, entropy display]
  2. [Theorem 3 proof]
  3. [Theorem 4 and Section 3.1]
  4. [Theorem 4 proof]
  5. [Section 3.1, assumptions]

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the local test is constructed explicitly and the applications verify standard testing conditions.

full rationale

The central derivation is self-contained. Theorem 1 builds explicit tests in four separate cases (chi-square threshold, linear score test, and two L1-based tests) and proves exponential type-I and type-II error bounds directly, without assuming the conclusion. Theorems 2 through 4 are standard testing-based arguments that combine this local test with metric entropy and prior-concentration conditions. The self-citations to Jeong and Ghosal (2021a, 2021b) are used only for KL-divergence formulas and compatibility-condition references, which are independently checkable and are not load-bearing for the main test construction. The high-dimensional regression step in Section 3.1 does contain a substantive mathematical gap: the displayed lower bound e^{-C1|S0| log p} is not supported by the preceding inequality when ||beta0||_2^2 can be as large as |S0| log^2 p, because the factor e^{-||beta0||_2^2} can dominate the target rate. That is a correctness concern, not a circularity, since no fitted parameter is renamed as a prediction and no result is equivalent to its input by construction. The paper's explicit acknowledgement that Theorem 1 requires epsilon < sigma0 and therefore cannot handle the shell approach is a stated limitation, not a hidden circular dependency. Overall, the derivation chain does not reduce to its own assumptions.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper introduces no free parameters fitted to data; all constants in the proofs are universal. The listed axioms are standard probability results and explicit prior conditions. No new entities (particles, fields, or dimensions) are postulated.

assumptions (6)
  • standard math Chi-square tail bounds (Laurent and Massart 2000, Lemma 1): Pr(chi^2_{k,0} >= k + 2 sqrt(kt) + 2t) <= e^{-t} and Pr(chi^2_{k,0} <= k - 2 sqrt(kt)) <= e^{-t}.
    Used in Theorem 1 Case 1 for both type-I and type-II error bounds.
  • standard math Monotonicity of the noncentral chi-square CDF in the noncentrality parameter (Marcum Q-function).
    Used in Case 1 type-II to bound Pr(chi^2_{n,lambda} <= t) by Pr(chi^2_{n,0} <= t).
  • standard math Sub-Gaussian concentration for sums of independent sub-Gaussian variables (Wainwright 2019, Proposition 2.5): Pr({|sum X_i| > t}) <= e^{-c t^2/n}.
    Used in Cases 3 and 4 to control the concentration of the L1 norm of a standard normal vector.
  • domain assumption KL divergence and variance expansions for Gaussian location-scale models, e.g., Theorem 9 of Jeong and Ghosal (2021b).
    Used in Theorem 3 to translate prior concentration in L2 and sigma into a KL-neighborhood condition. The calculation is standard and independently verifiable.
  • standard math B-spline approximation theory: any function in the Holder class H^alpha can be approximated by B-splines with sup-norm error C J^{-alpha} and bounded coefficients.
    Used in Section 3.2 to compute the prior concentration for the nonparametric regression application.
  • domain assumption Sufficient conditions on the prior of sigma^2 in Theorem 4: a priori independence of mu and sigma^2, a polynomial right tail, an L-Lipschitz density, and the lower bound g(sigma0^2) >= 2L sigma0 epsilon_n.
    These conditions define the class of variance priors covered by the theorem. They are explicitly stated and are mild for inverse-gamma and half-Cauchy priors under the stated growth ranges.

how reviews work

0 comments
Cite this review

Pith. "Pith review of $L_2$-norm posterior contraction in Gaussian models with unknown variance." pith.science (2026). https://pith.science/paper/K72CKLBQ

@misc{pith2026250600401,
  author       = {Pith},
  title        = {Pith review of: $L_2$-norm posterior contraction in Gaussian models with unknown variance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/K72CKLBQ}},
  note         = {Machine review of arXiv:2506.00401}
}
abstract

The testing-based approach is a fundamental tool for establishing posterior contraction rates. Although the Hellinger metric is attractive owing to the existence of a desirable test function, it is not directly applicable in Gaussian models, because translating the Hellinger metric into more intuitive metrics typically requires strong boundedness conditions. When the variance is known, this issue can be addressed by directly constructing a test function relative to the $L_2$-metric using the likelihood ratio test. However, when the variance is unknown, existing results are limited and rely on restrictive assumptions. To overcome this limitation, we derive a test function tailored to an unknown variance setting with respect to the $L_2$-metric and provide sufficient conditions for posterior contraction based on the testing-based approach. We apply this result to analyze high-dimensional regression and nonparametric regression.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

23 extracted references · 22 canonical work pages

  1. [1]

    Birg \'e , L. (1983). Robust testing for independent non identically distributed variables and M arkov chains. In Specifying Statistical Models: From Parametric to Non-Parametric, Using Bayesian or Non-Bayesian Approaches , pages 134--162. Springer

  2. [2]

    Birg \'e , L. (2006). Model selection via testing: an alternative to (penalized) maximum likelihood estimators. Annales de l'Institut Henri Poincar \'e , Probabilit \'e s et Statistiques , 42(3):273--325

  3. [3]

    Castillo, I., Schmidt-Hieber, J., and van der Vaart, A. (2015). Bayesian linear regression with sparse priors. The Annals of Statistics , 43(5):1986--2018

  4. [4]

    and van der Vaart, A

    Castillo, I. and van der Vaart, A. (2012). Needles and straw in a haystack: posterior concentration for possibly sparse sequences. The Annals of Statistics , 40(3):1367--1391

  5. [5]

    De Boor, C. (1978). A Practical Guide to Splines , volume 27. Springer-Verlag New York

  6. [6]

    L., Johnstone, I

    Donoho, D. L., Johnstone, I. M., Kerkyacharian, G., and Picard, D. (1995). Wavelet shrinkage: asymptopia? Journal of the Royal Statistical Society: Series B (Methodological) , 57(2):301--337

  7. [7]

    K., and van d er Vaart, A

    Ghosal, S., Ghosh, J. K., and van d er Vaart, A. W. (2000). Convergence rates of posterior distributions. The Annals of Statistics , 28(2):500--531

  8. [8]

    and van d er Vaart, A

    Ghosal, S. and van d er Vaart, A. (2007). Convergence rates of posterior distributions for noniid observations. The Annals of Statistics , 35(1):192--223

Show all 23 references
  1. [9]

    and van d er Vaart, A

    Ghosal, S. and van d er Vaart, A. (2017). Fundamentals of Nonparametric Bayesian Inference , volume 44. Cambridge University Press

  2. [10]

    and Ghosal, S

    Jeong, S. and Ghosal, S. (2021a). Posterior contraction in sparse generalized linear models. Biometrika , 108(2):367--379

  3. [11]

    and Ghosal, S

    Jeong, S. and Ghosal, S. (2021b). Unified B ayesian theory of sparse linear regression with nuisance parameters. Electronic Journal of Statistics , 15(1):3040--3111

  4. [12]

    and Massart, P

    Laurent, B. and Massart, P. (2000). Adaptive estimation of a quadratic functional by model selection. The Annals of Statistics , 28(5):1302--1338

  5. [13]

    Le Cam, L. (1973). Convergence of estimates under dimensionality restrictions. The Annals of Statistics , 1(2):38--53

  6. [14]

    and Barat, E

    Naulet, Z. and Barat, E. (2018). Some aspects of symmetric gamma process mixtures. Bayesian Analysis , 13(3):703--720

  7. [15]

    Ning, B., Jeong, S., and Ghosal, S. (2020). Bayesian linear regression for multivariate responses under group sparsity. Bernoulli , 26(3):2353--2382

  8. [16]

    Polson, N. G. and Ro c kov \'a , V. (2018). Posterior concentration for sparse deep learning. In Advances in Neural Information Processing Systems , volume 31, pages 930--939

  9. [17]

    and van der Pas, S

    Ro c kov \'a , V. and van der Pas, S. (2020). Posterior concentration for B ayesian regression trees and forests. The Annals of Statistics , 48(4):2108--2131

  10. [18]

    Schwartz, L. (1965). On B ayes procedures. Zeitschrift f \"u r Wahrscheinlichkeitstheorie und verwandte Gebiete , 4(1):10--26

  11. [19]

    and Ghosal, S

    Shen, W. and Ghosal, S. (2015). Adaptive B ayesian procedures using random series priors. Scandinavian Journal of Statistics , 42(4):1194--1213

  12. [20]

    and Liang, F

    Song, Q. and Liang, F. (2023). Nearly optimal B ayesian shrinkage for high-dimensional regression. Science China Mathematics , 66(2):409--442

  13. [21]

    van der Vaart, A. W. and van Zanten, J. H. (2008). Rates of contraction of posterior distributions based on G aussian process priors. The Annals of Statistics , 36(3):1435--1463

  14. [22]

    Wainwright, M. J. (2019). High-Dimensional Statistics: A Non-Asymptotic Viewpoint , volume 48. Cambridge University Press

  15. [23]

    Yoo, W. W. and Ghosal, S. (2016). Supremum norm posterior contraction and credible sets for nonparametric multivariate regression. The Annals of Statistics , 44(3):1069--1102

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.