REVIEW 2 major objections 4 minor 1 cited by
Sharp constants relating the sub-Gaussian norm and the sub-Gaussian parameter
T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper determines the exact constants in the inequalities relating the sub-Gaussian norm and the sub-Gaussian parameter: $\sqrt{3/8}\|X\|_{\psi_2} \le \sigma_X \le \sqrt{\log 2}\|X\|_{\psi_2}$, with both bounds sharp.
desk verdict Sharp constants for the psi2/sigma_X comparison, with mostly sound strategy and fixable lapses in the calculus; deserves a referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the moment set $H$ of centered probability measures with $\int e^{x^2}\,d\mu \le 2$, together with the linear functional $\mu \mapsto \int e^{sx}\,d\mu(x)$. The argument has four moves: (i) compactness and continuity on $H$ via Skorohod coupling and weak-compactness results, so the maximum is attained; (ii) a general theorem on extreme points of moment sets, applied to the constraints $\mu(1)=1$, $\mu(x)=0$, $\mu(e^{x^2})\le 2$, which forces any extremal maximizer to have at most three atoms; (iii) a Lagrange-multiplier/Rolle argument showing a three-atom maximizer cannot exist, leaving only two-point measures; and (iv) an explicit formula for the variance proxy of a binary random variable, after which the upper bound reduces to verifying that a certain function $F(u)$ stays above $2$ for $u\ge 1$. The lower bound is carried by a separate short argument: writing $E e^{X^2/K^2}$ as an integral against the standard Gaussian density and applying the definition of $\sigma_X$.
What would settle it
Fix any real $s$ and numerically maximize $\sum_{i=1}^3 p_i e^{s x_i}$ over three-point centered distributions with $\sum_i p_i e^{x_i^2} \le 2$; if any such maximum exceeds $e^{(\log 2)s^2/2}$, the upper bound fails. Equivalently, a single three- or four-point distribution satisfying the constraint whose moment generating function exceeds the binary supremum at some $s$ would refute the theorem.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the two classical inequalities relating $\|X\|_{\psi_2}$ and $\sigma_X$ are governed by sharp constants $\sqrt{3/8}$ and $\sqrt{\log 2}$, and that the extremes are realized by the two canonical sub-Gaussian examples: the standard Gaussian distribution for the lower bound and the symmetric Bernoulli (Rademacher) distribution for the upper bound. The proof of the upper bound shows something stronger: the supremum of $E e^{sX}$ over the set $H = \{\mu : \int x\,d\mu = 0,\ \int e^{x^2}\,d\mu \le 2\}$ is attained at a binary distribution for every real $s$. This reduces a potentially complex variational problem to a one-variable calculus inequality, which yields $\sigma_\mu \le \sqrt{\log 2}$ for all $\mu \in H$.
Load-bearing premise
The proof of the sharp upper bound rests on a general theorem about extreme points of moment sets applied to the constraint $\int e^{x^2}\,d\mu \le 2$; the paper does not verify all hypotheses of that theorem for this non-polynomial constraint, so if the theorem does not apply, the reduction to binary distributions collapses.
Editorial extensions
If this is right
- No universal inequality can improve the two constants: any centered sub-Gaussian $X$ satisfies $\sigma_X \le \sqrt{\log 2}\,\|X\|_{\psi_2}$ and $\|X\|_{\psi_2} \le \sqrt{8/3}\,\sigma_X$, and both bounds are attained.
- The upper-bound proof yields a parameter-free recipe: whenever $E e^{X^2/K^2}\le 2$, one immediately gets $E e^{sX}\le e^{(K^2\log 2)s^2/2}$ for all real $s$, so the variance proxy is at most $\sqrt{\log 2}\,K$.
- The lower bound shows that a bound on the variance proxy alone, $\sigma_X\le \sigma$, implies $\|X\|_{\psi_2}\le \sqrt{8/3}\,\sigma$, transferring Gaussian-type MGF control back to Orlicz-norm control with a concrete constant.
- The extremal laws are the standard Gaussian and Rademacher distributions, so these are the test cases for any refinement of related inequalities for centered sub-Gaussian variables.
Reading between the lines
- The same reduction method, replacing $e^{x^2}$ with $e^{|x|^p}$ for $p>2$, would plausibly yield sharp constants between $\|X\|_{\psi_p}$ and a suitable $p$-th-order variance proxy; the extremal support size would then depend on $p$, since the moment-set dimension increases.
- Because the sharp upper bound is attained by a discrete law, the supremum over $H$ of $E e^{sX}$ is non-smooth in $s$ near the extremal direction; numerical schemes that approximate the sup by fine grids should approach the binary bound from below, with the Rademacher law as the limiting extremizer.
- The two sharpness examples bracket the full ratio interval, but the paper does not say whether every value in $[\sqrt{3/8},\sqrt{\log 2}]$ is realized; mixtures of a Gaussian and a Rademacher law, suitably scaled, would be natural candidates to test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper identifies the sharp universal constants relating the sub-Gaussian norm \|X\|_{\psi_2} and the variance-proxy parameter \sigma_X for centered real-valued random variables. Theorem 1.1 claims that \sqrt{3/8}\,\|X\|_{\psi_2} \le \sigma_X \le \sqrt{\log 2}\,\|X\|_{\psi_2}, with sharpness attained by the standard normal distribution and the Rademacher distribution. The lower bound is proved by a Gaussian integral comparison; the upper bound is proved by reducing the maximization of the moment generating function over the set H = {centered laws with E e^{X^2} \le 2} to binary laws, using compactness, an extreme-point reduction based on Winkler's moment-set theorems, and a calculus lemma for binary laws. The introduction also surveys earlier known constants and situates the new result in the literature.
Significance. If the proof is completed, these appear to be the first sharp constants in a pair of inequalities that are widely used in high-dimensional probability and concentration theory. The numerical improvement from the previous best upper constant (about 1.12) to \sqrt{\log 2} \approx 0.83 is clean, and the two sharpness examples (standard Gaussian and Rademacher) are explicit and convincing. The lower-bound proof is short, correct, and self-contained. The paper also gives a useful overview of earlier bounds and references. However, the upper-bound proof as written relies on a moment-set theorem whose hypotheses are not checked, and it contains incorrect derivative identities in the binary calculus lemma; these points must be fixed before the result can be regarded as fully established.
major comments (2)
- [2.1.2 (Lemmas 2.6 and 2.7)] The reduction to H2 \cup H3 is the load-bearing step for the upper bound, but the paper does not state or verify the hypotheses of Winkler's theorems [16, Theorems 2.1(a) and 3.1] as applied to H in (2.7). The set H lives on noncompact R, the moment functions x and e^{x^2} are unbounded, and the exponential constraint is an inequality; these are exactly the features that require assumptions in moment-set theorems. In addition, Theorem 3.1 is invoked for a Borel-set-valued representation \bar\mu(B) = \int \nu(B)\,\pi(d\nu), which is stronger than the representation of continuous affine functionals supplied by Choquet's theorem. Please either quote the precise theorems used and verify their hypotheses for H, or replace this step with a direct proof that every extreme point of H has support of size at most three (for example, using the three moment conditions and Lemma A.1), together with Choquet's theorem applied to the continuous affine functional \mu \mapsto M_\mu(s).
- [2.2 (proof of Proposition 2.9)] The displayed derivative identities used to prove h2(u) \ge 0 and h(u) \ge 0 are incorrect. For h2(u) = 2c(u^3 - u - 2u\ln u) - (1+u)(u-1)^2 with c = \ln 2, one computes h2''(1) = 8c - 4 > 0, not 0, and h2'''(u) = 2c(6 + 2/u^2) - 6. For h(u) = 2c(u^3 + u^2 - u - 1 - 2u^2\ln u - 2u\ln u), the correct third derivative is h'''(u) = 2c(6 + 2/u^2 - 4/u), not 2c(6 - 2/u^2 - 4/u). The nonnegativity conclusions survive the correction: h2''(1) > 0 and h2''' > 0 on [1,\infty), while h'''(u) \ge 8c > 0 for u \ge 1. Thus the constants are not in doubt, but the proof as printed is not rigorous and the equations in this subsection must be corrected.
minor comments (4)
- [2.1.3 (Lemma 2.8)] The Lagrange multiplier argument writes the multiplier \lambda_2 for the inequality constraint g2 \le 0 even when this constraint is inactive; the proof works with \lambda_2 = 0 in that case, but this should be stated explicitly for completeness.
- [2.1.3 (Lemma 2.8)] The sentence 'Lemma 2.8 also shows that H3 is not a closed subset of the compact set H' appears before the lemma and is not part of the lemma's statement; it would be clearer placed after the proof.
- [2.1.3 (Lemma 2.8)] The symbol H3 is used both for the set of three-point measures in (2.2) and for the parameter set {(p,x) \in G3 : g0=g1=0, g2\le 0} inside the proof of Lemma 2.8; using a different symbol (e.g., \tilde H_3) would avoid confusion.
- [3] In the sentence 'This above inequality reveals...' the word 'above' is redundant; this is a minor typographical point.
Circularity Check
No load-bearing circularity; the only self-citation is minor and the sharp constants are derived from independent inputs.
full rationale
The claimed constants are not encoded in the definitions. The feasible set H in (2.1) is just the centering condition and the Orlicz-norm constraint E(e^{X^2}) ≤ 2; it does not involve the sub-Gaussian parameter or the target constants. The lower bound is proved independently in Section 3 via a Gaussian MGF integral and Fubini's theorem. The upper-bound proof reduces the maximization of E e^{sX} over H to finite support using Winkler's theorem [16] (Lemmas 2.6 and 2.7), then to two-point support (Lemma 2.8 and Proposition 2.1), and finally verifies the two-point variance-proxy inequality using the external formula [3, Theorem 3.1]. Neither [16] nor [3] assumes or encodes the sharp constants sqrt(3/8) or sqrt(log 2). The only self-citation is [7] in Lemma 2.5, where 'uniformly integrable, and hence tight' is cited to justify compactness; this is a standard implication and is not load-bearing for the target result. Possible correctness concerns, such as unverified hypotheses of Winkler's theorem or derivative slips in Section 2.2, would be matters of proof rigor, not circularity. Overall, no step of the derivation reduces by construction to its own input, so the central claim is grounded independently.
Assumptions & free parameters
assumptions (4)
- domain assumption Winkler's theorem on extreme points of moment sets [16]
- domain assumption Binary sub-Gaussian parameter formula of Buldygin and Moskvichova [3, Theorem 3.1]
- standard math Standard measure-theoretic probability facts: Skorohod coupling, Fatou's lemma, Prokhorov's theorem, uniform integrability criteria
- domain assumption Characterization of uniform integrability and tightness from [7]
Cite this review
Pith. "Pith review of Sharp constants relating the sub-Gaussian norm and the sub-Gaussian parameter." pith.science (2026). https://pith.science/paper/R3WYX7IA
@misc{pith2026250705928,
author = {Pith},
title = {Pith review of: Sharp constants relating the sub-Gaussian norm and the sub-Gaussian parameter},
year = {2026},
howpublished = {\url{https://pith.science/paper/R3WYX7IA}},
note = {Machine review of arXiv:2507.05928}
}
abstract
We determine the optimal constants in the classical inequalities relating the sub-Gaussian norm \(\|X\|_{\psi_2}\) and the sub-Gaussian parameter \(\sigma_X\) for centered real-valued random variables. We show that \(\sqrt{3/8} \cdot \|X\|_{\psi_2} \le \sigma_X \le \sqrt{\log 2} \cdot \|X\|_{\psi_2}\), and that both bounds are sharp, attained by the standard Gaussian and Rademacher distributions, respectively.
Forward citations
Cited by 1 Pith paper
-
Predictively Oriented Posteriors
A new posterior family scores the posterior predictive directly, giving slower concentration but better predictive performance under model misspecification.
Reference graph
Works this paper leans on
-
[1]
St´ ephane Boucheron, G´ abor Lugosi, and Pascal Massart,Concentration inequalities: A nonasymptotic theory of independence, Oxford University Press, 2013
work page 2013
-
[2]
V. V. Buldygin and Y. V. Kozachenko, Metric characterization of random variables and random processes, American Mathematical Society, 2000
work page 2000
-
[3]
V. V. Buldygin and K. K. Moskvichova, The sub-Gaussian norm of a binary random variable, Theory of Probability and Mathematical Statistics 86 (2013), 33–49
work page 2013
-
[4]
Djalil Chafa ¨ ı, Olivier Gu´ edon, Guillaume Lecu´ e, and Alain Pajor,Interactions be- tween compressed sensing random matrices and high dimensional geometry, Soci´ et´ e Math´ ematique de France, 2012
work page 2012
-
[5]
Stewart N. Ethier and Thomas G. Kurtz, Markov processes: Characterization and convergence, Wiley, 1986. 11
work page 1986
-
[6]
Olav Kallenberg, Foundations of modern probability, second ed., Springer, 2002
work page 2002
-
[7]
Lasse Leskel¨ a and Matti Vihola,Stochastic order characterization of uniform inte- grability and tightness, Statistics & Probability Letters 83 (2013), no. 1, 382–389
work page 2013
-
[8]
Yang Li, A simple note on the basic properties of subgaussian random variables, 2024+, https://arxiv.org/abs/2407.07348
work page Pith review arXiv 2024
Show all 17 references
-
[9]
E. J. McShane, The Lagrange multiplier rule, American Mathematical Monthly 80 (1973), no. 8, 922–925
1973
-
[10]
David Pollard, Probability tools, tricks, and miracles, http://www.stat.yale.edu/ ~pollard/Books/Pttm/, 2024
2024
-
[11]
Omar Rivasplata, Subgaussian random variables: An expository note , https://personalpages.manchester.ac.uk/staff/omar.rivasplata/pubs/ note/subgaussians.pdf, 2012
2012
-
[12]
Walter Rudin, Principles of mathematical analysis, third ed., McGraw–Hill, 1976
1976
-
[13]
edu/~rvan/APC550.pdf, 2016
Ramon van Handel, Probability in high dimension, https://web.math.princeton. edu/~rvan/APC550.pdf, 2016
2016
-
[14]
Roman Vershynin, High-dimensional probability, Cambridge University Press, 2018
2018
-
[15]
Wainwright, High-dimensional statistics, Cambridge University Press, 2019
Martin J. Wainwright, High-dimensional statistics, Cambridge University Press, 2019
2019
-
[16]
4, 581–587
Gerhard Winkler, Extreme points of moment sets, Mathematics of Operations Re- search 13 (1988), no. 4, 581–587
1988
-
[17]
Huiming Zhang and Song Xi Chen, Concentration inequalities for statistical infer- ence, Communications in Mathematical Research 37 (2021), no. 1, 1–85. 12
2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.