Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Sharp constants relating the sub-Gaussian norm and the sub-Gaussian parameter

T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper determines the exact constants in the inequalities relating the sub-Gaussian norm and the sub-Gaussian parameter: $\sqrt{3/8}\|X\|_{\psi_2} \le \sigma_X \le \sqrt{\log 2}\|X\|_{\psi_2}$, with both bounds sharp.

desk verdict Sharp constants for the psi2/sigma_X comparison, with mostly sound strategy and fixable lapses in the calculus; deserves a referee. read the letter →

arxiv 2507.05928 v1 pith:R3WYX7IA submitted 2025-07-08 math.PR math.STstat.TH

classification math.PRmath.STstat.TH MSC 60E1546E30
keywords sub-GaussiannormvarianceproxysharpconstantsoptimalOrliczLuxemburgmomentgeneratingfunctionRademacherdistribution
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper determines the exact values of the universal constants connecting the two standard measures of sub-Gaussian behaviour: the Orlicz-style norm $\|X\|_{\psi_2}$ and the variance proxy $\sigma_X$. Its main theorem states that for every centered real-valued random variable $X$, $\sqrt{3/8}\,\|X\|_{\psi_2} \le \sigma_X \le \sqrt{\log 2}\,\|X\|_{\psi_2}$, and that both inequalities are sharp. The upper bound is the genuinely new contribution: the authors prove that, among centered variables normalized to have $\|X\|_{\psi_2}\le 1$, the moment generating function is maximized by a two-point (binary) distribution, namely the Rademacher law in the sharp case. The lower bound, attained by the standard Gaussian, is proved by a short Gaussian-integral argument. Together they pin the ratio $\sigma_X/\|X\|_{\psi_2}$ to the interval $[\sqrt{3/8},\sqrt{\log 2}]$ with no room for improvement.

What carries the argument

The load-bearing object is the moment set $H$ of centered probability measures with $\int e^{x^2}\,d\mu \le 2$, together with the linear functional $\mu \mapsto \int e^{sx}\,d\mu(x)$. The argument has four moves: (i) compactness and continuity on $H$ via Skorohod coupling and weak-compactness results, so the maximum is attained; (ii) a general theorem on extreme points of moment sets, applied to the constraints $\mu(1)=1$, $\mu(x)=0$, $\mu(e^{x^2})\le 2$, which forces any extremal maximizer to have at most three atoms; (iii) a Lagrange-multiplier/Rolle argument showing a three-atom maximizer cannot exist, leaving only two-point measures; and (iv) an explicit formula for the variance proxy of a binary random variable, after which the upper bound reduces to verifying that a certain function $F(u)$ stays above $2$ for $u\ge 1$. The lower bound is carried by a separate short argument: writing $E e^{X^2/K^2}$ as an integral against the standard Gaussian density and applying the definition of $\sigma_X$.

What would settle it

Fix any real $s$ and numerically maximize $\sum_{i=1}^3 p_i e^{s x_i}$ over three-point centered distributions with $\sum_i p_i e^{x_i^2} \le 2$; if any such maximum exceeds $e^{(\log 2)s^2/2}$, the upper bound fails. Equivalently, a single three- or four-point distribution satisfying the constraint whose moment generating function exceeds the binary supremum at some $s$ would refute the theorem.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the two classical inequalities relating $\|X\|_{\psi_2}$ and $\sigma_X$ are governed by sharp constants $\sqrt{3/8}$ and $\sqrt{\log 2}$, and that the extremes are realized by the two canonical sub-Gaussian examples: the standard Gaussian distribution for the lower bound and the symmetric Bernoulli (Rademacher) distribution for the upper bound. The proof of the upper bound shows something stronger: the supremum of $E e^{sX}$ over the set $H = \{\mu : \int x\,d\mu = 0,\ \int e^{x^2}\,d\mu \le 2\}$ is attained at a binary distribution for every real $s$. This reduces a potentially complex variational problem to a one-variable calculus inequality, which yields $\sigma_\mu \le \sqrt{\log 2}$ for all $\mu \in H$.

Load-bearing premise

The proof of the sharp upper bound rests on a general theorem about extreme points of moment sets applied to the constraint $\int e^{x^2}\,d\mu \le 2$; the paper does not verify all hypotheses of that theorem for this non-polynomial constraint, so if the theorem does not apply, the reduction to binary distributions collapses.

Editorial extensions

If this is right

  • No universal inequality can improve the two constants: any centered sub-Gaussian $X$ satisfies $\sigma_X \le \sqrt{\log 2}\,\|X\|_{\psi_2}$ and $\|X\|_{\psi_2} \le \sqrt{8/3}\,\sigma_X$, and both bounds are attained.
  • The upper-bound proof yields a parameter-free recipe: whenever $E e^{X^2/K^2}\le 2$, one immediately gets $E e^{sX}\le e^{(K^2\log 2)s^2/2}$ for all real $s$, so the variance proxy is at most $\sqrt{\log 2}\,K$.
  • The lower bound shows that a bound on the variance proxy alone, $\sigma_X\le \sigma$, implies $\|X\|_{\psi_2}\le \sqrt{8/3}\,\sigma$, transferring Gaussian-type MGF control back to Orlicz-norm control with a concrete constant.
  • The extremal laws are the standard Gaussian and Rademacher distributions, so these are the test cases for any refinement of related inequalities for centered sub-Gaussian variables.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same reduction method, replacing $e^{x^2}$ with $e^{|x|^p}$ for $p>2$, would plausibly yield sharp constants between $\|X\|_{\psi_p}$ and a suitable $p$-th-order variance proxy; the extremal support size would then depend on $p$, since the moment-set dimension increases.
  • Because the sharp upper bound is attained by a discrete law, the supremum over $H$ of $E e^{sX}$ is non-smooth in $s$ near the extremal direction; numerical schemes that approximate the sup by fine grids should approach the binary bound from below, with the Rademacher law as the limiting extremizer.
  • The two sharpness examples bracket the full ratio interval, but the paper does not say whether every value in $[\sqrt{3/8},\sqrt{\log 2}]$ is realized; mixtures of a Gaussian and a Rademacher law, suitably scaled, would be natural candidates to test.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper identifies the sharp universal constants relating the sub-Gaussian norm \|X\|_{\psi_2} and the variance-proxy parameter \sigma_X for centered real-valued random variables. Theorem 1.1 claims that \sqrt{3/8}\,\|X\|_{\psi_2} \le \sigma_X \le \sqrt{\log 2}\,\|X\|_{\psi_2}, with sharpness attained by the standard normal distribution and the Rademacher distribution. The lower bound is proved by a Gaussian integral comparison; the upper bound is proved by reducing the maximization of the moment generating function over the set H = {centered laws with E e^{X^2} \le 2} to binary laws, using compactness, an extreme-point reduction based on Winkler's moment-set theorems, and a calculus lemma for binary laws. The introduction also surveys earlier known constants and situates the new result in the literature.

Significance. If the proof is completed, these appear to be the first sharp constants in a pair of inequalities that are widely used in high-dimensional probability and concentration theory. The numerical improvement from the previous best upper constant (about 1.12) to \sqrt{\log 2} \approx 0.83 is clean, and the two sharpness examples (standard Gaussian and Rademacher) are explicit and convincing. The lower-bound proof is short, correct, and self-contained. The paper also gives a useful overview of earlier bounds and references. However, the upper-bound proof as written relies on a moment-set theorem whose hypotheses are not checked, and it contains incorrect derivative identities in the binary calculus lemma; these points must be fixed before the result can be regarded as fully established.

major comments (2)
  1. [2.1.2 (Lemmas 2.6 and 2.7)] The reduction to H2 \cup H3 is the load-bearing step for the upper bound, but the paper does not state or verify the hypotheses of Winkler's theorems [16, Theorems 2.1(a) and 3.1] as applied to H in (2.7). The set H lives on noncompact R, the moment functions x and e^{x^2} are unbounded, and the exponential constraint is an inequality; these are exactly the features that require assumptions in moment-set theorems. In addition, Theorem 3.1 is invoked for a Borel-set-valued representation \bar\mu(B) = \int \nu(B)\,\pi(d\nu), which is stronger than the representation of continuous affine functionals supplied by Choquet's theorem. Please either quote the precise theorems used and verify their hypotheses for H, or replace this step with a direct proof that every extreme point of H has support of size at most three (for example, using the three moment conditions and Lemma A.1), together with Choquet's theorem applied to the continuous affine functional \mu \mapsto M_\mu(s).
  2. [2.2 (proof of Proposition 2.9)] The displayed derivative identities used to prove h2(u) \ge 0 and h(u) \ge 0 are incorrect. For h2(u) = 2c(u^3 - u - 2u\ln u) - (1+u)(u-1)^2 with c = \ln 2, one computes h2''(1) = 8c - 4 > 0, not 0, and h2'''(u) = 2c(6 + 2/u^2) - 6. For h(u) = 2c(u^3 + u^2 - u - 1 - 2u^2\ln u - 2u\ln u), the correct third derivative is h'''(u) = 2c(6 + 2/u^2 - 4/u), not 2c(6 - 2/u^2 - 4/u). The nonnegativity conclusions survive the correction: h2''(1) > 0 and h2''' > 0 on [1,\infty), while h'''(u) \ge 8c > 0 for u \ge 1. Thus the constants are not in doubt, but the proof as printed is not rigorous and the equations in this subsection must be corrected.
minor comments (4)
  1. [2.1.3 (Lemma 2.8)] The Lagrange multiplier argument writes the multiplier \lambda_2 for the inequality constraint g2 \le 0 even when this constraint is inactive; the proof works with \lambda_2 = 0 in that case, but this should be stated explicitly for completeness.
  2. [2.1.3 (Lemma 2.8)] The sentence 'Lemma 2.8 also shows that H3 is not a closed subset of the compact set H' appears before the lemma and is not part of the lemma's statement; it would be clearer placed after the proof.
  3. [2.1.3 (Lemma 2.8)] The symbol H3 is used both for the set of three-point measures in (2.2) and for the parameter set {(p,x) \in G3 : g0=g1=0, g2\le 0} inside the proof of Lemma 2.8; using a different symbol (e.g., \tilde H_3) would avoid confusion.
  4. [3] In the sentence 'This above inequality reveals...' the word 'above' is redundant; this is a minor typographical point.

Circularity Check

0 steps flagged · score 2.0 of 10

No load-bearing circularity; the only self-citation is minor and the sharp constants are derived from independent inputs.

full rationale

The claimed constants are not encoded in the definitions. The feasible set H in (2.1) is just the centering condition and the Orlicz-norm constraint E(e^{X^2}) ≤ 2; it does not involve the sub-Gaussian parameter or the target constants. The lower bound is proved independently in Section 3 via a Gaussian MGF integral and Fubini's theorem. The upper-bound proof reduces the maximization of E e^{sX} over H to finite support using Winkler's theorem [16] (Lemmas 2.6 and 2.7), then to two-point support (Lemma 2.8 and Proposition 2.1), and finally verifies the two-point variance-proxy inequality using the external formula [3, Theorem 3.1]. Neither [16] nor [3] assumes or encodes the sharp constants sqrt(3/8) or sqrt(log 2). The only self-citation is [7] in Lemma 2.5, where 'uniformly integrable, and hence tight' is cited to justify compactness; this is a standard implication and is not load-bearing for the target result. Possible correctness concerns, such as unverified hypotheses of Winkler's theorem or derivative slips in Section 2.2, would be matters of proof rigor, not circularity. Overall, no step of the derivation reduces by construction to its own input, so the central claim is grounded independently.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central derivation is largely self-contained. The non-standard black boxes are Winkler's theorem for the extreme point structure of moment sets and Buldygin-Moskvichova's formula for the binary sub-Gaussian parameter. No free parameters are fit to data and no new entities are postulated.

assumptions (4)
  • domain assumption Winkler's theorem on extreme points of moment sets [16]
    Used in Lemmas 2.6 and 2.7 to reduce the MGF maximization over H to finitely supported measures and to represent points as mixtures of extreme points. The paper cites the theorem but does not verify all hypotheses for the constraint integral exp(x^2) dmu <= 2.
  • domain assumption Binary sub-Gaussian parameter formula of Buldygin and Moskvichova [3, Theorem 3.1]
    Used in Proposition 2.9 to compute the variance proxy sigma_mu for centered two-point measures. Accepted as a black box without proof in the paper.
  • standard math Standard measure-theoretic probability facts: Skorohod coupling, Fatou's lemma, Prokhorov's theorem, uniform integrability criteria
    Used in Section 2.1.1 to prove continuity of the MGF functional and compactness of H in the weak topology, via Kallenberg [6] and Leskela-Vihola [7].
  • domain assumption Characterization of uniform integrability and tightness from [7]
    Used in Lemma 2.5 to infer tightness of H from a uniform integrability bound; this is a self-citation by one of the authors, though the result is a standard mathematical fact.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sharp constants relating the sub-Gaussian norm and the sub-Gaussian parameter." pith.science (2026). https://pith.science/paper/R3WYX7IA

@misc{pith2026250705928,
  author       = {Pith},
  title        = {Pith review of: Sharp constants relating the sub-Gaussian norm and the sub-Gaussian parameter},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/R3WYX7IA}},
  note         = {Machine review of arXiv:2507.05928}
}
abstract

We determine the optimal constants in the classical inequalities relating the sub-Gaussian norm \(\|X\|_{\psi_2}\) and the sub-Gaussian parameter \(\sigma_X\) for centered real-valued random variables. We show that \(\sqrt{3/8} \cdot \|X\|_{\psi_2} \le \sigma_X \le \sqrt{\log 2} \cdot \|X\|_{\psi_2}\), and that both bounds are sharp, attained by the standard Gaussian and Rademacher distributions, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predictively Oriented Posteriors

    stat.ME 2025-10 conditional novelty 6.0 of 10

    A new posterior family scores the posterior predictive directly, giving slower concentration but better predictive performance under model misspecification.

Reference graph

Works this paper leans on

17 extracted references · 17 canonical work pages · cited by 1 Pith paper

  1. [1]

    St´ ephane Boucheron, G´ abor Lugosi, and Pascal Massart,Concentration inequalities: A nonasymptotic theory of independence, Oxford University Press, 2013

  2. [2]

    V. V. Buldygin and Y. V. Kozachenko, Metric characterization of random variables and random processes, American Mathematical Society, 2000

  3. [3]

    V. V. Buldygin and K. K. Moskvichova, The sub-Gaussian norm of a binary random variable, Theory of Probability and Mathematical Statistics 86 (2013), 33–49

  4. [4]

    Djalil Chafa ¨ ı, Olivier Gu´ edon, Guillaume Lecu´ e, and Alain Pajor,Interactions be- tween compressed sensing random matrices and high dimensional geometry, Soci´ et´ e Math´ ematique de France, 2012

  5. [5]

    Ethier and Thomas G

    Stewart N. Ethier and Thomas G. Kurtz, Markov processes: Characterization and convergence, Wiley, 1986. 11

  6. [6]

    Olav Kallenberg, Foundations of modern probability, second ed., Springer, 2002

  7. [7]

    1, 382–389

    Lasse Leskel¨ a and Matti Vihola,Stochastic order characterization of uniform inte- grability and tightness, Statistics & Probability Letters 83 (2013), no. 1, 382–389

  8. [8]

    Yang Li, A simple note on the basic properties of subgaussian random variables, 2024+, https://arxiv.org/abs/2407.07348

Show all 17 references
  1. [9]

    E. J. McShane, The Lagrange multiplier rule, American Mathematical Monthly 80 (1973), no. 8, 922–925

  2. [10]

    David Pollard, Probability tools, tricks, and miracles, http://www.stat.yale.edu/ ~pollard/Books/Pttm/, 2024

  3. [11]

    Omar Rivasplata, Subgaussian random variables: An expository note , https://personalpages.manchester.ac.uk/staff/omar.rivasplata/pubs/ note/subgaussians.pdf, 2012

  4. [12]

    Walter Rudin, Principles of mathematical analysis, third ed., McGraw–Hill, 1976

  5. [13]

    edu/~rvan/APC550.pdf, 2016

    Ramon van Handel, Probability in high dimension, https://web.math.princeton. edu/~rvan/APC550.pdf, 2016

  6. [14]

    Roman Vershynin, High-dimensional probability, Cambridge University Press, 2018

  7. [15]

    Wainwright, High-dimensional statistics, Cambridge University Press, 2019

    Martin J. Wainwright, High-dimensional statistics, Cambridge University Press, 2019

  8. [16]

    4, 581–587

    Gerhard Winkler, Extreme points of moment sets, Mathematics of Operations Re- search 13 (1988), no. 4, 581–587

  9. [17]

    Huiming Zhang and Song Xi Chen, Concentration inequalities for statistical infer- ence, Communications in Mathematical Research 37 (2021), no. 1, 1–85. 12

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.