Pith. sign in

REVIEW 2 major objections 2 minor 12 references

Estimation of the sub-Gaussian parameter

T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read An empirical maximizer of the weighted cumulant function consistently estimates the sub-Gaussian parameter at near root-n rates when the maximizer exists.

desk verdict The paper supplies the first rates and minimax lower bounds for estimating the sub-Gaussian variance proxy via constrained empirical maximization of the cumulant function L. read the letter →

arxiv 2606.06384 v1 pith:NYVJZT4B submitted 2026-06-04 math.ST stat.MEstat.MLstat.TH

classification math.STstat.MEstat.MLstat.TH
keywords sub-Gaussianparametervarianceproxyempiricalcumulantmaximizationconsistencyratesminimaxoptimalitytailestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper defines the sub-Gaussian parameter as the supremum of L(λ) = (2/λ²) log E[exp(λX)] and constructs an estimator by constrained maximization of the corresponding empirical function. It proves consistency with rate O_p(n^{-1/2+ε}) whenever L attains its supremum, improving to the parametric O_p(n^{-1/2}) rate when that argmax is bounded; the same estimator is shown to be minimax optimal inside this subclass. The authors also establish necessity of the assumption by proving that the unrestricted minimax risk remains Ω(1) and that the risk interpolates between Ω(1/log n) and Ω(1) under successively weaker tail-growth conditions on L. When the underlying distribution is not sub-Gaussian the estimator diverges at a rate governed by the actual tails, and the method is applied to construct p-values in a large-scale permutation test for gene-ontology enrichment.

What carries the argument

Constrained maximizer of the empirical weighted cumulant generating function L_n(λ).

What would settle it

A sequence of sub-Gaussian distributions whose argmax of L tends to infinity, together with a demonstration that the estimator fails to converge faster than a constant.

Watch

Extended reading notes

Core claim

The estimator obtained by constrained maximization of the empirical version of L is consistent at rate O_p(n^{-1/2+ε}) whenever L possesses a maximizer and improves to O_p(n^{-1/2}) when the argmax is bounded; this estimator is minimax optimal on that subclass, while the minimax risk over all sub-Gaussian laws is Ω(1).

Load-bearing premise

That the population function L attains its supremum at some finite λ (or at a bounded λ).

Editorial extensions

If this is right

  • Root-n consistency and minimax optimality hold inside the subclass where the argmax of L is bounded.
  • The estimator diverges to infinity at a tail-determined rate when the distribution is not sub-Gaussian.
  • The method supplies a practical alternative to peaks-over-threshold for constructing p-values in permutation tests.
  • Minimax lower bounds form a continuum from Ω(1) down to Ω(1/log n) as tail-growth assumptions on L are strengthened.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Many common sub-Gaussian families satisfy the bounded-argmax condition in practice, making the parametric rate attainable without further restrictions.
  • The interpolation between constant and logarithmic minimax lower bounds suggests a hierarchy of tail classes that could be used for adaptive estimation.
  • The same estimator could be inserted into concentration bounds or risk analyses that currently treat the sub-Gaussian parameter as known.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper defines the sub-Gaussian parameter ξ²_* as the supremum of L(λ) = (2/λ²) log E[exp(λX)] for mean-zero X. It proposes an estimator obtained by constrained maximization of the empirical version of L, proves consistency with rates O_p(n^{-1/2+ε}) when L attains its maximum and O_p(n^{-1/2}) when the argmax is additionally bounded, establishes matching minimax lower bounds that are Ω(1) without the maximizer assumption and that interpolate between Ω(1/log n) and Ω(1) under progressively stronger tail-growth conditions on L, shows the estimator diverges at a tail-controlled rate when the distribution is not sub-Gaussian, and applies the estimator to construct p-values in a Gene Ontology enrichment permutation test.

Significance. If the stated rates, minimax optimality, and necessity results hold, the work supplies the first explicit estimation theory and optimality analysis for the sub-Gaussian parameter, a quantity that appears throughout concentration inequalities yet has lacked rigorous estimators. The continuum of minimax lower bounds and the clean separation of regimes according to properties of L are technically notable; the non-sub-Gaussian divergence result and the concrete statistical application further strengthen the contribution.

major comments (2)
  1. [Abstract] Abstract: the rate O_p(n^{-1/2+ε}) is stated to hold whenever L has a maximizer, yet the abstract does not indicate whether the ε arises from a uniform integrability argument or from a truncation that depends on the location of the maximizer; this distinction is load-bearing for whether the result extends to unbounded argmax locations.
  2. [Abstract] The minimax lower-bound construction that yields Ω(1) without the maximizer assumption is described only at the level of the abstract; the precise least-favorable family and the reduction to the empirical-process term used for the upper bound should be cross-referenced to confirm that the lower-bound construction does not inadvertently impose a bounded argmax.
minor comments (2)
  1. [Abstract] Abstract contains two typographical errors: 'consistent bound the rates' should read 'consistent and bound the rates'; 'an maximizer' should read 'a maximizer'.
  2. The precise form of the constraint imposed on the empirical maximization (e.g., the radius of the λ-ball or the penalty) is not stated in the abstract and should appear explicitly when the estimator is first defined.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the careful reading and constructive comments on the abstract. We address each point below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the rate O_p(n^{-1/2+ε}) is stated to hold whenever L has a maximizer, yet the abstract does not indicate whether the ε arises from a uniform integrability argument or from a truncation that depends on the location of the maximizer; this distinction is load-bearing for whether the result extends to unbounded argmax locations.

    Authors: The ε arises from a truncation argument in the proof (see the argument leading to Theorem 3.1) that depends on the (possibly unbounded) location of the maximizer. The result does extend to unbounded argmax locations provided a maximizer exists, with the ε loss; the bounded-argmax case recovers the root-n rate without truncation. We will revise the abstract to indicate the origin of ε and this distinction. revision: yes

  2. Referee: [Abstract] The minimax lower-bound construction that yields Ω(1) without the maximizer assumption is described only at the level of the abstract; the precise least-favorable family and the reduction to the empirical-process term used for the upper bound should be cross-referenced to confirm that the lower-bound construction does not inadvertently impose a bounded argmax.

    Authors: The Ω(1) lower bound is proved in Section 4.2 via a least-favorable family of distributions where L does not attain its supremum (explicitly constructed so that the argmax is at infinity). The reduction to the empirical-process term is the same as in the upper-bound analysis and does not impose boundedness. We will add a cross-reference to Section 4.2 in the abstract. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The derivation defines ξ²* explicitly as sup L(λ) with L the weighted CGF, constructs the estimator as constrained max of the empirical L, and derives consistency/rates/minimax optimality from standard concentration arguments conditioned on explicit properties of L (existence of maximizer or bounded argmax). Matching lower bounds of Ω(1) are shown without those assumptions, establishing necessity rather than circularity. No equation reduces a claimed rate or optimality result to a fitted parameter by construction, no self-citation chain is load-bearing for the central claims, and the analysis remains self-contained against external probabilistic benchmarks.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper rests on the standard definition of the cumulant generating function and the mean-zero assumption for X; no free parameters or invented entities are introduced in the abstract.

assumptions (2)
  • domain assumption X is a mean-zero random variable.
    Explicitly stated in the definition of L(λ).
  • standard math Standard properties of the expectation, logarithm, and supremum hold for the cumulant function.
    Invoked in the definition of ξ²_* and the empirical analogue.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Estimation of the sub-Gaussian parameter." pith.science (2026). https://pith.science/paper/NYVJZT4B

@misc{pith2026260606384,
  author       = {Pith},
  title        = {Pith review of: Estimation of the sub-Gaussian parameter},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NYVJZT4B}},
  note         = {Machine review of arXiv:2606.06384}
}
abstract

The sub-Gaussian parameter (also called the variance proxy) of a mean-zero random variable $X$ is defined as $\xi^2_* = \sup_{\lambda \in \mathbb{R}} L(\lambda)$ where $L(\lambda) = \frac{2}{\lambda^2} \log \mathbb{E} e^{\lambda X}$ is a weighted cumulant generating function. Despite the ubiquity of sub-Gaussian random variables, the estimation of $\xi^2_*$ has received little attention and is not yet well understood. In this work, we study a natural estimator of $\xi^2_*$ based on constrained maximization of the empirical analogue of $L$. We prove that the estimator is consistent bound the rates of convergence under assumptions on $L$: if $L$ has an maximizer, then our bound is $O_p(n^{-1/2 + \varepsilon})$ for any $\varepsilon > 0$; if the argmax of $L$ is also bounded, then the bound improves to $O_p(n^{-1/2})$. We show that our assumptions on $L$ are necessary by proving that the minimax risk over all sub-Gaussian distributions is $\Omega(1)$; imposing increasingly strong assumptions on the tail growth of $L$ yields a continuum of classes whose minimax lower bound interpolates between $\Omega(1/\log n)$ and $\Omega(1)$. Root-n rate is possible if we restrict to a subclass of distributions where $L$ attains its supremum in a bounded region, in which case our estimator is minimax optimal. If the underlying distribution is not sub-Gaussian, we show that our estimator goes to infinity with a divergence rate controlled by the tail of the distribution. Finally, we apply our estimator in a Gene Ontology (GO) enrichment study to construct p-values for a large-scale permutation test, showing that it can serve as a reliable alternative to the peaks-over-threshold approach, particularly in regimes where the peaks-over-threshold method is of uncertain validity.

Figures

Figures reproduced from arXiv: 2606.06384 by the authors.

Figure 1
Figure 1. Log-log plots for the mean absolute deviation for truncated ˆξ 2 n and untruncated ˆξ 2 n,U . Similarly, we shall write ψ(λ; P) := log M(λ), ψn(λ) = log Mn(λ), and ψ˜ n(λ) = log M˜ n(λ). When P is fixed, we will abbreviate M(λ) ≡ M(λ; P) and ψ(λ) ≡ ψ(λ; P). Proposition 2 (Feuerverger (1989), Theorems 2.1 and 2.4). Let I be the largest interval on which M(λ) is finite, and define ψ˜ n(λ) = log M˜ n(λ). Let k ∈ N, and… view at source ↗
Figure 2
Figure 2. Plots of L(λ; P). In Example 1(a), arg maxL exists so δ is eventually negative. In (c), L is increasing away from 0 so that δ(C) > 0 for all C > 0. 2.3 Inference Under even stronger assumptions, we can also derive the limiting distribution of ˆξ 2 n . Proposition 4. For λ ̸= 0, define V (λ) := 4 λ4  M(2λ) M(λ) 2 − 2λψ′ (λ) + λ 2σ 2 − 1  . Also, let V (0) := EX4 − σ 4 . Suppose that L is uniquely maximized at λ ∗ ,… view at source ↗
Figure 3
Figure 3. Distribution of √ n( ˆξ 2 n − ξ 2 ∗ ) for four choices of P. Proposition 4 gives asymptotic normality for all cases but N(0, 1). In practice, the hypotheses of Theorem 2 and Proposition 4 are difficult to check. One heuristic approach is to plot Ln on [−Cn, Cn], and check if the maximizers are in the interior. An affirmative answer would be consistent with the hypotheses in Theorem 2(c) and Proposition 4; on the oth… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

12 extracted references · 3 canonical work pages

  1. [1]

    Optimal sub-Gaussian variance proxy for 3-mass distributions,

    Atouani, S., Marchal, O. and Arbel, J. (2025). Optimal sub-gaussian variance proxy for 3-mass distributions. URL:https://arxiv.org/abs/2510.06132 Crane, H. and Xu, M. (2024). Root and community inference on the latent growth process of a network,Journal of the Royal Statistical Society: series B (statistical methodology)86: 825–865. Degen, M., Embrechts, ...

  2. [2]

    and Shamir, R

    URL:https://doi.org/10.1214/26-ECP761 Levi, H., Elkon, R. and Shamir, R. (2021). Domino: a network-based active module identification algorithm with reduced rate of false calls,Molecular systems biology17(1): e9593. Liu, J., Xu, M. and Xing, J. (2026). Beyond single algorithms: A framework for validating and aggregating active modules in genetic interacti...

  3. [3]

    H., Cook, H., Kuhn, M., Wyder, S., Simonovic, M., Santos, A., Doncheva, N

    URL:https://doi.org/10.1038/s41598-024-53346-z Szklarczyk, D., Morris, J. H., Cook, H., Kuhn, M., Wyder, S., Simonovic, M., Santos, A., Doncheva, N. T., Roth, A., Bork, P., Jensen, L. J. and von Mering, C. (2017). The STRING database in 2017: quality-controlled protein-protein association networks, made broadly accessible,Nucleic Acids Res.45(D1): D362–D3...

  4. [4]

    Then, for any λ∈ [−Cn, Cn], there isπλ ∈ Πn such that |λ−π λ|< n −1, and the following holds by the triangle inequality: |M(k) n (λ)−M (k)(λ)| ≤ |M (k) n (λ)−M (k) n (πλ)|+|M (k) n (πλ)−M (k)(πλ)|+|M (k)(πλ)−M (k)(λ)|.(4) Let us first address the middle term of (4). By Proposition 13 and then (3), P max π∈Πn |M(k) n (π)−M (k)(π)|> n −1/2+ε ≤NP |M(k) n (π)...

  5. [5]

    Then, for anyλ∈R, L(λ;P)≤ 2 λ2 loge |λ| ≤ 2 |λ|

    Proof of Proposition 5.LetPbe supported on[−1,1]. Then, for anyλ∈R, L(λ;P)≤ 2 λ2 loge |λ| ≤ 2 |λ| . Combining this with the fact thatL(0;P) =σ2(P)≥0, we have that, for anyC >0, δP (C) = sup |λ|≥C L(λ;P)−sup |λ|≤C L(λ;P)≤sup |λ|≥C L(λ;P)≤sup |λ|≥C 2 |λ| ≤ 2 C . The first claim of the Proposition thus follows. Now supposeσ 2(P)≥ 3 2 δ0. LetC 1 := 8δ−1 0 and...

  6. [6]

    (2025) gives exact expressions for the sub-Gaussian parameter of three-point distributions whenp≤1/6, they are still rather difficult to work with

    While Atouani et al. (2025) gives exact expressions for the sub-Gaussian parameter of three-point distributions whenp≤1/6, they are still rather difficult to work with. Instead, it will be easier to settle for an estimate ofξ2 ∗(P )via Proposition

  7. [7]

    Define q0 = 1 2 and qn := q0 + 1/(4√n)

    For q∈ (0, 1), definePq := (1 −q )δ0 + (q/2)δ1 + (q/2)δ−1. Define q0 = 1 2 and qn := q0 + 1/(4√n). It is clear thatξ2 ∗(Pq0)and ξ2 ∗(Pqn)are both no larger than 1 since the distributions are supported on[ −1, 1]. Moreover, by Proposition 5,Pq0 , Pqn ∈ P (Ξ, C0, r−). Then, by Proposition 18 and the fact thatq 0 andq n are both greater than1/3, we have |ξ2 ...

  8. [8]

    As for its empirical analogueLn, first let λ∗ ∈ Λ∗(P )

    Therefore, none of the maximizers ofLcan be outside of the interval[−C0, C0]. As for its empirical analogueLn, first let λ∗ ∈ Λ∗(P ). By Proposition 3(a), there exists a sequence ηn = o(1)dependent only onΞand δ0 and C0 such that, with probability at least1−η n, sup |λ|≤C0 Ln(λ)≥L n(λ∗)≥L(λ ∗)−sup |λ|≤C0 |Ln(λ)−L(λ)|> ξ 2 ∗ − δ0 2 .(26) By Proposition 3(b...

Show all 12 references
  1. [9]

    Ifthemean-zeroGaussianprocess GP withcovariancefunction( f, g) 7→E P f(X)g(X)−EP f(X)EP g(X) has a version that is a tight Borel measurable element ofℓ∞(F), then F is P -pre-Gaussian.If F isP-pre-Gaussian for eachP∈ P, and it holds that sup P∈P EP ∥GP ∥F <∞andlim δ→0 sup P∈P E...

  2. [10]

    Moreover, ifCn is made larger thanC, then we must have˜ξ2 n(P) =ξ 2 ∗(P)for allP∈ P(Ξ, C 0, r−)

    Proof of Theorem 3.According to Proposition 8, we may makenlarge enough that the event Λ∗ n ∪Λ ∗(P)⊂[−C 0, C0] occurs with high probability, uniformly overP(Ξ, C0, r−). Moreover, ifCn is made larger thanC, then we must have˜ξ2 n(P) =ξ 2 ∗(P)for allP∈ P(Ξ, C 0, r−). Therefore w...

  3. [11]

    Provided n be large enough, λ will be contained in[−Cn, Cn], and by a law of large numbers for infinite means (Durrett (2019), Theorem 2.3.8), ˜Mn(λ) a.s

    Proof of Proposition 11.First suppose M fails to exist at someλ, and EX <∞ . Provided n be large enough, λ will be contained in[−Cn, Cn], and by a law of large numbers for infinite means (Durrett (2019), Theorem 2.3.8), ˜Mn(λ) a.s. − − → ∞, and then ˆξ2 n ≥2λ −2ψn(λ) = 2λ−2 lo...

  4. [12]

    Note that∆ n = |X(n) − ¯X| ∨ |X (1) − ¯X|

    Proof of Proposition 12.For anyλ∈[0, C n]we have ˆξ2 n ≥ 2 λ2 log 1 n nX i=1 eλ(Xi− ¯X) ! ≥ 2 λ2 log 1 n e(X(n)− ¯X)λ = 2 λ (X(n) − ¯X)− logn λ , 30 and on the other hand ifλ∈[−Cn,0]then ˆξ2 n ≥ 2 λ2 log 1 n nX i=1 eλ(Xi− ¯X) ! ≥ 2 λ2 log 1 n e(X(1)− ¯X)λ ≥ 2 |λ| |X(1) − ¯X| −...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.