REVIEW 2 major objections 5 minor 36 references
Bounds on the Excess Minimum Risk via Generalized Information Divergence Measures
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper proves that the excess minimum risk of estimating a target from a degraded observation is bounded by Rényi, Jensen–Shannon, and Sibson divergence measures, with sub-Gaussian parameters that may depend on the target.
desk verdict Two of the three generalized-divergence bounds are real contributions; the Sibson bound has a genuine proof gap that needs an added boundedness assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the variational representation of divergences: the Donsker–Varadhan formula for KL divergence, a variational characterization of Rényi divergence, and a variational characterization of Sibson mutual information. Each representation expresses a divergence as a supremum over test functions, which turns a sub-Gaussian tail bound on the loss into a quadratic inequality that controls the difference of expected losses. The auxiliary-distribution method—introducing a tilted intermediate distribution and splitting the gap into two KL terms—carries the Jensen–Shannon and Sibson proofs. The target-dependent sub-Gaussian parameter $\sigma^2(y)$ enters through the quadratic term of these inequalities, and Jensen and Cauchy–Schwarz steps collect it into the factor $E[\sigma^2(Y)]$.
What would settle it
Compute the inequality from Lemma 6 with unbounded $\gamma^2$: for $\gamma^2(U^*) = e^U$ with $U$ standard normal and $t=2$, $\log E[e^{t \gamma^2(U^*)}] > t \log E[e^{\gamma^2(U^*)}]$, so the proof step fails; a counterexample to (70) under the stated assumptions, or a revised theorem with an added boundedness condition, would settle the scope of the Sibson bound. For the Rényi and Jensen–Shannon bounds, a direct numerical comparison of the bound against the true excess risk for a nonconstant $\sigma^2(y)$ model would detect any violation.
Extended reading notes
Core claim
Under the Markov condition $Y \to X \to Z$, the paper establishes, for every $\alpha \in (0,1)$, inequalities of the form $$L^*_l(Y|Z) - L^*_l(Y|X) \le \sqrt{\frac{2E[\$sigma^{2}$(Y)]}{\$\alpha$}\, D_\$\alpha$(P_{X|Y,Z}\|P_{X|Z}|P_{Y,Z})}$$ for Rényi divergence, with analogous bounds using the $\alpha$-Jensen–Shannon divergence and Sibson's $\alpha$-mutual information. The Rényi bound recovers the earlier mutual-information excess-risk bound as $\alpha \to 1$; the Jensen–Shannon bound recovers it as $\alpha \to 0$ and gives a Lautum-information bound as $\alpha \to 1$; the Sibson bound recovers it in the constant-parameter case. For bounded losses the constants become explicit in terms of $\|l\|_\infty$. The paper thus claims that the excess risk is governed by how strongly the presence of $Y$ changes the conditional law of $X$ given $Z$, measured at any order $\alpha$.
Load-bearing premise
The bounds all assume that for every fixed target value the loss of the optimal estimator concentrates around its conditional mean at least as quickly as a Gaussian with variance parameter $\sigma^2(y)$, and that these parameters have finite average; the Sibson version also requires a tilted moment condition on $\gamma^2(Y^*)$.
Editorial extensions
If this is right
- For every $\alpha \in (0,1)$, the excess risk is bounded by a divergence term with a constant that is explicit in $\alpha$ and in the average sub-Gaussian parameter, giving a tunable family of bounds.
- The Rényi bound recovers the known mutual-information bound in the limit $\alpha \to 1$, the Jensen–Shannon bound recovers it as $\alpha \to 0$, and the Jensen–Shannon limit $\alpha \to 1$ produces a Lautum-information bound.
- For bounded losses, all three theorems give corollaries with the explicit pre-factor $\|l\|_\infty/\sqrt{2}$, so the bounds are computable for finite-alphabet channels.
- In the numerical examples, at least one $\alpha$-parameterized bound is tighter than the mutual-information bound over a range of $\alpha$, and in the $q$-ary symmetric channel the advantage grows with alphabet size.
- Because $\sigma^2$ may depend on $Y$, the theorems cover losses such as $l(y,y')=\min\{|y-y'|,|y-c|\}$ whose natural sub-Gaussian parameter is $(\lvert y\rvert+c)^2/4$.
Reading between the lines
- The paper's approach suggests that any divergence with a Donsker–Varadhan-type variational representation can generate a corresponding excess-risk bound, so $f$-divergences and maximal leakage are natural candidates for the same treatment.
- Because the bounds are explicit in $\alpha$, one can optimize $\alpha$ for a given channel and loss; the pre-factor grows or shrinks with $\alpha$ while the divergence term moves in the opposite direction, so an interior optimum should exist in typical models.
- The target-dependent $\sigma^2(y)$ formulation aligns with estimation problems where uncertainty scales with the target magnitude, such as estimating unbounded signals under clipped losses; the Gaussian-channel examples are a first illustration of this regime.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the excess minimum risk L*_l(Y|Z) - L*_l(Y|X) for random vectors forming the Markov chain Y -> X -> Z. It derives upper bounds in terms of Rényi divergence (Theorem 1), α-Jensen-Shannon divergence (Theorem 2), and Sibson mutual information (Theorem 3), with the stated goal of generalizing the mutual-information bound of Györfi et al. (2023) and removing the constant-sub-Gaussian-parameter assumption made in prior work by Modak et al. and Aminian et al. Theorems 1 and 2 are proved by variational representations of the corresponding divergences; Theorem 3 uses the auxiliary-distribution method and a preliminary sub-Gaussian lemma (Lemma 6). The paper also gives three numerical examples showing that the Rényi and α-Jensen-Shannon bounds can be tighter than the mutual-information bound for some ranges of α. The central technical weakness is in the proof of Lemma 6 and hence Theorem 3, where an invalid exponential-moment inequality is used; the Rényi and α-Jensen-Shannon results do not depend on this step.
Significance. If Theorems 1 and 2 are correct, the paper makes a useful contribution: it extends the excess-risk bounds of Györfi et al. to nonconstant, target-dependent sub-Gaussian parameters and provides a family of α-parameterized bounds whose limits recover the mutual-information bound. The proofs of Theorems 1 and 2 appear sound, the α → 0 and α → 1 limits are correct consistency checks, and the numerical examples, while simple, illustrate the claimed tightening. The Sibson-mutual-information part (Theorem 3 and Corollary 3) is not established as stated because Lemma 6 relies on a false exponential-moment inequality; this is a load-bearing gap in one of the three advertised families of bounds. Since the remaining results are substantial and the flaw is localized to the Sibson section, the paper is worth pursuing after a major revision that either repairs the proof under genuinely justified assumptions or removes the unproven Sibson claim.
major comments (2)
- [Section 3.3, proof of Lemma 6, inequality (59)] The step leading to (59) replaces log E_{P_U*}[e^{(λ'^2/2) γ^2(U*)}] by (λ'^2/2) log E_{P_U*}[e^{γ^2(U*)}]. This uses the inequality log E[e^{tX}] ≤ t log E[e^X] for X = γ^2(U*) and t = λ'^2/2 ≥ 0. That inequality holds for 0 ≤ t ≤ 1 but is generally false for t > 1; for example, if X takes values 0 and 2 with equal probability and t = 2, then log E[e^{2X}] ≈ 3.33 > 2 log E[e^X] ≈ 2.87. Since λ' ranges over all real numbers in the Donsker–Varadhan application leading to (61), the subsequent quadratic-in-λ argument is not justified. Consequently Lemma 6, Theorem 3, and Corollary 3 are not established under the stated assumptions; a repair requires an additional strong condition that makes the exponential-moment comparison valid for all λ' or a different proof technique.
- [Section 3.3, Theorem 3 and Corollary 3] Because Theorem 3 is proved by applying Lemma 6 conditionally on Z = z and then integrating, the invalidity of Lemma 6 propagates directly to Theorem 3. Corollary 3 is a specialization of Theorem 3, so it inherits the same unsupported step. The abstract and introduction advertise bounds based on Sibson's mutual information; until this gap is fixed, that part of the claimed contribution should be either corrected or explicitly withdrawn.
minor comments (5)
- [Lemma 5] The statement that αX is |α|σ_X^2-sub-Gaussian is numerically incorrect; the correct statement is that αX is α^2σ_X^2-sub-Gaussian. Since the only scaling used in the paper is by −1, this error does not affect the main theorems, but it should be corrected.
- [Theorem 1 and Theorem 2 hypotheses] The assumptions state 'for all y ∈ R' and 'σ^2 : R → R', but Y takes values in R^p; these should read 'for all y ∈ R^p' and 'σ^2 : R^p → R'.
- [Lemma 2 opening sentence] The lemma says the random variables are 'defined on the same probability' but the word 'space' is missing; this typo should be fixed.
- [Remark 2 and Remark 4] The equalities D_KL(P_{X|Y,Z} || P_{X|Z} | P_{Y,Z}) = I(X;Y) - I(Z;Y) and the analogous α → 0 and α → 1 limits rely on the Markov-chain condition Y → X → Z; it would be helpful to state this explicitly at the point of use.
- [Section 4, Example 1] The text says that for q = 2, 3, 5 the input distributions are 'explicitly specified in the figure captions,' but the distributions appear in the subcaptions of Figure 1 in the main text rather than in the figure captions; the cross-reference should be corrected.
Circularity Check
No circularity: the Renyi, alpha-Jensen-Shannon, and Sibson bounds are each derived from variational characterizations and sub-Gaussian hypotheses, with the alpha-to-1 limits recovering the Gyorfi-Linder-Walk mutual-information bound only as a consistency check; self-citations are prior work being generalized, and no parameter is fitted to the target quantity.
full rationale
The derivation chain is self-contained and non-circular. Theorem 1 follows from Lemma 2, which is proved directly from the variational characterization of the Renyi divergence (Lemma 1, from Birrell et al. [19]) and the assumed conditional sub-Gaussianity of l(y, f(X)); the excess-risk difference enters only at the end via E[l(Y, f(X))] = L*_l(Y|X) and E[l(Y-bar, f(X-bar))] >= L*_l(Y|Z) (Eqs. (25)-(26)), which bound the target quantity rather than define it. Theorem 2 follows from Lemma 3 via the Donsker-Varadhan formula and the alpha-convex combination P^(alpha)_{X|Z,Y=y} = alpha P_{X|Z,Y=y} + (1-alpha) P_{X|Z}; Theorem 3 follows from Lemma 6 via the Sibson variational representation [10, 32] and the tilted auxiliary distribution of Definition 11, with the Sibson-optimal marginal P_{Y*|Z} solved from Eq. (8)/(71) rather than assumed to equal the bound. The alpha->1 and alpha->0 limits in Remarks 2-4 recover the mutual-information bound of [15, Theorem 3] as a consistency check; the theorem of [15] is never used as a proof input, so the self-citations (Linder co-authors [15], and [1] is a prior conference version of this work) are not load-bearing. No parameter is fitted to any datum: the numerical examples evaluate the closed-form bounds and compare them against the mutual-information bound (Figures 1-3). For the record, one flagged defect is a correctness issue, not circularity: in the proof of Lemma 6 (Section 3.3, inequality (59)), the step 'log E_{P_U*P_V}[exp(lambda' g(U*,V))] <= (lambda'^2/2) log E_{P_U*}[e^{gamma^2(U*)}]' requires log E[e^{t gamma^2}] <= t log E[e^{gamma^2}] with t = lambda'^2/2, which fails for t > 1 unless gamma^2(U*) is essentially bounded, so Theorem 3 is not established under the stated assumptions; Lemma 5's scaling 'alpha X is |alpha| sigma_X^2-sub-Gaussian' should be alpha^2 sigma_X^2. These defects do not reduce the claimed bounds to their inputs by construction, so the circularity score remains 0.
Assumptions & free parameters
assumptions (7)
- domain assumption Markov chain Y -> X -> Z and the data processing inequality for expected risk make the excess minimum risk nonnegative.
- standard math Variational characterization of Renyi divergence from Birrell et al. [19, Theorem 3.1].
- standard math Donsker-Varadhan variational formula for KL divergence.
- standard math Sibson mutual information can be written with the minimizing distribution P_U* via the density in Definition 6.
- standard math Hoeffding's lemma gives sub-Gaussianity of bounded loss functions.
- standard math The set of sub-Gaussian random variables is closed under scaling and addition, as stated in Lemma 5.
- ad hoc to paper The proof of Lemma 6 assumes log E[e^{t X}] <= t log E[e^{X}] for all t >= 0 with X = gamma^2(U*).
Cite this review
Pith. "Pith review of Bounds on the Excess Minimum Risk via Generalized Information Divergence Measures." pith.science (2026). https://pith.science/paper/242GMMDC
@misc{pith2026250524117,
author = {Pith},
title = {Pith review of: Bounds on the Excess Minimum Risk via Generalized Information Divergence Measures},
year = {2026},
howpublished = {\url{https://pith.science/paper/242GMMDC}},
note = {Machine review of arXiv:2505.24117}
}
abstract
Given finite-dimensional random vectors $Y$, $X$, and $Z$ that form a Markov chain in that order (i.e., $Y \to X \to Z$), we derive upper bounds on the excess minimum risk using generalized information divergence measures. Here, $Y$ is a target vector to be estimated from an observed feature vector $X$ or its stochastically degraded version $Z$. The excess minimum risk is defined as the difference between the minimum expected loss in estimating $Y$ from $X$ and from $Z$. We present a family of bounds that generalize the mutual information based bound of Gy\"orfi et al. (2023), using the R\'enyi and $\alpha$-Jensen-Shannon divergences, as well as Sibson's mutual information. Our bounds are similar to those developed by Modak et al. (2021) and Aminian et al. (2024) for the generalization error of learning algorithms. However, unlike these works, our bounds do not require the sub-Gaussian parameter to be constant and therefore apply to a broader class of joint distributions over $Y$, $X$, and $Z$. We also provide numerical examples under both constant and non-constant sub-Gaussianity assumptions, illustrating that our generalized divergence based bounds can be tighter than the one based on mutual information for certain regimes of the parameter $\alpha$.
Figures
Reference graph
Works this paper leans on
-
[1]
Bounding excess minimum risk via r´ enyi’s divergence,
A. Omanwar, F. Alajaji, and T. Linder, “Bounding excess minimum risk via r´ enyi’s divergence,” in 2024 International Symposium on Information Theory and Its Appli- cations (ISITA). IEEE, 2024, pp. 59–63
work page 2024
-
[2]
On measures of entropy and information,
A. R´ enyi, “On measures of entropy and information,” in the Fourth Berkeley Sym- posium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics , vol. 4. University of California Press, 1961, pp. 547–562
work page 1961
-
[3]
On a generalization of the Jensen–Shannon divergence and the Jensen– Shannon centroid,
F. Nielsen, “On a generalization of the Jensen–Shannon divergence and the Jensen– Shannon centroid,” Entropy, vol. 22, no. 2, p. 221, 2020
work page 2020
-
[4]
Divergence measures based on the shannon entropy,
J. Lin, “Divergence measures based on the shannon entropy,” IEEE Transactions on Information Theory, vol. 37, no. 1, pp. 145–151, 1991
work page 1991
-
[5]
R. Sibson, “Information radius,” Zeitschrift f¨ ur Wahrscheinlichkeitstheorie und ver- wandte Gebiete , vol. 14, no. 2, pp. 149–160, 1969
work page 1969
-
[6]
Generalized cutoff rates and renyi’s information measures,
I. Csiszar, “Generalized cutoff rates and renyi’s information measures,” IEEE Trans- actions on Information Theory , vol. 41, no. 1, pp. 26–34, 1995
work page 1995
-
[7]
Information-theoretic analysis of generalization capability of learning algorithms,
A. Xu and M. Raginsky, “Information-theoretic analysis of generalization capability of learning algorithms,” Advances in neural information processing systems , vol. 30, 2017
work page 2017
-
[8]
Tightening mutual information-based bounds on generalization error,
Y. Bu, S. Zou, and V. V. Veeravalli, “Tightening mutual information-based bounds on generalization error,” IEEE Journal on Selected Areas in Information Theory , vol. 1, no. 1, pp. 121–130, 2020
work page 2020
Show all 36 references
-
[9]
Robust generalization via f- mutual in- formation,
A. R. Esposito, M. Gastpar, and I. Issa, “Robust generalization via f- mutual in- formation,” in 2020 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2020, pp. 2723–2728
2020
-
[10]
Variational characterizations of sibson’s α-mutual information,
——, “Variational characterizations of sibson’s α-mutual information,” in 2024 IEEE International Symposium on Information Theory (ISIT). IEEE, 2024, pp. 2110–2115
2024
-
[11]
Generalization error bounds via r´ enyi-, f-divergences and maximal leakage,
——, “Generalization error bounds via r´ enyi-, f-divergences and maximal leakage,” IEEE Transactions on Information Theory , vol. 67, no. 8, pp. 4986–5004, 2021
2021
-
[12]
R´ enyi divergence based bounds on generalization error,
E. Modak, H. Asnani, and V. M. Prabhakaran, “R´ enyi divergence based bounds on generalization error,” in 2021 IEEE Information Theory Workshop (ITW) . IEEE, 2021, pp. 1–6. 23
2021
-
[13]
Understanding estimation and generalization error of generative adversarial networks,
K. Ji, Y. Zhou, and Y. Liang, “Understanding estimation and generalization error of generative adversarial networks,” IEEE transactions on Information Theory , vol. 67, no. 5, pp. 3114–3129, 2021
2021
-
[14]
Minimum excess risk in bayesian learning,
A. Xu and M. Raginsky, “Minimum excess risk in bayesian learning,” IEEE Trans- actions on Information Theory , vol. 68, no. 12, pp. 7935–7955, 2022
2022
-
[15]
Lossless transformations and excess risk bounds in statistical inference,
L. Gy¨ orfi, T. Linder, and H. Walk, “Lossless transformations and excess risk bounds in statistical inference,” Entropy, vol. 25, no. 10, p. 1394, 2023
2023
-
[16]
Information-theoretic analysis of min- imax excess risk,
H. Hafez-Kolahi, B. Moniri, and S. Kasaei, “Information-theoretic analysis of min- imax excess risk,” IEEE Transactions on Information Theory , vol. 69, no. 7, pp. 4659–4674, 2023
2023
-
[17]
Information- theoretic characterizations of generalization error for the gibbs algorithm,
G. Aminian, Y. Bu, L. Toni, M. R. Rodrigues, and G. W. Wornell, “Information- theoretic characterizations of generalization error for the gibbs algorithm,” IEEE Transactions on Information Theory , vol. 70, no. 1, pp. 632–655, 2023
2023
-
[18]
Learning algorithm general- ization error bounds via auxiliary distributions,
G. Aminian, S. Masiha, L. Toni, and M. R. Rodrigues, “Learning algorithm general- ization error bounds via auxiliary distributions,” IEEE Journal on Selected Areas in Information Theory, 2024
2024
-
[19]
Variational representations and neural network estimation of r´ enyi divergences,
J. Birrell, P. Dupuis, M. A. Katsoulakis, L. Rey-Bellet, and J. Wang, “Variational representations and neural network estimation of r´ enyi divergences,” SIAM Journal on Mathematics of Data Science , vol. 3, no. 4, pp. 1093–1116, 2021
2021
-
[20]
Robust bounds on risk-sensitive functionals via r´ enyi divergence,
R. Atar, K. Chowdhary, and P. Dupuis, “Robust bounds on risk-sensitive functionals via r´ enyi divergence,” SIAM/ASA Journal on Uncertainty Quantification , vol. 3, no. 1, pp. 18–33, 2015
2015
-
[21]
A variational characterization of r´ enyi divergences,
V. Anantharam, “A variational characterization of r´ enyi divergences,” IEEE Trans- actions on Information Theory , vol. 64, no. 11, pp. 6979–6989, 2018
2018
-
[22]
Information-type measures of difference of probability distributions and indirect observations,
I. Csisz´ ar, “Information-type measures of difference of probability distributions and indirect observations,” Studia Sci. Math. Hungarica , vol. 2, pp. 299–318, 1967
1967
-
[23]
Divergence measures based on the shannon entropy,
J. Lin, “Divergence measures based on the shannon entropy,” IEEE Transactions on Information theory, vol. 37, no. 1, pp. 145–151, 1991
1991
-
[24]
Pac-bayesian model averaging,
D. A. McAllester, “Pac-bayesian model averaging,” in Twelfth Annual Conference on Computational Learning Theory (COLT). ACM, 1999, pp. 164–170
1999
-
[25]
Properties of variational approximations of gibbs posteriors,
P. Alquier, J. Ridgway, and N. Chopin, “Properties of variational approximations of gibbs posteriors,” Journal of Machine Learning Research , vol. 17, no. 1, pp. 8372– 8414, 2016
2016
-
[26]
Generalization bounds via wasserstein distance based algorith- mic stability,
A. Lopez and V. Jog, “Generalization bounds via wasserstein distance based algorith- mic stability,” in 37th International Conference on Machine Learning (ICML) , 2020, pp. 6326–6335
2020
-
[27]
Addressing gan training instabil- ities via tunable classification losses,
M. Welfert, G. R. Kurri, K. Otstot, and L. Sankar, “Addressing gan training instabil- ities via tunable classification losses,” IEEE Journal on Selected Areas in Information Theory, 2024
2024
-
[28]
Generative adversarial nets,
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural in- formation processing systems, vol. 27, 2014
2014
-
[29]
Asymptotic evaluation of certain markov process expectations for large time. iv,
M. Donsker and S. Varadhan, “Asymptotic evaluation of certain markov process expectations for large time. iv,” Communications on Pure and Applied Mathematics , vol. 36, no. 2, pp. 183–212, 1983
1983
-
[30]
R´ enyi divergence and kullback-leibler divergence,
T. Van Erven and P. Harremos, “R´ enyi divergence and kullback-leibler divergence,” IEEE Transactions on Information Theory , vol. 60, no. 7, pp. 3797–3820, 2014
2014
-
[31]
α-mutual information,
S. Verd´ u, “α-mutual information,” in 2015 Information Theory and Applications Workshop (ITA). IEEE, 2015, pp. 1–6. 24
2015
-
[32]
Sibson’s α-mutual information and its variational representations,
A. R. Esposito, M. Gastpar, and I. Issa, “Sibson’s α-mutual information and its variational representations,” arXiv preprint arXiv:2405.08352 , 2024
2024 arXiv
-
[33]
Probability inequalities for sums of bounded random variables,
W. Hoeffding, “Probability inequalities for sums of bounded random variables,” The collected works of Wassily Hoeffding , pp. 409–426, 1994
1994
-
[34]
Lautum information,
D. P. Palomar and S. Verd´ u, “Lautum information,”IEEE transactions on informa- tion theory, vol. 54, no. 3, pp. 964–975, 2008
2008
-
[35]
Sub-gaussian random variables,
V. V. Buldygin and Y. V. Kozachenko, “Sub-gaussian random variables,” Ukrainian Mathematical Journal, vol. 32, pp. 483–489, 1980
1980
-
[36]
Subgaussian random variables: An expository note,
O. Rivasplata, “Subgaussian random variables: An expository note,” Internet publi- cation, PDF, vol. 5, 2012. 25
2012
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.