Pith. sign in

REVIEW 1 major objections 4 minor 89 references

Deviation Inequalities for R\'{e}nyi Divergence Estimators via Variational Expression

T0 review · 1 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper proves exponential deviation inequalities for Gaussian-smoothed plug-in and neural estimators of Rényi divergences, with tails of the form $2e^{-(nz^2\wedge nz)}$, without requiring compact support or densities bounded away from

desk verdict The compact-support results are solid and worth refereeing, but the sub-Gaussian one-sample KL claim has a real conditioning-bias gap that needs fixing before it is taken as established. read the letter →

arxiv 2508.09382 v3 pith:CN5GPRVS submitted 2025-08-12 cs.IT math.IT

classification cs.ITmath.IT MSC 62G0562G2094A1760E15
keywords RényidivergenceKLdeviationinequalitysmoothedplug-inestimatorneuralempiricalprocessdifferentialprivacyauditvariationalexpression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Estimating a Rényi divergence from samples is a workhorse task in information theory, statistics, and machine learning, but the probability law of the estimation error is poorly understood: prior tail bounds required compactly supported distributions with densities bounded away from zero. This paper removes those restrictions. Using the variational (dual) expressions for KL and Rényi divergences, the authors bound the estimation error by the supremum of an empirical process over a Gaussian-smoothed function class, then apply sharp exponential concentration inequalities to obtain explicit tail bounds: the probability that the error exceeds a constant times $n^{-1/2}+z$ is at most $2e^{-(nz^2\wedge nz)}$. The bounds hold for compactly supported measures (both smoothed plug-in and neural estimators, at every Rényi order) and, for the smoothed KL estimator, for sub-Gaussian measures in the one-sample setting. As an application, the neural estimator becomes a test statistic for auditing Rényi differential privacy with finite-sample type I and type II error guarantees.

What carries the argument

The central objects are the variational (dual) expressions for the divergences — $D_{\mathrm{KL}}(\mu\|\nu)=\sup_f[E_\mu[f]-\log E_\nu[e^f]]$ and its affine variant, plus the analogous order-$\alpha$ expression — whose suprema are attained at the log-density ratio $f^\star=\log(d\mu/d\nu)$. The workhorse move is to bound the estimator's error by suprema of empirical processes indexed by Gaussian-smoothed function classes such as $\{f*\varphi_\sigma : |f(x)|\le \tfrac12 b(1+\|x\|)\}$ and $\{e^f*\varphi_\sigma\}$, which have integrable envelopes for compactly supported (and truncated sub-Gaussian) data. Concentration comes from exponential inequalities for empirical-process suprema — one for u

What would settle it

Simulate the two-sample smoothed plug-in estimator with $\mu$ uniform on the unit ball $B_d(1)$, $\nu$ a fixed translate of $\mu$, fixed $d$ (say $d=2$) and $\sigma=1$. For $n=10^3$ to $10^6$, draw many independent sample pairs, compute the deviation $|D_{\mathrm{KL}}(\hat\mu_n*\gamma_\sigma \| \hat\nu_n*\gamma_\sigma) - D_{\mathrm{KL}}(\mu*\gamma_\sigma \| \nu*\gamma_\sigma)|$, and check that the empirical exceedance probability over the claimed threshold $\Omega_{r,d,\sigma}(n^{-1/2}+z)$ stays at or below $2e^{-(nz^2\wedge nz)}$ across a grid of $z>0$. A single $(\mu,\nu,r,\sigma,z)$ for whi

Watch

Extended reading notes

Core claim

Exponential deviation inequalities are proved for Gaussian-smoothed plug-in and neural estimators of KL and Rényi divergences, with explicit constants. The tail shape is $P(|D-\widehat D_n|\ge\Omega_{r,d,\sigma}(n^{-1/2}+z))\le 2e^{-(nz^2\wedge nz)}$: Theorem 1 for the smoothed KL plug-in estimator on compactly supported measures (with a one-sample version for sub-Gaussian measures), Theorem 3 for Rényi order $\alpha$. Neural (variational) estimators obey the same inequality (Theorem 4), with an added $(M^{2\alpha}+M^{2|\alpha-1|})\delta$ term for the sup-norm error of approximating the optimal function $f^\star=\log(d\mu/d\nu)$ by the neural class. Applications: one-sided concentration from

Load-bearing premise

The neural-estimator and privacy-audit results stand or fall on the premise that the true log-density ratio $f^\star$ can be uniformly approximated by the chosen neural network class on the support of the data, with an approximation error $\delta$ that vanishes as the sample size grows; the smoothed plug-in results rest on the parallel premise that $f^\star$ lies almost surely in a Gaussian-smoothed function class with an integrable envelope.

Editorial extensions

If this is right

  • Smoothed plug-in estimators of KL and Rényi divergences come with finite-sample exponential tail bounds for compactly supported distributions, a regime where earlier results were limited to finitely many atoms, compact support, or densities bounded away from zero.
  • The one-sided concentration bound from the expectation makes the smoothed plug-in divergence a valid tool in random-coding existence arguments over Gaussian channels: with high probability the empirical-codebook divergence is within $O(n^{-1/2})$ of its expectation, the ingredient wiretap and covert-communication proofs need.
  • The Rényi-DP audit test attains type I error at most $2e^{-n^\tau}$ together with type II error at most $2e^{-\vartheta}$ for an explicit exponent $\vartheta$ — the first finite-sample guarantee of this kind, where earlier audits were either partial or only asymptotic.
  • Under Lipschitz and bounded-ratio density assumptions, the smoothing bias can be traded against estimation variance (Corollary 1), giving a deviation bound from the unsmoothed divergence at rate $n^{-s/(d+2s+4)}$.
  • The statistical version of the unbounded-envelope concentration inequality (Lemma 1) is a standalone tool for constructing confidence intervals for empirical-process suprema when the underlying distribution is unknown.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The dimension dependence in the constants is exponential in $d$ (order $d^{6d+9}(9e^2)^d$ and $\sigma^{-d/2}$ for the smoothed classes); the authors flag this as sub-optimal. A sharper entropy bound on a smoother function class than the Hölder class would be the natural path to making these inequalities useful where $d$ grows with $n$ — e.g., wiretap codes whose codeword length scales with codeboo
  • Lemma 2 is stated for a general smoothing kernel $\Phi$, not only Gaussians; the same route should yield exponential deviation inequalities for plug-in estimators smoothed with any kernel whose convolution keeps $f^\star$ in a class with integrable envelope and controlled covering entropy — a directly testable extension (e.g., compactly supported or Student-$t$ kernels).
  • Because Rényi divergence is non-decreasing in its order, an auditor could run the neural test at several orders $\alpha$, or at large $\alpha$ to approximate an $\epsilon$-DP audit; the explicit type II exponent $\vartheta$ in (57) says how many output samples are needed to separate a mechanism at divergence level $\epsilon$ from one at $\epsilon+\Delta$. This multi-order audit strategy is not in
  • The two-sample sub-Gaussian case is left open for Rényi orders $\alpha \ne 1$, and for KL the two-sample sub-Gaussian bound is not exponential: the truncation argument inflates the envelope norms exponentially in the truncation level. Closing that gap would require a genuinely different function-class construction — a concrete open problem the paper states.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The paper derives exponential deviation inequalities for Rényi divergence estimators. For Gaussian-smoothed plug-in estimators, Theorem 1(i) gives a two-sample bound for compactly supported distributions, Theorem 1(ii) gives a one-sample bound for sub-Gaussian distributions, and Theorem 3 extends the compact-support bound to Rényi order α. Theorem 2 provides a Rademacher-complexity version, and Lemma 1 is a statistical Fuk–Nagaev inequality for unbounded envelopes. For neural estimators, Theorem 4 bounds |D̂_{α,G}(μ̂_n,ν̂_n) − D_α(μ‖ν)| uniformly over pairs in P̄_{M,ρ}(G,δ), and Proposition 2 converts this into finite-sample type I/II guarantees for Rényi-DP auditing. The proofs use variational representations, empirical-process concentration inequalities, and entropy estimates for Gaussian-smoothed Hölder classes.

Significance. If the results hold, this is a useful contribution: it provides the first exponential tail bounds for smoothed plug-in and neural Rényi-divergence estimators under unbounded (sub-Gaussian) marginals and for neural estimators, with explicit dependence on dimension and distributional parameters. Lemma 1 is a self-contained extension of Talagrand/Bartlett–Bousquet–Mendelson inequalities to unbounded envelopes. The entropy estimates in Lemma 3 and the detailed constant chasing are strengths. The DP-audit application gives finite-sample guarantees that are absent from earlier asymptotic treatments. However, one load-bearing sub-Gaussian claim in Theorem 1(ii) has a conditioning-bias gap that needs to be fixed before the advertised extension is established.

major comments (1)
  1. [§4.1, Theorem 1(ii), after Eq. (40)] The proof conditions on E_n(t), replaces μ by the truncated law μ̃, and then invokes the one-sample bound (28). This does not control the deviation from D_KL(μ*γ_σ‖ν*γ_σ). If (28) is applied with underlying distribution μ̃, the target is D_KL(μ̃*γ_σ‖ν*γ_σ); if one keeps the centering at μ, the samples are no longer i.i.d. μ under E_n(t). A missing term such as |D_KL(μ̃*γ_σ‖ν*γ_σ) − D_KL(μ*γ_σ‖ν*γ_σ)|, or equivalently sup_{f∈F_t}|E_{μ̃}f − E_μ f|, is not bounded. Since the envelope in (38) is quadratic in ‖x‖ with t = T_{n,p,κ,L}, this bias is generally nonzero. Sub-Gaussianity supplies tail estimates that could bound it, but no such estimate appears in the paper. Thus Theorem 1(ii) is not established, and the advertised sub-Gaussian extension is load-bearing for the contribution.
minor comments (4)
  1. [§4.5, Proposition 1(17b)] For α>1 the proof after Eq. (50) yields an upper bound of order (α−1)M^{2α+1}+αM^{2α}, not M^{2α+1}+α/(α−1)M^{2α}. The displayed bound is smaller than what is proved when α>2. This statement is not used in the main deviation theorems, so it does not affect the central claims, but it should be corrected.
  2. [§3.3, Assumption 3] The notation P_{M,β}(G_n,δ_n) is ambiguous; the class P̄_{M,ρ}(G,δ) defined earlier depends on ρ, not β. Please clarify whether β is a typo for ρ or a separate smoothness parameter.
  3. [§4.1, Part (ii)] In the one-sample proof, the display involving f*_{μ̂_n*φ_σ,ν̂_n*φ_σ} should use ν (the known measure), not ν̂_n. This is a notational slip that can confuse the reader.
  4. [Theorem 1(ii), Eq. (10)] The exponent in the third term, n^{pτ/(d+12)}/log(n+1), yields a subexponential rather than a pure exponential decay for fixed p,τ. The word 'exponential' is used loosely for the whole bound; consider rephrasing or deriving a cleaner lower bound.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the deviation inequalities are forward empirical-process bounds with explicit constants; neural approximation and DP gaps appear as declared assumptions, not fitted inputs.

full rationale

The paper's main inequalities (Theorems 1-4, Proposition 2) are derived from the Donsker-Varadhan/Renyi variational formulas (Eqs. (1)-(2)) by bounding the difference between empirical and population expectations by suprema of empirical processes. The constants Ω, ξ, λ, c are explicit functions of r,d,σ,M,α,β and are not fitted to data. Lemma 2 is the engine: it converts a deviation of the smoothed plug-in KL estimator into a sum of empirical-process suprema, then applies Talagrand or Fuk-Nagaev concentration; no target quantity is re-used as an input. Proposition 1 (the smoothing approximation bound) is proved in Section 4.5 via Taylor expansion, not imported as an unproved self-citation, even though it generalizes a prior lemma. Theorem 4 is explicitly conditional on the neural class G approximating f* within δ: the class P_bar_{M,rho}(G,δ) is defined as those pairs for which such approximation holds, and δ appears as an assumed approximation error, not as a fitted parameter; the cited neural approximation results [6,44,75] support non-vacuity but are not needed to make the theorem's conditional statement true. Proposition 2's DP audit bound is just Theorem 4 applied under the null/alternative; its type-II exponent depends on the unknown gap D_α(μ1||ν1)−ε, which is a property of the hypotheses, not a circular construct. The truncation argument in Theorem 1(ii) raises a possible correctness concern (conditioning on E_n(t) changes the law to tilde μ while the target remains D_KL(μ*γ||ν*γ), potentially missing a bias), but this is an omitted-bias/validity issue, not a circularity: it does not make any claimed prediction equivalent by construction to an input. No fitted-input-called-prediction, self-definitional, or self-citation-load-bearing steps were found.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The paper's results are theorems derived from standard tools (variational formulas, concentration inequalities, entropy bounds); no parameters are fitted to data. The main assumptions are domain restrictions (bounded density ratios, Lipschitz regularity, sub-Gaussian moment conditions) and the existence of a neural class approximating the variational optimizer up to a user-supplied delta.

assumptions (7)
  • standard math Variational expressions for KL and Renyi divergences (Eqs. (1) and (2)) are valid and the suprema are attained at f* = log dmu/dnu.
    Invoked at the start of the proofs of Lemma 2 and Theorem 3; cited to Donsker-Varadhan and Birrell et al., and Lemma 7 proves a variant.
  • standard math Talagrand's concentration inequality (Theorem 6) and Fuk-Nagaev inequality (Theorem 7) hold with the stated constants.
    Used in Lemma 2 and Theorems 1 and 3 to convert empirical process bounds into deviation inequalities.
  • standard math Covering entropy of Holder balls is bounded as in Theorem 8, with the explicit prefactor derived in Appendix B.
    Used in Lemma 3 to bound entropy of Gaussian-smoothed function classes.
  • domain assumption The neural class G satisfies Assumption 2 (uniform L2 covering entropy bound) for beta > d/2, and the distribution class P_bar_{M,rho}(G,delta) is non-empty with controllable delta.
    Load-bearing for Theorem 4 and Proposition 2; the paper cites prior approximation results for examples but leaves delta open.
  • domain assumption For Corollary 1, the densities p_mu and p_nu belong to the Lipschitz class Lip_{s,1,M}(rho) and have bounded ratios M.
    Needed for Proposition 1, the approximation bound between smoothed and unsmoothed Renyi divergence.
  • domain assumption In Theorem 1(ii), each X from mu is sigma^2-sub-Gaussian and satisfies ||X||_{psi_p} <= L for some p >= 2.
    Used for the truncation argument and the tail bound of max_i ||X_i||.
  • domain assumption For Lemma 2 in general, the optimizer f* of the smoothed variational problem lies a.s. in the chosen function class F_n with integrable envelope.
    This is the step converting the deviation to an empirical process supremum; the paper verifies it for compact support and for the truncated sub-Gaussian case.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deviation Inequalities for R\'{e}nyi Divergence Estimators via Variational Expression." pith.science (2026). https://pith.science/paper/CN5GPRVS

@misc{pith2026250809382,
  author       = {Pith},
  title        = {Pith review of: Deviation Inequalities for R\'enyi Divergence Estimators via Variational Expression},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CN5GPRVS}},
  note         = {Machine review of arXiv:2508.09382}
}
read the original abstract

R\'enyi divergences play a pivotal role in information theory, statistics, and machine learning. While several estimators of these divergences have been proposed in the literature with their consistency properties established and minimax convergence rates quantified, existing accounts of probabilistic bounds governing the estimation error are relatively underdeveloped. Here, we make progress in this regard by establishing exponential deviation inequalities for smoothed plug-in estimators and neural estimators by relating the error to an appropriate empirical process and leveraging tools from empirical process theory. In particular, our approach does not require the underlying distributions to be compactly supported or have densities bounded away from zero, an assumption prevalent in existing results. The deviation inequality also leads to a one-sided concentration bound from the expectation, which is useful in random-coding arguments over continuous alphabets in information theory with potential applications to physical-layer security. As another concrete application, we consider a hypothesis testing framework for auditing R\'{e}nyi differential privacy using the neural estimator as a test statistic and obtain non-asymptotic performance guarantees for such a test.

Figures

Figures reproduced from arXiv: 2508.09382 by the authors.

Figure 1
Figure 1. Random coding over a Gaussian channel where each codeword Xi of a random codebook {Xi} n i=1 of size n is a vector of length d generated i.i.d. according to µ. A power constraint on each input codeword naturally translates to a bound on its Euclidean norm. where c¯d,r := (3ed) 3(d+2)√ dr(r + 1) d 2 +3 and cd,s := R Rd ∥z∥ s φ1(z)dz. The proof of Corollary 1 (see Section 4.4) relies on the fact that the difference be… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 77 canonical work pages

  1. [1]

    Acharya, S

    J. Acharya, S. Bhadane, P. Indyk, and Z. Sun. Estimating entropy of distributions in constant space. InAdvances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019

  2. [2]

    Adamczak

    R. Adamczak. A tail inequality for suprema of unbounded empirical processes with applications to Markov chains.Electronic Journal of Probability, 13:1000 – 1034, 2008

  3. [3]

    Adamczak

    R. Adamczak. A few remarks on the operator norm of random Toeplitz matrices.Journal of Theoretical Proba- bility, 23:85 – 108, 04 2010

  4. [4]

    R. Agrawal. Finite-sample concentration of the multinomial in relative entropy.IEEE Transactions on Informa- tion Theory, 66(10):6297–6302, 2020

  5. [5]

    Antos and I

    A. Antos and I. Kontoyiannis. Convergence properties of functional estimates for discrete distributions.Random Structures & Algorithms, 19(3-4):163–193, 2001

  6. [6]

    A. R. Barron. Universal approximation bounds for superpositions of a sigmoidal function.IEEE Transactions on Information Theory, 39(3):930–945, May 1993

  7. [7]

    P. L. Bartlett, O. Bousquet, and S. Mendelson. Local Rademacher complexities.The Annals of Statistics, 33(4):1497 – 1537, 2005

  8. [8]

    M. I. Belghazi, A. Baratin, S. Rajeshwar, S. Ozair, Y. Bengio, A. Courville, and D. Hjelm. Mutual information neural estimation. InProceedings of the 35th International Conference on Machine Learning, volume 80, pages 531–540, Stockholm Sweden, Jul. 2018. 44 S. SREEKUMAR AND K. KATO

Show all 89 references
  1. [9]

    T. B. Berrett and R. J. Samworth. Efficient functional estimation and the super-oracle phenomenon.The Annals of Statistics, 51(2):668 – 690, 2023

  2. [10]

    T. B. Berrett, R. J. Samworth, and M. Yuan. Efficient multivariate entropy estimation viak-nearest neighbour distances.The Annals of Statistics, 47(1):288–318, Feb. 2019

  3. [11]

    P. J. Bickel and Y. Ritov. Estimating integrated squared density derivatives: Sharp best order of convergence estimates.Sankhy¯ a: The Indian Journal of Statistics, Series A (1961-2002), 50(3):381–393, 1988

  4. [12]

    Birge and P

    L. Birge and P. Massart. Estimation of integral functionals of a density.The Annals of Statistics, 23(1):11 – 29, 1995

  5. [13]

    Birrell, P

    J. Birrell, P. Dupuis, M. A. Katsoulakis, L. Rey-Bellet, and J. Wang. Variational representations and neural network estimation of R´ enyi divergences.SIAM Journal on Mathematics of Data Science, 3(4):1093–1116, 2021

  6. [14]

    Block, Z

    A. Block, Z. Jia, Y. Polyanskiy, and A. Rakhlin. Rate of convergence of the smoothed empirical Wasserstein distance.Annales de l’Institut Henri Poincar´ e, to appear, 2025

  7. [15]

    Bousquet

    O. Bousquet. A Bennett concentration inequality and its application to suprema of empirical processes.Comptes Rendus Mathematique, 334(6):495–500, 2002

  8. [16]

    Y. Bu, S. Zou, Y. Liang, and V. V. Veeravalli. Estimation of KL divergence: optimal minimax rate.IEEE Transactions on Information Theory, 64(4):2648–2674, 2018

  9. [17]

    Bulinski and D

    A. Bulinski and D. Dimitrov. Statistical estimation of the Kullback-Leibler divergence.Mathematics, 9(5):1–36, March 2021

  10. [18]

    H. Cai, S. Kulkarni, and S. Verdu. Universal divergence estimation for finite-alphabet sources.IEEE Transactions on Information Theory, 52(8):3456–3475, 2006

  11. [19]

    T. M. Cover and J. A. Thomas.Elements of Information Theory. NewYork: Wiley, 1991

  12. [20]

    G. A. Darbellay and I. Vajda. Estimation of the information by an adaptive partitioning of the observation space.IEEE Transactions on Information Theory, 45(4):1315–1321, 1999

  13. [21]

    Delattre and N

    S. Delattre and N. Fournier. On the Kozachenko–Leonenko entropy estimator.Journal of Statistical Planning and Inference, 185:69–93, 2017

  14. [22]

    R. A. DeVore and G. G. Lorentz.Constructive Approximation. Springer Berlin, Heidelberg, 1993

  15. [23]

    Domingo-Enrich and Y

    C. Domingo-Enrich and Y. Mroueh. Auditing differential privacy in high dimensions with the kernel quantum R´ enyi divergence.arXiv:2205.13941, 2022

  16. [24]

    M. D. Donsker and S. R. S. Varadhan. Asymptotic evaluation of certain Markov process expectations for large time, i.Communications on Pure and Applied Mathematics, 28(1):1–47, 1975

  17. [25]

    Dwork, K

    C. Dwork, K. Kenthapadi, F. McSherry, I. Mironov, and M. Naor. Our data, ourselves: privacy via distributed noise generation. InAdvances in Cryptology - EUROCRYPT 2006, pages 486–503. Springer Berlin Heidelberg, 2006

  18. [26]

    Dwork, F

    C. Dwork, F. McSherry, K. Nissim, and A. Smith. Calibrating noise to sensitivity in private data analysis. In Proceedings of the Third Theory of Cryptography Conference, pages 265–284. Springer Berlin Heidelberg, 2006

  19. [27]

    Elbr¨ achter, D

    D. Elbr¨ achter, D. Perekrestenko, P. Grohs, and H. B¨ olcskei. Deep neural network approximation theory.IEEE Transactions on Information Theory, 67(5):2581–2623, 2021

  20. [28]

    G. B. Folland. Remainder estimates in Taylor’s theorem.The American Mathematical Monthly, 97(3):233–235, 1990

  21. [29]

    D. K. Fuk and S. V. Nagaev. Probability inequalities for sums of independent random variables.Theory of Probability & Its Applications, 16(4):643–660, 1971

  22. [30]

    W. Gao, S. Oh, and P. Viswanath. Demystifying fixedk-nearest neighbor information estimators.IEEE Trans- actions on Information Theory, 64(8):5629–5661, 2018

  23. [31]

    Goldfeld, P

    Z. Goldfeld, P. Cuff, and H. H. Permuter. Wiretap channels with random states non-causally available at the encoder.IEEE Transactions on Information Theory, 66(3):1497–1519, 2020

  24. [32]

    Goldfeld, K

    Z. Goldfeld, K. Greenewald, J. Niles-Weed, and Y. Polyanskiy. Convergence of smoothed empirical measures with applications to entropy estimation.IEEE Transactions on Information Theory, 66(7):4368–4391, Jul. 2020

  25. [33]

    M. N. Goria, N. N. Leonenko, V. V. Mergel, and P. L. N. Inverardi. A new class of random vector entropy estimators and its applications in testing statistical hypotheses.Journal of Nonparametric Statistics, 17(3):277– 297, 2005

  26. [34]

    F. R. Guo and T. S. Richardson. Chernoff-type concentration of empirical probabilities in relative entropy.IEEE Transactions on Information Theory, 67(1):549–558, 2021

  27. [35]

    Haje Hussein and Y

    F. Haje Hussein and Y. Golubev. On entropy estimation by m-spacing method.Journal of Mathematical Sciences, 163(3), 2009

  28. [36]

    Y. Han, J. Jiao, and T. Weissman. Minimax estimation of divergences between discrete distributions.IEEE Journal on Selected Areas in Information Theory, 1(3):814–823, 2020

  29. [37]

    Y. Han, J. Jiao, T. Weissman, and Y. Wu. Optimal rates of entropy estimation over Lipschitz balls.The Annals of Statistics, 48(6):3228–3250, Dec. 2020. DEVIATION INEQUALITIES FOR R ´ENYI DIVERGENCE ESTIMATORS 45

  30. [38]

    M. M. Hossain, A. Wisler, and K. R. Moon. Nonparametric estimation of non-smooth divergences. InProceed- ings of the 33rd ACM International Conference on Information and Knowledge Management, CIKM ’24, page 3787–3791, New York, NY, USA, 2024. Association for Computing Machinery

  31. [39]

    Jagielski, J

    M. Jagielski, J. Ullman, and A. Oprea. Auditing differentially private machine learning: How private is private SGD? InProceedings of Advances in Neural Information Processing Systems, volume 33, pages 22205–22216, 2020

  32. [40]

    J. Jiao, K. Venkat, Y. Han, and T. Weissman. Minimax estimation of functionals of discrete distributions.IEEE Transactions on Information Theory, 61(5):2835–2885, 2015

  33. [41]

    C. Jin, P. Netrapalli, R. Ge, S. M. Kakade, and M. I. Jordan. A short note on concentration inequalities for random vectors with subGaussian norm.arXiv:1902.03736, Feb. 2019

  34. [42]

    Kandasamy, A

    K. Kandasamy, A. Krishnamurthy, B. Poczos, L. Wasserman, and J. M. Robins. Nonparametric von Mises estimators for entropies, divergences and mutual informations. InProceedings of Advances in Neural Information Processing Systems, volume 28, pages 397–405, Montr´ eal, Canada, Dec. 2015

  35. [43]

    Kerkyacharian and D

    G. Kerkyacharian and D. Picard. Estimating nonquadratic functionals of a density using Haar wavelets.The Annals of Statistics, 24(2):485 – 507, 1996

  36. [44]

    J. M. Klusowski and A. R. Barron. Approximation by combinations of ReLU and squared ReLU ridge functions withℓ 1 andℓ 0 controls.IEEE Transactions on Information Theory, 64(12):7649–7656, Oct. 2018

  37. [45]

    Kontoyiannis and M

    I. Kontoyiannis and M. Skoularidou. Estimating the directed information and testing for causality.IEEE Trans- actions on Information Theory, 62(11):6053–6067, 2016

  38. [46]

    L. F. Kozachenko and N. N. Leonenko. Sample estimate of the entropy of a random vector.Problems of Infor- mation Transmission, 23(2):9–16, 1987

  39. [47]

    Kraskov, H

    A. Kraskov, H. St¨ ogbauer, and P. Grassberger. Estimating mutual information.Physical Review E, 69, Jun 2004

  40. [48]

    Krishnamurthy, K

    A. Krishnamurthy, K. Kandasamy, B. P´ oczos, and L. Wasserman. Nonparametric estimation of R´ enyi diver- gence and friends. InProceedings of the 31st International Conference on International Conference on Machine Learning - Volume 32, ICML’14, page II–919–II–927, 2014

  41. [49]

    B. Laurent. Efficient estimation of integral functionals of a density.The Annals of Statistics, 24(2):659 – 681, 1996

  42. [50]

    Ledoux and M

    M. Ledoux and M. Talagrand.Probability in Banach Spaces. Springer-Verlag Berlin Heidelberg, 1991

  43. [51]

    Leonenko, L

    N. Leonenko, L. Pronzato, and V. Savani. A class of R´ enyi information estimators for multidimensional densities. The Annals of Statistics, 36(5):2153 – 2182, 2008

  44. [52]

    Leung-Yan-Cheong and M

    S. Leung-Yan-Cheong and M. Hellman. The Gaussian wire-tap channel.IEEE Transactions on Information Theory, 24(4):451–456, 1978

  45. [53]

    H. Liu, L. Wasserman, and J. Lafferty. Exponential concentration for mutual information estimation with ap- plication to forests. InAdvances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012

  46. [54]

    J. Liu, P. Cuff, and S. Verd´ u.E γ-resolvability.IEEE Transactions on Information Theory, 63(5):2629–2658, 2017

  47. [55]

    Mardia, J

    J. Mardia, J. Jiao, E. T´ anczos, R. D. Nowak, and T. Weissman. Concentration inequalities for the empirical distribution of discrete distributions: beyond the method of types.Information and Inference: A Journal of the IMA, 9(4):813–850, 11 2019

  48. [56]

    I. Mironov. R´ enyi differential privacy. In2017 IEEE 30th Computer Security Foundations Symposium (CSF), pages 263–275, 2017

  49. [57]

    K. R. Moon, K. Sricharan, K. Greenewald, and A. O. Hero. Ensemble estimation of information divergence. Entropy, 20(8), Aug. 2018

  50. [58]

    Nguyen, M

    X. Nguyen, M. J. Wainwright, and M. I. Jordan. Estimating divergence functionals and the likelihood ratio by convex risk minimization.IEEE Transactions on Information Theory, 56(11):5847–5861, Oct. 2010

  51. [59]

    Noshad, K

    M. Noshad, K. R. Moon, S. Y. Sekeh, and A. O. Hero. Direct estimation of information divergence using nearest neighbor ratios. InProceedings of the 2017 IEEE International Symposium on Information Theory, pages 903– 907, Jun. 2017

  52. [60]

    P´ al, B

    D. P´ al, B. P´ oczos, and C. Szepesv´ ari. Estimation of R´ enyi entropy and mutual information based on generalized nearest-neighbor graphs. InProceedings of the 24th International Conference on Neural Information Processing Systems - Volume 2, NIPS’10, page 1849–1857, Red H...

  53. [61]

    Paninski

    L. Paninski. Estimation of entropy and mutual information.Neural Computation, 15(6):1191–1253, June 2003

  54. [62]

    Perez-Cruz

    F. Perez-Cruz. Kullback-Leibler divergence estimation of continuous distributions. InProceedings of the 2008 IEEE International Symposium on Information Theory, pages 1666–1670, Toronto, ON, Canada, Jul. 2008

  55. [63]

    Poczos and J

    B. Poczos and J. Schneider. On the estimation ofα-divergences. InProceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, volume 15 ofProceedings of Machine Learning Research, pages 609–617, Fort Lauderdale, FL, USA, 11–13 Apr 2011. PML...

  56. [64]

    A. R´ enyi. On measures of entropy and information. InProceedings of the Fourth Berkeley Symposium on Math- ematical Statistics and Probability, volume 1, pages 547–561, Berkeley, 1961. University of California Press

  57. [65]

    H. Robbins. A remark on Stirling’s formula.The American Mathematical Monthly, 62(1):26–29, 1955

  58. [66]

    J. J. Ryu, S. Ganguly, Y.-H. Kim, Y.-K. Noh, and D. D. Lee. Nearest neighbor density functional estimation from inverse Laplace transform.IEEE Transactions on Information Theory, 68(6):3511–3551, 2022

  59. [67]

    Schmidt-Hieber

    J. Schmidt-Hieber. Nonparametric regression using deep neural networks with ReLU activation function.The Annals of Statistics, 48(4):1875 – 1897, 2020

  60. [68]

    Singh, N

    H. Singh, N. Misra, V. Hnizdo, A. Fedorowicz, and E. Demchuk. Nearest neighbor estimates of entropy.American Journal of Mathematical and Management Sciences, 23(3-4):301–321, 2003

  61. [69]

    Singh and B

    S. Singh and B. P´ oczos. Exponential concentration of a density functional estimator. InProceedings of Advances in Neural Information Processing Systems, volume 27, page 3032–3040, Montreal, Canada, Dec. 2014

  62. [70]

    Singh and B

    S. Singh and B. Poczos. Generalized exponential concentration inequality for R´ enyi divergence estimation. In Proceedings of the 31st International Conference on Machine Learning, volume 32 (1) ofProceedings of Machine Learning Research, pages 333–341, Bejing, China, 22–24 Ju...

  63. [71]

    Singh and B

    S. Singh and B. P´ oczos. Finite-sample analysis of fixed-k nearest neighbor density functional estimators. In Proceedings of Advances in Neural Information Processing Systems, volume 29, pages 1225–1233, Barcelona, Spain, Dec. 2016

  64. [72]

    Sreekumar and M

    S. Sreekumar and M. Berta. Limit distribution theory for quantum divergences.IEEE Transactions on Infor- mation Theory, 71(1):459–484, 2025

  65. [73]

    Sreekumar, A

    S. Sreekumar, A. Bunin, Z. Goldfeld, H. H. Permuter, and S. Shamai. The secrecy capacity of cost-constrained wiretap channels.IEEE Transactions on Information Theory, 67(3):1433–1445, 2021

  66. [74]

    Sreekumar and Z

    S. Sreekumar and Z. Goldfeld. Soft-covering via constant-composition superposition codes. In2021 IEEE Inter- national Symposium on Information Theory (ISIT), pages 2876–2881, 2021

  67. [75]

    Sreekumar and Z

    S. Sreekumar and Z. Goldfeld. Neural estimation of statistical divergences.Journal of Machine Learning Research, 23(126):1–75, 2022

  68. [76]

    Sreekumar, Z

    S. Sreekumar, Z. Goldfeld, and K. Kato. Limit distribution theory for f-divergences.IEEE Transactions on Information Theory, 70(2):1233–1267, 2024

  69. [77]

    Sricharan, R

    K. Sricharan, R. Raich, and A. O. Hero. Estimation of nonlinear functionals of densities with confidence.IEEE Transactions on Information Theory, 58(7):4135–4159, 2012

  70. [78]

    D. Tsur, Z. Aharoni, Z. Goldfeld, and H. Permuter. Neural estimation and optimization of directed information over continuous spaces.IEEE Transactions on Information Theory, 69(8):4777–4798, 2023

  71. [79]

    D. Tsur, Z. Aharoni, Z. Goldfeld, and H. Permuter. Data-driven optimization of directed information over discrete alphabets.IEEE Transactions on Information Theory, 70(3):1652–1670, 2024

  72. [80]

    A. B. Tsybakov and E. C. van der Meulen. Root-n consistent estimators of entropy for densities with unbounded support.Scandinavian Journal of Statistics, 23(1):75–83, 1996

  73. [81]

    Valiant and P

    G. Valiant and P. Valiant. Estimating the unseen: an n/log(n)-sample estimator for entropy and support size, shown optimal via new CLTs. InProceedings of the Forty-Third Annual ACM Symposium on Theory of Com- puting, STOC ’11, page 685–694, New York, NY, USA, 2011. Association...

  74. [82]

    Valiant and P

    G. Valiant and P. Valiant. Estimating the unseen: Improved estimators for entropy and other properties.J. ACM, 64(6), Oct. 2017

  75. [83]

    A. W. van der Vaart and J. A. Wellner.Weak Convergence and Empirical Processes. Springer, New York, 1996

  76. [84]

    van Erven and P

    T. van Erven and P. Harremo¨ es. R´ enyi divergence and Kullback-Leibler divergence.IEEE Transactions on Information Theory, 60(7):3797–3820, Jul. 2014

  77. [85]

    Q. Wang, S. R. Kulkarni, and S. Verdu. Divergence estimation of continuous distributions based on data- dependent partitions.IEEE Transactions on Information Theory, 51(9):3064–3074, Sep. 2005

  78. [86]

    Wisler, K

    A. Wisler, K. Moon, and V. Berisha. Direct ensemble estimation of density functionals. InProceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 2866–2870, Apr. 2018

  79. [87]

    Wu and P

    Y. Wu and P. Yang. Minimax rates of entropy estimation on large alphabets via best polynomial approximation. IEEE Transactions on Information Theory, 62(6):3702–3720, 2016

  80. [88]

    Yarotsky

    D. Yarotsky. Error bounds for approximations with deep ReLU networks.Neural Networks, 94:103–114, Oct. 2017

  81. [89]

    Zhao and L

    P. Zhao and L. Lai. Minimax optimal estimation of KL divergence for continuous distributions.IEEE Transac- tions on Information Theory, 66(12):7787–7811, 2020. (S. Sreekumar)L2S, CNRS, CentraleSup ´elec, University of Paris-Saclay, France Email address:sreejith.sreekumar@centr...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.