REVIEW 3 major objections 4 minor 27 references
Maximum entropy on projective space yields one optimizer for many entropies, and an acceptance ellipsoid fixes the deformation parameter.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 20:39 UTC pith:QF5IMZD5
load-bearing objection A sound, honest methods paper: the calibration identity is a clean design tool, but the universality framing is definitional and the acceptance-region scope is narrower than the title implies. the 3 major comments →
Projective Maximum Entropy: Universality and Acceptance-Region Calibration
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the normalized power functional J_γ([r]) = ∫ r̄^{1+γ} dμ₀, defined on the projective space of nonnegative measures, is the common variational object behind a broad class of entropy and scoring-rule constructions. For fixed moment constraints, every admissible strictly monotone transform of J_γ yields exactly the same maximizer, a universality result that unifies the maximum-entropy implications of Tsallis and Rényi entropies, Hölder composite scores, and related homogeneous divergences. Under mean and covariance constraints, this common optimizer is the q-exponential density: for γ>0 it is the compactly supported q-Gaussian with support {Q ≤ d+2+2/γ}, and fo
What carries the argument
The central object is the projective power functional J_γ([r]) = ∫ r̄^{1+γ} dμ₀, defined on the quotient of nonnegative measures by positive rescaling, with r̄ the normalized representative of the ray [r]. This functional depends only on the shape of the unnormalized measure, not its total mass, and it is the common diagonal ordering behind the listed entropy and scoring-rule constructions. The universality theorem (Theorem 1) states that for any admissible monotone transform of J_γ, the optimizer under linear moment constraints is the same; the proof uses strict convexity (γ>0) or strict concavity (−1<γ<0) of t↦t^{1+γ}. The q-exponential form then follows from the variational Lagrangian, an
Load-bearing premise
The calibration theorem presupposes that the prescribed acceptance region is exactly the Mahalanobis ellipsoid A_{μ,V}(R) formed from the same mean μ and covariance V that appear in the moment constraints; if the externally given admissible region has a different center, shape, or topology, the formula γ_R = 2/(R²−d−2) does not apply.
What would settle it
Take d=1, μ=0, V=1, and choose a radius R²>3, say R²=4, so γ_R=2. Construct the density q_R(x) = C(1−x²/4)_+²? (with the appropriate normalization from Eq. (21)) and verify numerically that it has mean 0 and covariance 1. Then maximize the projective power functional 1−∫p³ dx over all one-dimensional densities with mean 0 and covariance 1; if any density with these moments yields a strictly larger value than q_R, the universal maximum-entropy claim is false. Alternatively, choose R²=3 (the lower boundary) and check whether a density with mean 0 and covariance 1 can have support contained in th
If this is right
- If the calibration result holds, a prescribed Mahalanobis acceptance region directly determines a unique bounded-support reference density with matching mean and covariance, so the deformation parameter is no longer an abstract tuning constant.
- The same maximizer is shared by Tsallis entropy, Rényi entropy, Hölder composite scores, pseudo-spherical scores, Bregman–Hölder potentials, and dual homogeneous potentials under the same moment constraints, implying that maximum-entropy selection is insensitive to which of these functionals is used.
- The family interpolates continuously between a uniform density on a small ellipsoid (as R² ↓ d+2) and the Gaussian density (as R² → ∞), providing a tractable spectrum of reference distributions.
- Affine equivariance of the calibrated family means that applying an affine transformation to the data simply transforms the parameters μ and V in the same way, which is desirable for multivariate reference construction.
- For negative deformation, the Student-type density gives a heavy-tailed reference whose tail thickness is controlled by the same deformation parameter through ν = −2/γ−d, so the sign of γ distinguishes bounded-support from heavy-tailed design.
Where Pith is reading between the lines
- If the calibration theorem extends to other families of admissible regions, one could in principle calibrate a deformation parameter from a box or a union of ellipsoids; the paper explicitly does not provide such a construction, so this remains an open direction rather than a paper claim.
- The universality theorem suggests that, for the purpose of maximum-entropy selection, the diagonal entropy ordering is the only relevant feature of a scoring rule; off-diagonal behavior, influence functions, and scale identification are independent and must be chosen separately.
- The identity γ_R = 2/(R²−d−2) makes R² an identifiable parameter of the q-Gaussian family, so in a data-rich setting one could estimate the effective acceptance radius from the shape of the fitted density rather than specifying it externally.
- A testable extension would be to replace the mean and covariance constraints with escort moments or robust estimating equations; the paper notes that universality would then need to be re-examined because the feasible set may itself depend on the deformation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a projective maximum-entropy framework on the quotient space of nonnegative measures. It defines a projective power functional J_γ and a class of γ-admissible entropies as its strictly monotone transforms (Def. 2). Theorem 1 claims that all such entropies have the same moment-constrained optimizer; Theorem 2 gives a q-exponential form under linear constraints, conditional on KKT regularity. Theorem 3 solves the mean–covariance problem explicitly: for γ>0 the unique maximizer is a compactly supported q-Gaussian, and for −2/(d+2)<γ<0 it is a Student-type density, with global optimality proved by convexity/concavity in Appendix A. Theorem 4 inverts the support-radius identity R²=d+2+2/γ to set γ_R=2/(R²−d−2), obtaining a bounded-support reference density whose support is exactly a prescribed Mahalanobis ellipsoid A_{μ,V}(R), with no additional support constraint. The paper also discusses affine equivariance, limiting cases, and connections to existing scoring rules and unnormalized models.
Significance. If the claims are taken as stated, the paper provides a clean unification of known power-entropy extremal results and a practical calibration formula for bounded-support reference distributions. The mean–covariance theorem is proven carefully: the beta-integral calculations in Appendix A are correct, and the global optimality argument via convexity/concavity is a genuine strength. The calibration formula is an exact inversion of Eq. (17) and is derived cleanly. However, the universality theorem is essentially definitional, since γ-admissibility is defined as monotone transformation of J_γ, and the acceptance-region calibration is restricted to Mahalanobis ellipsoids with the same center and shape as the moment constraints. These limitations materially affect how the paper's central claims should be read, but they do not undermine the correctness of the core derivations.
major comments (3)
- [§6, Theorem 4 and §1] The calibration theorem is an exact inversion of the support-radius identity (17) for the density family (14). It applies only when the prescribed acceptance region is exactly A_{μ,V}(R) = {Q_{μ,V}(x) ≤ R²}, i.e., a Mahalanobis ellipsoid centered at μ with shape V, where (μ,V) are simultaneously the moment constraints. The paper's title and Section 1 motivation (safety envelopes, feasible design domains, acceptance regions) suggest a more general admissible-region calibration. No construction is given for non-ellipsoidal regions, displaced ellipsoids, or ellipsoids with different center/orientation than the moment-constraint pair, and Remark 4's threshold R²>d+2 further restricts the geometry. This is not an internal inconsistency, but it is a load-bearing scope restriction. The abstract and introduction should state this explicitly, and ideally the paper should discuss whether approxima
- [§3, Definition 2 and Theorem 1] Theorem 1 is a direct consequence of the definition of γ-admissibility. Since H_{γ,F} is defined as F_γ(J_γ) with F_γ strictly monotone, and the reference functional S_γ is also a monotone function of J_γ, the equality of arg max follows immediately on any feasible set. The proof is one line and contains no new mathematical content. The substantive contribution is the identification that Tsallis, Rényi, H"older composite, pseudo-spherical, and related entropies all fall into this equivalently-ordered class (Table 1, Corollary 1). Please reframe Theorem 1 as a lemma or remark arising from the definition, and place the emphasis on the catalog of examples and the separation between diagonal ordering and off-diagonal behavior.
- [§4, Theorem 2] Theorem 2 is stated conditionally on 'the usual variational and KKT regularity conditions,' but no constraint qualification is verified for the general linear-constraint problem. As stated, it is a formal necessary-condition result, not a fully validated characterization. This is not fatal because Theorems 3 and 4 are proven independently via convexity/concavity in Appendix A and do not rely on Theorem 2. Still, the paper should either add explicit constraint-qualification hypotheses (e.g., Slater-type conditions for the positive-part constraint) or explicitly label Theorem 2 as a formal derivation that holds when the appropriate regularity conditions are satisfied.
minor comments (4)
- [Figure 1 legend] The legend shows entries such as 'R² = 3.2, = 10' — the γ symbol appears to be missing due to rendering. Please check the figure text.
- [§2.3, Definition 2] The symbol F_γ is used both for the transform and for the resulting functional. Using φ_γ or another letter for the transform would reduce confusion.
- [§3.1, Eq. (6)] The affinity ρ_γ is introduced only for γ>0. The later negative-branch discussion in Corollary 2 and Remark 2 would benefit from a brief sentence clarifying that the negative case is handled by the reverse-H"older form, so that the affinity notation is not misleading.
- [References] Reference [17] has a capitalization typo: 'boltzmann–gibbs statistics' should be 'Boltzmann–Gibbs statistics.'
Circularity Check
Universality theorem is definitional by construction; the mean-covariance optimizer proof and the calibration inversion are independent, so the circularity is partial.
specific steps
-
self definitional
[Definition 2 (Eq. 5) and Theorem 1 (Eq. 12)]
"The functional Hγ,F ([r]) =F γ (Jγ ([r])) (5) is called aγ-admissible projective entropy."
The class of γ-admissible entropies is defined as strictly monotone functions of the same normalized power functional Jγ. Theorem 1 then concludes that every member has exactly the same optimizer as Sγ = (1−Jγ)/γ. The proof simply restates that monotone transformations preserve the argmax of Jγ. Thus the advertised 'universality' is built into the definition of the class rather than derived from an independent variational principle. It is a true but tautological statement: the equivalence holds by construction because Hγ,F was defined to be Fγ(Jγ).
full rationale
The paper has three linked contributions. The third and headline-result, acceptance-region calibration, is not circular in a damaging sense: it is a transparent algebraic inversion of the support-radius identity R² = d+2+2/γ (Eq. 17), and it is explicitly restricted to Mahalanobis ellipsoids sharing the same μ and V as the moment constraints (Remark 5 notes the restriction is not technical). That scope limitation is a correctness/scope concern, not a hidden circularity. The mean-covariance optimizer proof in Appendix A is also self-contained: it uses beta integrals and convexity/concavity of t^{1+γ} to establish global optimality, and it does not depend on the paper's own prior work. The genuinely constructional element is Theorem 1: the universality claim reduces to the definition of γ-admissible entropy as Fγ(Jγ). This does not infect the independent proof of the q-Gaussian/Student optimizer or the algebraic calibration step, so the circularity is partial rather than total, but one of the three principal contributions is true by definition rather than by substantive derivation.
Axiom & Free-Parameter Ledger
free parameters (1)
- R² (squared Mahalanobis radius of acceptance region)
axioms (4)
- domain assumption Standard measure-theoretic setup and integrability of power integrals (Section 2.1)
- domain assumption Existence and KKT regularity of the maximizer in the constrained variational problem (Theorem 2)
- standard math Convexity of the feasible set C_{τ,γ} (Theorem 1)
- standard math Beta-integral evaluations and Hölder inequality (Appendix A)
read the original abstract
Maximum-entropy reference distributions are usually constructed on the normalized probability simplex. This formulation is less natural for unnormalized statistical models, in which positive multiples represent the same shape, and it does not directly explain how a prescribed admissible region should determine the deformation parameter of a bounded-support reference distribution. We formulate maximum entropy on the projective space of nonnegative measures and establish three results of statistical relevance. First, a universality theorem shows that every admissible monotone transform of the same normalized power functional has exactly the same optimizer under linear moment constraints. The result unifies the maximum-entropy implications of Tsallis and R\'enyi entropies, H\"older composite scores, pseudo-spherical scores, Bregman--H\"older constructions, and related homogeneous divergences without asserting a new distribution family. Second, the common optimizer is characterized as a $q$-exponential density; under mean and covariance constraints it is a compactly supported $q$-Gaussian for positive deformation and a Student-type density for negative deformation. Third, a prescribed Mahalanobis acceptance region with squared radius $R^2>d+2$ uniquely determines the deformation parameter $\gamma_R=2/(R^2-d-2)$. The resulting affine-equivariant reference density is the unique projective maximum-entropy solution, and its support coincides with the specified ellipsoid without an additional support constraint. This provides a principled method for constructing bounded-support statistical reference distributions from robust location and scatter estimates or from externally specified admissible regions.
Reference graph
Works this paper leans on
-
[1]
Physical Review 106(4), 620–630 (1957) https://doi.org/10.1103/PhysRev.106.620
Jaynes, E.T.: Information theory and statistical mechanics. Physical Review 106(4), 620–630 (1957) https://doi.org/10.1103/PhysRev.106.620
-
[2]
IEEE Transactions on Information Theory51(2), 473–478 (2005) https://doi.org/10.1109/TIT.2004
Lutwak, E., Yang, D., Zhang, G.: Cram´ er–Rao and moment-entropy inequali- ties for R´ enyi entropy and generalized Fisher information. IEEE Transactions on Information Theory51(2), 473–478 (2005) https://doi.org/10.1109/TIT.2004. 840871
doi:10.1109/tit.2004 2005
-
[3]
Johnson, O., Vignat, C.: Some results concerning maximum R´ enyi entropy distri- butions. Annales de l’Institut Henri Poincar´ e, Probabilit´ es et Statistiques43(3), 339–351 (2007) https://doi.org/10.1016/j.anihpb.2006.05.001
-
[4]
Entropy22(11), 1244 (2020) https://doi.org/10.3390/ e22111244
Reeves, G.: A two-moment inequality with applications to R´ enyi entropy and mutual information. Entropy22(11), 1244 (2020) https://doi.org/10.3390/ e22111244
2020
-
[5]
Entropy13(6), 1170–1185 (2011) https://doi.org/10.3390/e13061170
Amari, S.-i., Ohara, A.: Geometry ofq-exponential family of probability distri- butions. Entropy13(6), 1170–1185 (2011) https://doi.org/10.3390/e13061170
-
[6]
Information Geometry1, 39–78 (2018) https://doi.org/10.1007/ s41884-018-0012-6
Wong, T.-K.L.: Logarithmic divergences from optimal transport and R´ enyi geometry. Information Geometry1, 39–78 (2018) https://doi.org/10.1007/ s41884-018-0012-6
2018
-
[7]
Wong, T.-K.L., Zhang, J.: Tsallis and R´ enyi deformations linked via a newλ- duality. IEEE Transactions on Information Theory68(8), 5353–5373 (2022) https: //doi.org/10.1109/TIT.2022.3159385
arXiv 2022
-
[8]
MATSUZOE, H.: INV ARIANT DUALLY FLAT STRUCTURES ON ¡italic¿q¡/italic¿-EXPONENTIAL F AMILIES, pp. 195–210. https://doi.org/10. 1142/9789811296710 0013 . https://www.worldscientific.com/doi/abs/10.1142/ 9789811296710 0013
-
[9]
Journal of the American Statistical Association102(477), 359–378 (2007)
Gneiting, T., Raftery, A.E.: Strictly proper scoring rules, prediction, and esti- mation. Journal of the American Statistical Association102(477), 359–378 (2007)
2007
-
[10]
Bernoulli24(1), 53–79 (2018) https://doi.org/10.3150/16-BEJ857
Ovcharov, E.Y.: Proper scoring rules and Bregman divergence. Bernoulli24(1), 53–79 (2018) https://doi.org/10.3150/16-BEJ857
-
[11]
Biometrika85(3), 549–559 (1998)
Basu, A., Harris, I.R., Hjort, N.L., Jones, M.C.: Robust and efficient estimation 19 by minimising a density power divergence. Biometrika85(3), 549–559 (1998)
1998
-
[12]
Journal of Multivariate Analysis99(9), 2053–2081 (2008)
Fujisawa, H., Eguchi, S.: Robust parameter estimation with a small bias against heavy contamination. Journal of Multivariate Analysis99(9), 2053–2081 (2008)
2053
-
[13]
Bernoulli20(4), 2278–2304 (2014)
Kanamori, T., Fujisawa, H.: Affine invariant divergences associated with proper composite scoring rules and their applications. Bernoulli20(4), 2278–2304 (2014)
2014
-
[14]
Entropy16, 2611–2628 (2014)
Kanamori, T.: Scale-invariant divergences for density functions. Entropy16, 2611–2628 (2014)
2014
-
[15]
Information Geometry6, 81–106 (2023)
Hino, H., Eguchi, S.: Active learning by query by committee with robust divergences. Information Geometry6, 81–106 (2023)
2023
-
[16]
Matsuzoe, H., Takatsu, A.: Gauge freedom of entropies onq-Gaussian measures. In: Nielsen, F. (ed.) Progress in Information Geometry: Theory and Applications. Signals and Communication Technology, pp. 127–152. Springer, Cham (2021). https://doi.org/10.1007/978-3-030-65459-7 6
-
[17]
Journal of Statistical Physics52, 479–487 (1988)
Tsallis, C.: Possible generalization of boltzmann–gibbs statistics. Journal of Statistical Physics52, 479–487 (1988)
1988
-
[18]
In: Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, vol
R´ enyi, A.: On measures of entropy and information. In: Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, vol. 1, pp. 547– 561 (1961)
1961
-
[19]
measuring information and uncertainty
Good, I.J.: Comment on “measuring information and uncertainty”. In: Godambe, V.P., Sprott, D.A. (eds.) Foundations of Statistical Inference, p. 337. Holt, Rinehart and Winston, ??? (1971)
1971
-
[20]
The Annals of Statistics40(1), 561–592 (2012) https://doi.org/10.1214/12-AOS971
Parry, M., Dawid, A.P., Lauritzen, S.: Proper local scoring rules. The Annals of Statistics40(1), 561–592 (2012) https://doi.org/10.1214/12-AOS971
-
[21]
Biometrika102(3), 559–572 (2015)
Kanamori, T., Fujisawa, H.: Robust estimation under heavy contamination using unnormalized models. Biometrika102(3), 559–572 (2015)
2015
-
[22]
Journal of Machine Learning Research18(56), 1–26 (2017)
Takenouchi, T., Kanamori, T.: Statistical inference with unnormalized discrete models and localized homogeneous divergences. Journal of Machine Learning Research18(56), 1–26 (2017)
2017
-
[23]
Bayesian Analysis10(2), 479–499 (2015) https://doi.org/10.1214/15-BA942
Dawid, A.P., Musio, M.: Bayesian model selection based on proper scoring rules. Bayesian Analysis10(2), 479–499 (2015) https://doi.org/10.1214/15-BA942
-
[24]
Journal of Machine Learning Research6, 695–709 (2005)
Hyv¨ arinen, A.: Estimation of non-normalized statistical models by score match- ing. Journal of Machine Learning Research6, 695–709 (2005)
2005
-
[25]
In: Proceedings of the 20 Twenty Third International Conference on Artificial Intelligence and Statistics
Uehara, M., Kanamori, T., Takenouchi, T., Matsuda, T.: A unified statistically efficient estimation framework for unnormalized models. In: Proceedings of the 20 Twenty Third International Conference on Artificial Intelligence and Statistics. Proceedings of Machine Learning Research, vol. 108, pp. 809–819. PMLR, ??? (2020)
2020
-
[26]
In: Advances in Neural Information Processing Systems, vol
Yu, L., Song, J., Song, Y., Ermon, S.: Pseudo-spherical contrastive divergence. In: Advances in Neural Information Processing Systems, vol. 34 (2021)
2021
-
[27]
In: Proceedings of the 42nd Inter- national Conference on Machine Learning
Ryu, J.J., Shah, A., Wornell, G.W.: A unified view on learning unnormalized distributions via noise-contrastive estimation. In: Proceedings of the 42nd Inter- national Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 267, pp. 52444–52474. PMLR, ??? (2025) 21
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.