REVIEW 3 major objections 5 minor 14 references
Information geometry of Bayes computations
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This note argues that the nonparametric dually affine statistical bundle formalism—a Banach-space version of classical information geometry in which the Fisher score is a velocity—is the natural setting for Bayes computations, yielding…
desk verdict Concise note with correct-looking Bayes computations, but the core KL gradient formulas are deferred to an unpublished paper; fixable and worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the statistical bundle $S E(\mu) = \{(p, v) : p \in E(\mu), v \in B, E_p[v] = 0\}$, a vector bundle over the maximal exponential model $E(\mu) = \{ e^{u - \psi(u)} \mu : u \in S \subset B\}$, whose fibers are the tangent spaces of centered random variables. Two affine charts carry the argument: the exponential chart $s_p(q) = \log(q/p) - E_p[\log(q/p)]$ and the mixture chart $\eta_p(q) = q/p - 1$, together with the parallel transports $eU^p_q v = v - E_p[v]$ and $mU^p_q w = (q/p) w$. The Weyl cocycle identities make these charts a genuine affine atlas, so the Fisher score is the velocity in the moving frame. All of the paper's identities—the KL gradients, marginalization as conditional expectation, and conditioning as centered score—are direct computations in these charts.
What would settle it
A concrete test is to represent a joint density by two different versions that agree almost everywhere but differ at a point $x_0$, and check whether Eq. (10) produces the same conditional score; if the two versions give different results, the formula is not well-defined in the standard $L^1$ equivalence-class setting, showing the central claim depends on the pointwise-everywhere regularity assumption.
Extended reading notes
Core claim
The paper's central claim is that every basic Bayes operation has a clean expression in the statistical bundle over the maximal exponential model. For a joint density $p_{12}$, the velocity of marginalization is the conditional expectation of the joint score given the first coordinate, Eq. (9). The velocity of conditioning at a point $x$ is the joint score evaluated at $x$ minus its conditional average under the conditional density, Eq. (10) — in other words, the centered score. The KL divergence is the natural divergence of the affine charts: the first natural gradient is the negative exponential chart $s_p(q)$, and the second is the negative mixture chart $\eta_q(p)$, Eqs. (6)-(7). An exponential factorization identity and the KL chain rule follow from the same charts, and a worked exponential-family example recovers the classical score formulas as special cases.
Load-bearing premise
The load-bearing premise is that densities in the model are defined pointwise everywhere on the outcome space, so that conditioning at a fixed observation $x$ is meaningful; if densities are only defined up to null sets, the conditioning map and the derivative formula in Eq. (10) require a version or fail.
Editorial extensions
If this is right
- The marginalization derivative (Eq. 9) reveals the conditional expectation operator as the tangent map of a geometric projection, so Bayesian marginalization becomes a first-order calculus operation.
- The gradient formulas (Eqs. 6-7) yield a coordinate-free natural gradient for optimizing over densities, which can be used directly in variational Bayes and other divergence-based inference.
- The exponential decomposition of the joint density gives a chart-based proof of the KL chain rule, the additive structure central to evidence lower bound computations.
- For exponential families, the derived formulas reduce to the classical facts that the score is the centered sufficient statistic and the posterior score is its conditional expectation, so the bundle calculus subsumes the parametric results.
- The paper's final remark notes that applying the derived differential equations in applied statistics such as variational Bayes depends on choosing an exponential family with suitable special characters.
Reading between the lines
- If the pointwise-regularity assumption can be relaxed to a.e.-defined densities using a version of conditional expectation, the same derivative formulas would extend to general Polish spaces; this is an extension the paper does not make.
- Read geometrically, Eq. (10) says a Bayesian update moves the posterior in the direction of the residual of the joint score after removing its conditional mean—a new interpretation of the information contributed by an observation.
- The same bundle calculus applies to any map between density models defined by integration or conditioning, so it could also serve marginalization in graphical models or transport-based inference, not just Bayes.
- A direct test of the framework would be to implement the natural gradient in a nonparametric variational Bayes scheme where the approximating family is an infinite-dimensional exponential model and check whether the centered-score derivative matches a finite-difference estimate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This note presents the author's statistical-bundle formalism as a setting for nonparametric dually affine information geometry. After recalling the exponential and mixture charts, it states formulas for the total natural gradient of the KL divergence (Eqs. (6)-(7)), derives the derivative of the marginalization map as a conditional expectation (Eq. (9)) and the derivative of the conditioning map as a centered conditional score (Eq. (10)), and applies these to exponential families in Section 3.4. The paper is written as a concise research note; the abstract and final remark claim that these computations demonstrate the usefulness of the statistical-bundle formalism for Bayes computations and variational Bayes.
Significance. The derivations that are actually carried out in the note are correct and clean: Eq. (9) and Eq. (10) follow from the chart calculus, and the exponential-family example is internally consistent. If the gradient formulas in Eqs. (6)-(7) were proved in the note or available in a published source, the paper would be a useful, mostly self-contained reference for the nonparametric dually affine calculus of Bayes computations. In its current form, however, the central KL-gradient claim rests on an unpublished, in-progress reference [11], so the standalone significance is diminished. The note is also explicitly restricted to pointwise-defined densities (e.g., continuous densities on a bounded domain), and that limitation should be reflected in the abstract and final remarks.
major comments (3)
- [Section 2, Eqs. (6)-(7)] The total natural gradient formulas are the main load-bearing claim of the note, but they are asserted as a 'direct computation' and their derivation is deferred to [11], which the reference list describes as 'In progress'. These formulas are not decorative: Section 3.4 uses both equations to compute parametric KL derivatives, and the final remark presents the KL gradient as a main advantage of the formalism. As the manuscript stands, the central claim is not independently verifiable from the text. Please include a self-contained derivation from the chart definitions in Eqs. (3)-(4), or cite a published proof; this is the main point that must be fixed.
- [Section 3.2, Eq. (10)] The derivation of the conditioning derivative relies on pointwise evaluation of densities at a fixed observation x, as the text acknowledges ('well-defined in our setup because the densities are defined everywhere'). This restricts the nonparametric claim made in the abstract and final remark, since the usual L^1 setting identifies densities up to null sets. Please state this limitation at the first mention of the nonparametric setup and clarify whether Eq. (10) is independent of the chosen density version; if it is not, the scope of the claim should be reduced accordingly.
- [Section 3.4] The parametric gradient computations are displayed only after the phrase 'from eq. (14) with eq. (7) and eq. (6), respectively'. The reader must reconstruct the pairing step and the cancellation of the mean term to see why the displayed integrals have no E[T] contribution. Once Eqs. (6)-(7) are proved, this is a short chain-rule computation, but the note should show at least one line of that pairing so Section 3.4 is self-contained.
minor comments (5)
- [Reference [11]] The title of reference [11] contains a typo ('Kullback-Leible r' should be 'Kullback-Leibler'). More importantly, because this reference is listed as 'In progress', it cannot carry the derivation of Eqs. (6)-(7) in a published manuscript.
- [Section 1] The 'Weyl's axioms' are invoked without being stated. Since the note is meant as a concise review, either state the axioms or give a precise pointer to the relevant part of [5].
- [Section 1 and Section 2] The notation '★' for the Fisher score is used repeatedly but defined only implicitly in the displayed curve computation; an explicit definition would help readers who are not already familiar with the statistical-bundle formalism.
- [Section 2] The term 'total natural gradient' is introduced through the displayed identity with arbitrary curves, but existence and uniqueness of the pair (grad1, grad2) are not discussed. A sentence noting that these are the ordinary derivatives in the exponential and mixture charts would make the definition transparent and would also clarify the meaning of the 'direct computation'.
- [Throughout] There are several typographical artifacts in the text (for example, 'expession' in Section 3.1 and the rendering of some formulas as 'dsp' in the extracted text); these should be cleaned up in the final version.
Circularity Check
No circularity: the Bayes derivative computations are derived in-paper; deferred proof of Eqs. (6)-(7) is a self-citation/rigor concern, not a circular reduction.
full rationale
The paper's advertised Bayes computations do not reduce by construction to their inputs. Section 3.1 derives Eq. (9), identifying the derivative of marginalization with a conditional expectation, directly from the mixture-chart expression of the marginalization map. Section 3.2 derives Eq. (10), identifying the bundle derivative of conditioning with a centered score, by an explicit chain-rule calculation in the mixture chart. Neither derivation assumes the conclusion or fits a parameter to a target quantity. The KL gradient identities (6)-(7) are introduced as a 'direct computation' with the proof deferred to the author's in-progress manuscript [11]; this is a self-citation and rigor concern for a conference note, but not a circularity, because the displayed definitions (3)-(4) of the exponential and mixture charts and the KL divergence determine these derivatives, and the formulas are not fitted parameters or restatements of the conclusion. The references to [5] review the framework and are not used to import an unverified uniqueness theorem or to forbid alternatives. The restriction to pointwise-defined densities, stated in Section 1 and again in Section 3.2, is a genuine scope limitation but not a circular step. Overall, the derivation chain is self-contained for the main Bayes computations, with one deferred external-looking proof that lowers standalone verifiability but does not create circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The maximal exponential model E(mu) contains all densities exp(u - psi(u)) * mu and S is the largest open convex domain such that an open exponential arc connects any two densities.
- domain assumption Densities are defined pointwise everywhere, not just up to null sets, so conditioning p_{2|1}(y|x) is well defined for each fixed x.
- ad hoc to paper The gradient formulas grad1 D(p||q) = -s_p(q) and grad2 D(p||q) = -eta_q(p) in Eqs. (6)-(7) hold as stated.
Cite this review
Pith. "Pith review of Information geometry of Bayes computations." pith.science (2026). https://pith.science/paper/BC4FP56B
@misc{pith2026250202160,
author = {Pith},
title = {Pith review of: Information geometry of Bayes computations},
year = {2026},
howpublished = {\url{https://pith.science/paper/BC4FP56B}},
note = {Machine review of arXiv:2502.02160}
}
read the original abstract
Amari's Information Geometry is a dually affine formalism for parametric probability models. The literature proposes various nonparametric functional versions. Our approach uses classical Weyl's axioms so that the affine velocity of a one-parameter statistical model equals the classical Fisher's score. In the present note, we first offer a concise review of the notion of a statistical bundle as a set of couples of probability densities and Fisher's scores. Then, we show how the nonparametric dually affine setup deals with the basic Bayes and Kullback-Leibler divergence computations.
Reference graph
Works this paper leans on
-
[11]
Pistone, G.: Constrained minima of the Kullback-Leible r divergence (2025). In progress
work page 2025
-
[5]
Dually affine Information Geometry modeled on a Banach space
Chirco, G., Pistone, G.: Dually affine Information Geometr y modeled on a Banach space (2022). DOI 10.48550/ARXIV.2204.00917. URL https://arxiv.org/abs/2204.00917. arXiv:2204.00917
work page Pith review arXiv doi:10.48550/arxiv.2204.00917 2022
-
[1]
American Mathematical Society (2000)
Amari, S., Nagaoka, H.: Methods of information geometry. American Mathematical Society (2000). Translated from the 1993 Japanese original by Daishi Harada
work page 2000
-
[2]
Amari, S.i.: Information geometry and its applications, Applied Mathematical Sciences , vol
-
[3]
Arnold, V.I.: Mathematical methods of classical mechani cs, Graduate Texts in Mathematics, vol. 60. Springer-Verlag, New York (1989). Translated from the 1974 Russian original by K. Vogtmann and A. Weinstein, Corrected reprint of the second (1989) edition
work page 1989
-
[4]
Ay, N., Jost, J., L ˆe, H.V., Schwachh ¨ofer, L.: Information geometry, Ergebnisse der Math- ematik und ihrer Grenzgebiete. 3. Folge. A Series of Modern S urveys in Mathematics [Results in Mathematics and Related Areas. 3rd Series. A Ser ies of Modern Surveys in Mathematics], vol. 64. Springer, Cham (2017). DOI 10.1007/978-3-319-56 478-4. URL https://doi....
-
[6]
Efron, B., Hastie, T.: Computer age statistical inferenc e, Institute of Mathematical Statis- tics (IMS) Monographs , vol. 5. Cambridge University Press, New York (2016). URL https://doi.org/10.1017/CBO9781316576533. Algorithms, evidence, and data science
-
[7]
Acta Mat hematica Sinica, En- glish Series 22(4), 1175–1182 (2006)
Egozcue, J.J., D ´ıaz–Barrero, J.L., Pawlowsky–Glahn, V.: Hilbert space of p robabil- ity density functions based on Aitchison geometry. Acta Mat hematica Sinica, En- glish Series 22(4), 1175–1182 (2006). DOI 10.1007/s10114-005-0678-2. UR L http://dx.doi.org/10.1007/s10114-005-0678-2
Show all 14 references
-
[8]
URL https://arxiv.org/abs/2411.03265
Khesin, B., Misio lek, G., Modin, K.: Information geometry of diffeomorphism groups (2024). URL https://arxiv.org/abs/2411.03265
2024 arXiv
-
[9]
Pistone, G.: Nonparametric information geometry. In: F. Nielsen, F. Barbaresco (eds.) Geo- metric science of information, Lecture Notes in Comput. Sci. , vol. 8085, pp. 5–36. Springer, Heidelberg (2013). First International Conference, GSI 20 13 Paris, France, August 28-30, 20...
2013
-
[10]
Pistone, G.: Statistical bundle of the transport model. In: F. Nielsen, F. Barbaresco (eds.) Geometric Science of Information, pp. 752–759. Springer In ternational Publishing, Cham (2021)
2021
-
[12]
Pistone, G., Sempi, C.: An infinite-dimensional geometr ic structure on the space of all the probability measures equivalent to a given one. Ann. Statist. 23(5), 1543–1561 (1995)
1995
-
[13]
URL https://arxiv.org/abs/1912.08003
Talska, R., Menafoglio, A., Hron, K., Egozcue, J.J., Pal area-Albaladejo, J.: Changing ref- erence measure in bayes spaces with applications to functio nal data analysis (2019). URL https://arxiv.org/abs/1912.08003
2019 arXiv
-
[194]
URL https://doi.org/10.1007/978-4-431-55978-8
Springer, [Tokyo] (2016). URL https://doi.org/10.1007/978-4-431-55978-8
2016 doi
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.