Pith. sign in

REVIEW 3 major objections 5 minor 14 references

Information geometry of Bayes computations

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This note argues that the nonparametric dually affine statistical bundle formalism—a Banach-space version of classical information geometry in which the Fisher score is a velocity—is the natural setting for Bayes computations, yielding…

desk verdict Concise note with correct-looking Bayes computations, but the core KL gradient formulas are deferred to an unpublished paper; fixable and worth refereeing. read the letter →

arxiv 2502.02160 v1 pith:BC4FP56B submitted 2025-02-04 math.ST stat.TH

classification math.STstat.TH MSC 62B0562F15
keywords informationgeometrystatisticalbundleduallyaffinestructureBayescomputationKullback-LeiblerdivergenceconditionalexpectationFisherscorenonparametricinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This note argues that the nonparametric dually affine statistical bundle formalism—a Banach-space version of classical information geometry in which the Fisher score is the velocity of a curve—is the natural setting for Bayes computations. It establishes that the derivative of the marginalization map is a conditional expectation, $d M_\#(p)[s_p] = E_p[s_p \mid \pi]$, and that the bundle derivative of conditioning at a fixed observation is a centered score, $d K_x(p)[s_p] = s_p(x,\cdot) - E_{p_{2|1}(\cdot|x)}[s_p(x,\cdot)]$. It also derives the total natural gradient of the Kullback-Leibler divergence as $\operatorname{grad}_1 D(p\|q) = -s_p(q)$ and $\operatorname{grad}_2 D(p\|q) = -\eta_q(p)$, i.e., the negative exponential and mixture charts. If these claims hold, Bayesian calculations can be understood geometrically without finite-dimensional parametric assumptions, offering coordinate-free tools for variational Bayes and related applied statistics.

What carries the argument

The central object is the statistical bundle $S E(\mu) = \{(p, v) : p \in E(\mu), v \in B, E_p[v] = 0\}$, a vector bundle over the maximal exponential model $E(\mu) = \{ e^{u - \psi(u)} \mu : u \in S \subset B\}$, whose fibers are the tangent spaces of centered random variables. Two affine charts carry the argument: the exponential chart $s_p(q) = \log(q/p) - E_p[\log(q/p)]$ and the mixture chart $\eta_p(q) = q/p - 1$, together with the parallel transports $eU^p_q v = v - E_p[v]$ and $mU^p_q w = (q/p) w$. The Weyl cocycle identities make these charts a genuine affine atlas, so the Fisher score is the velocity in the moving frame. All of the paper's identities—the KL gradients, marginalization as conditional expectation, and conditioning as centered score—are direct computations in these charts.

What would settle it

A concrete test is to represent a joint density by two different versions that agree almost everywhere but differ at a point $x_0$, and check whether Eq. (10) produces the same conditional score; if the two versions give different results, the formula is not well-defined in the standard $L^1$ equivalence-class setting, showing the central claim depends on the pointwise-everywhere regularity assumption.

Watch

Extended reading notes

Core claim

The paper's central claim is that every basic Bayes operation has a clean expression in the statistical bundle over the maximal exponential model. For a joint density $p_{12}$, the velocity of marginalization is the conditional expectation of the joint score given the first coordinate, Eq. (9). The velocity of conditioning at a point $x$ is the joint score evaluated at $x$ minus its conditional average under the conditional density, Eq. (10) — in other words, the centered score. The KL divergence is the natural divergence of the affine charts: the first natural gradient is the negative exponential chart $s_p(q)$, and the second is the negative mixture chart $\eta_q(p)$, Eqs. (6)-(7). An exponential factorization identity and the KL chain rule follow from the same charts, and a worked exponential-family example recovers the classical score formulas as special cases.

Load-bearing premise

The load-bearing premise is that densities in the model are defined pointwise everywhere on the outcome space, so that conditioning at a fixed observation $x$ is meaningful; if densities are only defined up to null sets, the conditioning map and the derivative formula in Eq. (10) require a version or fail.

Editorial extensions

If this is right

  • The marginalization derivative (Eq. 9) reveals the conditional expectation operator as the tangent map of a geometric projection, so Bayesian marginalization becomes a first-order calculus operation.
  • The gradient formulas (Eqs. 6-7) yield a coordinate-free natural gradient for optimizing over densities, which can be used directly in variational Bayes and other divergence-based inference.
  • The exponential decomposition of the joint density gives a chart-based proof of the KL chain rule, the additive structure central to evidence lower bound computations.
  • For exponential families, the derived formulas reduce to the classical facts that the score is the centered sufficient statistic and the posterior score is its conditional expectation, so the bundle calculus subsumes the parametric results.
  • The paper's final remark notes that applying the derived differential equations in applied statistics such as variational Bayes depends on choosing an exponential family with suitable special characters.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the pointwise-regularity assumption can be relaxed to a.e.-defined densities using a version of conditional expectation, the same derivative formulas would extend to general Polish spaces; this is an extension the paper does not make.
  • Read geometrically, Eq. (10) says a Bayesian update moves the posterior in the direction of the residual of the joint score after removing its conditional mean—a new interpretation of the information contributed by an observation.
  • The same bundle calculus applies to any map between density models defined by integration or conditioning, so it could also serve marginalization in graphical models or transport-based inference, not just Bayes.
  • A direct test of the framework would be to implement the natural gradient in a nonparametric variational Bayes scheme where the approximating family is an infinite-dimensional exponential model and check whether the centered-score derivative matches a finite-difference estimate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This note presents the author's statistical-bundle formalism as a setting for nonparametric dually affine information geometry. After recalling the exponential and mixture charts, it states formulas for the total natural gradient of the KL divergence (Eqs. (6)-(7)), derives the derivative of the marginalization map as a conditional expectation (Eq. (9)) and the derivative of the conditioning map as a centered conditional score (Eq. (10)), and applies these to exponential families in Section 3.4. The paper is written as a concise research note; the abstract and final remark claim that these computations demonstrate the usefulness of the statistical-bundle formalism for Bayes computations and variational Bayes.

Significance. The derivations that are actually carried out in the note are correct and clean: Eq. (9) and Eq. (10) follow from the chart calculus, and the exponential-family example is internally consistent. If the gradient formulas in Eqs. (6)-(7) were proved in the note or available in a published source, the paper would be a useful, mostly self-contained reference for the nonparametric dually affine calculus of Bayes computations. In its current form, however, the central KL-gradient claim rests on an unpublished, in-progress reference [11], so the standalone significance is diminished. The note is also explicitly restricted to pointwise-defined densities (e.g., continuous densities on a bounded domain), and that limitation should be reflected in the abstract and final remarks.

major comments (3)
  1. [Section 2, Eqs. (6)-(7)] The total natural gradient formulas are the main load-bearing claim of the note, but they are asserted as a 'direct computation' and their derivation is deferred to [11], which the reference list describes as 'In progress'. These formulas are not decorative: Section 3.4 uses both equations to compute parametric KL derivatives, and the final remark presents the KL gradient as a main advantage of the formalism. As the manuscript stands, the central claim is not independently verifiable from the text. Please include a self-contained derivation from the chart definitions in Eqs. (3)-(4), or cite a published proof; this is the main point that must be fixed.
  2. [Section 3.2, Eq. (10)] The derivation of the conditioning derivative relies on pointwise evaluation of densities at a fixed observation x, as the text acknowledges ('well-defined in our setup because the densities are defined everywhere'). This restricts the nonparametric claim made in the abstract and final remark, since the usual L^1 setting identifies densities up to null sets. Please state this limitation at the first mention of the nonparametric setup and clarify whether Eq. (10) is independent of the chosen density version; if it is not, the scope of the claim should be reduced accordingly.
  3. [Section 3.4] The parametric gradient computations are displayed only after the phrase 'from eq. (14) with eq. (7) and eq. (6), respectively'. The reader must reconstruct the pairing step and the cancellation of the mean term to see why the displayed integrals have no E[T] contribution. Once Eqs. (6)-(7) are proved, this is a short chain-rule computation, but the note should show at least one line of that pairing so Section 3.4 is self-contained.
minor comments (5)
  1. [Reference [11]] The title of reference [11] contains a typo ('Kullback-Leible r' should be 'Kullback-Leibler'). More importantly, because this reference is listed as 'In progress', it cannot carry the derivation of Eqs. (6)-(7) in a published manuscript.
  2. [Section 1] The 'Weyl's axioms' are invoked without being stated. Since the note is meant as a concise review, either state the axioms or give a precise pointer to the relevant part of [5].
  3. [Section 1 and Section 2] The notation '★' for the Fisher score is used repeatedly but defined only implicitly in the displayed curve computation; an explicit definition would help readers who are not already familiar with the statistical-bundle formalism.
  4. [Section 2] The term 'total natural gradient' is introduced through the displayed identity with arbitrary curves, but existence and uniqueness of the pair (grad1, grad2) are not discussed. A sentence noting that these are the ordinary derivatives in the exponential and mixture charts would make the definition transparent and would also clarify the meaning of the 'direct computation'.
  5. [Throughout] There are several typographical artifacts in the text (for example, 'expession' in Section 3.1 and the rendering of some formulas as 'dsp' in the extracted text); these should be cleaned up in the final version.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the Bayes derivative computations are derived in-paper; deferred proof of Eqs. (6)-(7) is a self-citation/rigor concern, not a circular reduction.

full rationale

The paper's advertised Bayes computations do not reduce by construction to their inputs. Section 3.1 derives Eq. (9), identifying the derivative of marginalization with a conditional expectation, directly from the mixture-chart expression of the marginalization map. Section 3.2 derives Eq. (10), identifying the bundle derivative of conditioning with a centered score, by an explicit chain-rule calculation in the mixture chart. Neither derivation assumes the conclusion or fits a parameter to a target quantity. The KL gradient identities (6)-(7) are introduced as a 'direct computation' with the proof deferred to the author's in-progress manuscript [11]; this is a self-citation and rigor concern for a conference note, but not a circularity, because the displayed definitions (3)-(4) of the exponential and mixture charts and the KL divergence determine these derivatives, and the formulas are not fitted parameters or restatements of the conclusion. The references to [5] review the framework and are not used to import an unverified uniqueness theorem or to forbid alternatives. The restriction to pointwise-defined densities, stated in Section 1 and again in Section 3.2, is a genuine scope limitation but not a circular step. Overall, the derivation chain is self-contained for the main Bayes computations, with one deferred external-looking proof that lowers standalone verifiability but does not create circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The computations rest on the statistical bundle formalism from the author's own earlier papers. The main structural assumptions are pointwise-defined continuous densities and open exponential arcs between all densities; the main unproved ingredient is the gradient formula (6)-(7), deferred to an unpublished reference. There are no free parameters or invented entities.

assumptions (3)
  • domain assumption The maximal exponential model E(mu) contains all densities exp(u - psi(u)) * mu and S is the largest open convex domain such that an open exponential arc connects any two densities.
    Section 1: 'We assume S is the largest open, convex domain, and it is such that an open exponential arc connects all couples of densities.' This is a strong connectedness and regularity condition on the model space.
  • domain assumption Densities are defined pointwise everywhere, not just up to null sets, so conditioning p_{2|1}(y|x) is well defined for each fixed x.
    Section 3.2: 'is well-defined in our setup because the densities are defined everywhere'; Section 1 restricts Omega to finite sets or bounded domains with continuous functions.
  • ad hoc to paper The gradient formulas grad1 D(p||q) = -s_p(q) and grad2 D(p||q) = -eta_q(p) in Eqs. (6)-(7) hold as stated.
    Asserted in Section 2 with 'A detailed derivation will appear in [11]', an unpublished reference by the same author. The paper does not prove these formulas.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Information geometry of Bayes computations." pith.science (2026). https://pith.science/paper/BC4FP56B

@misc{pith2026250202160,
  author       = {Pith},
  title        = {Pith review of: Information geometry of Bayes computations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BC4FP56B}},
  note         = {Machine review of arXiv:2502.02160}
}
read the original abstract

Amari's Information Geometry is a dually affine formalism for parametric probability models. The literature proposes various nonparametric functional versions. Our approach uses classical Weyl's axioms so that the affine velocity of a one-parameter statistical model equals the classical Fisher's score. In the present note, we first offer a concise review of the notion of a statistical bundle as a set of couples of probability densities and Fisher's scores. Then, we show how the nonparametric dually affine setup deals with the basic Bayes and Kullback-Leibler divergence computations.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

14 extracted references · 12 canonical work pages

  1. [11]

    In progress

    Pistone, G.: Constrained minima of the Kullback-Leible r divergence (2025). In progress

  2. [5]

    Dually affine Information Geometry modeled on a Banach space

    Chirco, G., Pistone, G.: Dually affine Information Geometr y modeled on a Banach space (2022). DOI 10.48550/ARXIV.2204.00917. URL https://arxiv.org/abs/2204.00917. arXiv:2204.00917

  3. [1]

    American Mathematical Society (2000)

    Amari, S., Nagaoka, H.: Methods of information geometry. American Mathematical Society (2000). Translated from the 1993 Japanese original by Daishi Harada

  4. [2]

    Amari, S.i.: Information geometry and its applications, Applied Mathematical Sciences , vol

  5. [3]

    Arnold, V.I.: Mathematical methods of classical mechani cs, Graduate Texts in Mathematics, vol. 60. Springer-Verlag, New York (1989). Translated from the 1974 Russian original by K. Vogtmann and A. Weinstein, Corrected reprint of the second (1989) edition

  6. [4]

    Ay, N., Jost, J., L ˆe, H.V., Schwachh ¨ofer, L.: Information geometry, Ergebnisse der Math- ematik und ihrer Grenzgebiete. 3. Folge. A Series of Modern S urveys in Mathematics [Results in Mathematics and Related Areas. 3rd Series. A Ser ies of Modern Surveys in Mathematics], vol. 64. Springer, Cham (2017). DOI 10.1007/978-3-319-56 478-4. URL https://doi....

  7. [6]

    Efron, B., Hastie, T.: Computer age statistical inferenc e, Institute of Mathematical Statis- tics (IMS) Monographs , vol. 5. Cambridge University Press, New York (2016). URL https://doi.org/10.1017/CBO9781316576533. Algorithms, evidence, and data science

  8. [7]

    Acta Mat hematica Sinica, En- glish Series 22(4), 1175–1182 (2006)

    Egozcue, J.J., D ´ıaz–Barrero, J.L., Pawlowsky–Glahn, V.: Hilbert space of p robabil- ity density functions based on Aitchison geometry. Acta Mat hematica Sinica, En- glish Series 22(4), 1175–1182 (2006). DOI 10.1007/s10114-005-0678-2. UR L http://dx.doi.org/10.1007/s10114-005-0678-2

Show all 14 references
  1. [8]

    URL https://arxiv.org/abs/2411.03265

    Khesin, B., Misio lek, G., Modin, K.: Information geometry of diffeomorphism groups (2024). URL https://arxiv.org/abs/2411.03265

  2. [9]

    Pistone, G.: Nonparametric information geometry. In: F. Nielsen, F. Barbaresco (eds.) Geo- metric science of information, Lecture Notes in Comput. Sci. , vol. 8085, pp. 5–36. Springer, Heidelberg (2013). First International Conference, GSI 20 13 Paris, France, August 28-30, 20...

  3. [10]

    Pistone, G.: Statistical bundle of the transport model. In: F. Nielsen, F. Barbaresco (eds.) Geometric Science of Information, pp. 752–759. Springer In ternational Publishing, Cham (2021)

  4. [12]

    Pistone, G., Sempi, C.: An infinite-dimensional geometr ic structure on the space of all the probability measures equivalent to a given one. Ann. Statist. 23(5), 1543–1561 (1995)

  5. [13]

    URL https://arxiv.org/abs/1912.08003

    Talska, R., Menafoglio, A., Hron, K., Egozcue, J.J., Pal area-Albaladejo, J.: Changing ref- erence measure in bayes spaces with applications to functio nal data analysis (2019). URL https://arxiv.org/abs/1912.08003

  6. [194]

    URL https://doi.org/10.1007/978-4-431-55978-8

    Springer, [Tokyo] (2016). URL https://doi.org/10.1007/978-4-431-55978-8

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.