{"id":"94ecfc86-f760-4d54-97f6-ee26f08e8cc3","arxiv_id":"2502.02160","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In nonparametric information geometry, marginalization and conditioning in Bayes computations are realized as derivatives of maps between statistical bundles.","lead":"Statistical bundles, a geometric structure on spaces of probability distributions, let Bayes' rule and Kullback-Leibler divergence computations appear as derivative operations. The note offers a concise framework that may eventually give variational Bayes algorithms a cleaner geometric footing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gradient formulas (6)-(7) are the load-bearing core of the note's central claim, but their derivation is deferred to an unpublished in-progress paper [11]; without a proof or public reference the main assertion is not independently verifiable in this preprint.","rationale":"The note's Bayes computations (9)-(10) are derived in the text and, under the stated continuous-density assumption, they check out: direct differentiation of p_{2|1}(y|x)=p(x,y)/p_X(x) along a score s gives s(x,y)-E[s|X=x], matching Equation (10), and the marginalization derivative is the conditional expectation, matching Equation (9). So the novel computations in Section 3 are internally sound. The weakness is the gradient theorem in Section 2: Equations (6)-(7) are not derived here and are deferred to an unpublished manuscript. Since the strongest claim explicitly advertises these gradient formulas and Section 3.4 applies them, the central claim's support is incomplete. A direct re-derivation from the charts is straightforward and likely confirms the formulas, so I do not call them false; I call the preprint's self-containedness insufficient. The pointwise-density restriction raised by the reader is real but explicitly assumed in Section 1, so I treat it as a scope limitation rather than a hidden flaw; it would become decisive only if the paper claimed to cover L^1 densities. Recommendation: keep the reader's CONDITIONAL verdict; the paper should either include the derivation of (6)-(7) or replace the in-progress citation with a public source.","tokens_in":10098,"tokens_out":16365,"duration_ms":158149,"concrete_test":"Re-derive Equation (6) from definitions (3)-(4) by taking an arbitrary smooth curve p_t in E(mu) with score s_t and computing d/dt D(p_t || q) = d/dt E_{p_t}[log(p_t/q)]; express the result as -<s_t, u_{p_t}(q)>_{p_t}. Similarly recover Equation (7) by varying the second argument and expressing the derivative as -<eta_{q_t}(p), s_{q_t}>_{q_t}. Run the same computation on a two-point sample space with explicit densities; if either identity fails, the central gradient claim of the note collapses. If both identities reproduce, the concern is only the missing self-contained proof and the verdict should remain CONDITIONAL pending a public derivation of (6)-(7).","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equations (6)-(7) state the total natural gradient of the KL divergence. They are introduced as a 'direct computation' but the derivation is relegated to reference [11], which is described as 'in progress.' These formulas are used in Section 3.4 to compute parametric KL gradients, so the note's advertised Bayes computations depend on them. The definitions (3)-(4) of exponential and mixture charts are sufficient to verify them by a finite-dimensional chain rule, so the gap is plausibly only presentational. However, as a standalone argument, the central claim rests on an unproven external result that the reader cannot check. A secondary scope limitation is explicit: Section 1 restricts to continuous pointwise-defined densities, so Equation (10)'s pointwise conditioning does not extend to the usual L^1 setting without a regular version; this is acknowledged in the text and is a limitation rather than an inconsistency.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This note presents the author's statistical-bundle formalism as a setting for nonparametric dually affine information geometry. After recalling the exponential and mixture charts, it states formulas for the total natural gradient of the KL divergence (Eqs. (6)-(7)), derives the derivative of the marginalization map as a conditional expectation (Eq. (9)) and the derivative of the conditioning map as a centered conditional score (Eq. (10)), and applies these to exponential families in Section 3.4. The paper is written as a concise research note; the abstract and final remark claim that these computations demonstrate the usefulness of the statistical-bundle formalism for Bayes computations and variational Bayes.","tokens_in":10269,"tokens_out":8404,"duration_ms":84788,"significance":"The derivations that are actually carried out in the note are correct and clean: Eq. (9) and Eq. (10) follow from the chart calculus, and the exponential-family example is internally consistent. If the gradient formulas in Eqs. (6)-(7) were proved in the note or available in a published source, the paper would be a useful, mostly self-contained reference for the nonparametric dually affine calculus of Bayes computations. In its current form, however, the central KL-gradient claim rests on an unpublished, in-progress reference [11], so the standalone significance is diminished. The note is also explicitly restricted to pointwise-defined densities (e.g., continuous densities on a bounded domain), and that limitation should be reflected in the abstract and final remarks.","major_comments":[{"comment":"The total natural gradient formulas are the main load-bearing claim of the note, but they are asserted as a 'direct computation' and their derivation is deferred to [11], which the reference list describes as 'In progress'. These formulas are not decorative: Section 3.4 uses both equations to compute parametric KL derivatives, and the final remark presents the KL gradient as a main advantage of the formalism. As the manuscript stands, the central claim is not independently verifiable from the text. Please include a self-contained derivation from the chart definitions in Eqs. (3)-(4), or cite a published proof; this is the main point that must be fixed.","section":"Section 2, Eqs. (6)-(7)"},{"comment":"The derivation of the conditioning derivative relies on pointwise evaluation of densities at a fixed observation x, as the text acknowledges ('well-defined in our setup because the densities are defined everywhere'). This restricts the nonparametric claim made in the abstract and final remark, since the usual L^1 setting identifies densities up to null sets. Please state this limitation at the first mention of the nonparametric setup and clarify whether Eq. (10) is independent of the chosen density version; if it is not, the scope of the claim should be reduced accordingly.","section":"Section 3.2, Eq. (10)"},{"comment":"The parametric gradient computations are displayed only after the phrase 'from eq. (14) with eq. (7) and eq. (6), respectively'. The reader must reconstruct the pairing step and the cancellation of the mean term to see why the displayed integrals have no E[T] contribution. Once Eqs. (6)-(7) are proved, this is a short chain-rule computation, but the note should show at least one line of that pairing so Section 3.4 is self-contained.","section":"Section 3.4"}],"minor_comments":[{"comment":"The title of reference [11] contains a typo ('Kullback-Leible r' should be 'Kullback-Leibler'). More importantly, because this reference is listed as 'In progress', it cannot carry the derivation of Eqs. (6)-(7) in a published manuscript.","section":"Reference [11]"},{"comment":"The 'Weyl's axioms' are invoked without being stated. Since the note is meant as a concise review, either state the axioms or give a precise pointer to the relevant part of [5].","section":"Section 1"},{"comment":"The notation '★' for the Fisher score is used repeatedly but defined only implicitly in the displayed curve computation; an explicit definition would help readers who are not already familiar with the statistical-bundle formalism.","section":"Section 1 and Section 2"},{"comment":"The term 'total natural gradient' is introduced through the displayed identity with arbitrary curves, but existence and uniqueness of the pair (grad1, grad2) are not discussed. A sentence noting that these are the ordinary derivatives in the exponential and mixture charts would make the definition transparent and would also clarify the meaning of the 'direct computation'.","section":"Section 2"},{"comment":"There are several typographical artifacts in the text (for example, 'expession' in Section 3.1 and the rendering of some formulas as 'dsp' in the extracted text); these should be cleaned up in the final version.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is best read as a research note that synthesizes the author's own tutorial [5] and Amari's textbook [2], with new organizational value in showing how Bayes computations look in the statistical-bundle formalism. The main obstacle to publication is that the central gradient formulas in Eqs. (6)-(7) are deferred to the author's own in-progress paper [11]; the editor may wish to require that the proof be included in this note before acceptance. The paper fits the scope of a short mathematical statistics note if that gap is closed and the pointwise-density limitation is clearly stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Giovanni Pistone's note shows that the statistical bundle formalism, which he has developed over a series of papers, gives a clean dually affine setting for basic Bayes computations. What is actually new are the derivative formulas: the bundle derivative of marginalization is a conditional expectation (Eq. 9), the derivative of conditioning at a fixed x is a centered score (Eq. 10), and the parametric examples in Section 3.4 are worked out in detail. Those derivations are carried out in the text and they look correct. The exponential-family example at the end is a nice concrete illustration.\n\nThe soft spot is almost exactly where the stress-test puts it. Equations (6)-(7), the total natural gradients of the KL divergence, are the load-bearing core of the note's advertised 'Bayes computations.' They are asserted as a 'direct computation' and deferred to reference [11], which is an in-progress paper by the author. So as a standalone note, the main claim is not independently verifiable. I think the stress-test is right that a finite-dimensional chain rule using the definitions (3)-(4) is sufficient to check them, which makes this a presentational gap rather than a fatal one. But the gap is real: a reader cannot check the central equations from this note alone. The fix is straightforward: prove (6)-(7) in a page, or cite a publicly available version of [11].\n\nThe second limitation is the pointwise-defined densities setting. Section 1 restricts Omega to a finite set or bounded domain with continuous densities, which makes conditioning at a fixed x well-defined. The paper says this explicitly, so it's an acknowledged scope restriction rather than a hidden assumption. It does mean Eq. (10) does not automatically extend to L^1 densities without a regular version.\n\nThe citation pattern is heavy on the author's own tutorial [5], but that is appropriate because the formalism comes from there. The derivations in this note are independent of [11] except for (6)-(7).\n\nThis paper is for information geometers and statisticians working on variational Bayes who want a nonparametric language for KL gradients. It is a legitimate short note, not a breakthrough and not a sloppy one. The correct next step is to send it to review and ask the author to close the gap on (6)-(7). I would accept it for that purpose.","headline":"Concise note with correct-looking Bayes computations, but the core KL gradient formulas are deferred to an unpublished paper; fixable and worth refereeing.","tokens_in":10777,"tokens_out":2946,"would_cite":false,"duration_ms":26482,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62B05","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"This note argues that the nonparametric dually affine statistical bundle formalism—a Banach-space version of classical information geometry in which the Fisher score is a velocity—is the natural setting for Bayes computations, yielding…","keywords":["information geometry","statistical bundle","dually affine structure","Bayes computation","Kullback-Leibler divergence","conditional expectation","Fisher score","nonparametric inference"],"falsifier":"A concrete test is to represent a joint density by two different versions that agree almost everywhere but differ at a point $x_0$, and check whether Eq. (10) produces the same conditional score; if the two versions give different results, the formula is not well-defined in the standard $L^1$ equivalence-class setting, showing the central claim depends on the pointwise-everywhere regularity assumption.","tokens_in":9844,"feed_emoji":"📐","tokens_out":11550,"duration_ms":104261,"temperature":0.7,"pith_summary":"This note argues that the nonparametric dually affine statistical bundle formalism—a Banach-space version of classical information geometry in which the Fisher score is the velocity of a curve—is the natural setting for Bayes computations. It establishes that the derivative of the marginalization map is a conditional expectation, $d M_\\#(p)[s_p] = E_p[s_p \\mid \\pi]$, and that the bundle derivative of conditioning at a fixed observation is a centered score, $d K_x(p)[s_p] = s_p(x,\\cdot) - E_{p_{2|1}(\\cdot|x)}[s_p(x,\\cdot)]$. It also derives the total natural gradient of the Kullback-Leibler divergence as $\\operatorname{grad}_1 D(p\\|q) = -s_p(q)$ and $\\operatorname{grad}_2 D(p\\|q) = -\\eta_q(p)$, i.e., the negative exponential and mixture charts. If these claims hold, Bayesian calculations can be understood geometrically without finite-dimensional parametric assumptions, offering coordinate-free tools for variational Bayes and related applied statistics.","feed_headline":"Bayesian conditioning is a centered score in this geometry","feed_subtitle":"A nonparametric geometry turns Bayes updates into conditional expectations and centered scores.","key_machinery":"The central object is the statistical bundle $S E(\\mu) = \\{(p, v) : p \\in E(\\mu), v \\in B, E_p[v] = 0\\}$, a vector bundle over the maximal exponential model $E(\\mu) = \\{ e^{u - \\psi(u)} \\mu : u \\in S \\subset B\\}$, whose fibers are the tangent spaces of centered random variables. Two affine charts carry the argument: the exponential chart $s_p(q) = \\log(q/p) - E_p[\\log(q/p)]$ and the mixture chart $\\eta_p(q) = q/p - 1$, together with the parallel transports $eU^p_q v = v - E_p[v]$ and $mU^p_q w = (q/p) w$. The Weyl cocycle identities make these charts a genuine affine atlas, so the Fisher score is the velocity in the moving frame. All of the paper's identities—the KL gradients, marginalization as conditional expectation, and conditioning as centered score—are direct computations in these charts.","core_discovery":"The paper's central claim is that every basic Bayes operation has a clean expression in the statistical bundle over the maximal exponential model. For a joint density $p_{12}$, the velocity of marginalization is the conditional expectation of the joint score given the first coordinate, Eq. (9). The velocity of conditioning at a point $x$ is the joint score evaluated at $x$ minus its conditional average under the conditional density, Eq. (10) — in other words, the centered score. The KL divergence is the natural divergence of the affine charts: the first natural gradient is the negative exponential chart $s_p(q)$, and the second is the negative mixture chart $\\eta_q(p)$, Eqs. (6)-(7). An exponential factorization identity and the KL chain rule follow from the same charts, and a worked exponential-family example recovers the classical score formulas as special cases.","pith_inferences":["If the pointwise-regularity assumption can be relaxed to a.e.-defined densities using a version of conditional expectation, the same derivative formulas would extend to general Polish spaces; this is an extension the paper does not make.","Read geometrically, Eq. (10) says a Bayesian update moves the posterior in the direction of the residual of the joint score after removing its conditional mean—a new interpretation of the information contributed by an observation.","The same bundle calculus applies to any map between density models defined by integration or conditioning, so it could also serve marginalization in graphical models or transport-based inference, not just Bayes.","A direct test of the framework would be to implement the natural gradient in a nonparametric variational Bayes scheme where the approximating family is an infinite-dimensional exponential model and check whether the centered-score derivative matches a finite-difference estimate."],"forward_implications":["The marginalization derivative (Eq. 9) reveals the conditional expectation operator as the tangent map of a geometric projection, so Bayesian marginalization becomes a first-order calculus operation.","The gradient formulas (Eqs. 6-7) yield a coordinate-free natural gradient for optimizing over densities, which can be used directly in variational Bayes and other divergence-based inference.","The exponential decomposition of the joint density gives a chart-based proof of the KL chain rule, the additive structure central to evidence lower bound computations.","For exponential families, the derived formulas reduce to the classical facts that the score is the centered sufficient statistic and the posterior score is its conditional expectation, so the bundle calculus subsumes the parametric results.","The paper's final remark notes that applying the derived differential equations in applied statistics such as variational Bayes depends on choosing an exponential family with suitable special characters."],"supporting_citations":[{"why":"Supplies the tutorial construction of the statistical bundle, its charts, and parallel transports that the paper builds on.","marker":"[5]"},{"why":"Provides the dually affine formalism, the natural gradient definition, and the Bayes-computation section that motivates the examples.","marker":"[2]"},{"why":"Classical reference for the dually affine information geometry that the nonparametric bundle generalizes.","marker":"[1]"},{"why":"Gives the Fisher score and exponential-family background used in the worked example.","marker":"[6]"},{"why":"The paper defers the detailed derivation of the KL gradient formulas to this companion work.","marker":"[11]"},{"why":"Establishes the infinite-dimensional geometric structure on densities equivalent to a given measure, foundational for the nonparametric model.","marker":"[12]"}],"fun_headline_variants":["Bayes computations as centered Fisher scores","Geometry of Bayes: conditioning equals centered score","Nonparametric geometry for Bayes and KL divergence","Centered scores unify Bayes operations","How Bayes updates become geometric scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that densities in the model are defined pointwise everywhere on the outcome space, so that conditioning at a fixed observation $x$ is meaningful; if densities are only defined up to null sets, the conditioning map and the derivative formula in Eq. (10) require a version or fail.","fun_headline_variants_meta":{"raw":{"variants":["Bayes computations as centered Fisher scores","Geometry of Bayes: conditioning equals centered score","Nonparametric geometry for Bayes and KL divergence","Centered scores unify Bayes operations","How Bayes updates become geometric scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000374,"raw_usage":{"total_tokens":1915,"prompt_tokens":779,"completion_tokens":1136,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":395,"completion_tokens_details":{"reasoning_tokens":1074}},"tokens_in":395,"tokens_out":1136,"duration_ms":10688,"temperature":1.0,"reasoning_tokens":1074,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:06:22.892048+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test is to represent a joint density by two different versions that agree almost everywhere but differ at a point $x_0$, and check whether Eq. (10) produces the same conditional score; if the two versions give different results, the formula is not well-defined in the standard $L^1$ equivalence-class setting, showing the central claim depends on the pointwise-everywhere regularity assumption.","supporting_citations":[{"cited_title":"Dually affine Information Geometry modeled on a Banach space","cited_arxiv_id":"2204.00917","evidence_quote":"Supplies the tutorial construction of the statistical bundle, its charts, and parallel transports that the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the dually affine formalism, the natural gradient definition, and the Bayes-computation section that motivates the examples."},{"cited_title":"American Mathematical Society (2000)","cited_arxiv_id":null,"evidence_quote":"Classical reference for the dually affine information geometry that the nonparametric bundle generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Fisher score and exponential-family background used in the worked example."},{"cited_title":"In progress","cited_arxiv_id":null,"evidence_quote":"The paper defers the detailed derivation of the KL gradient formulas to this companion work."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the infinite-dimensional geometric structure on densities equivalent to a given measure, foundational for the nonparametric model."}],"review_version":1}