Pith. sign in

REVIEW 2 major objections 5 minor 28 references

Parametric models without densities, scores, or Fisher information can still carry Riemannian metrics once laws are tempered distributions and instruments extract Godambe information.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 06:01 UTC pith:6I4LOVEG

load-bearing objection Solid, usable extension of information geometry via Godambe instruments on tempered distributions; examples carry the weight and the math is clean under stated premises. the 2 major comments →

arxiv 2607.11246 v1 pith:6I4LOVEG submitted 2026-07-13 math.ST stat.MEstat.TH

Weak Information Geometry: Riemannian Structures from Distributional Inference Functions and Stein Discrepancies

classification math.ST stat.MEstat.TH MSC 62B1162F1053C2060G52
keywords information geometryGodambe informationtempered distributionsweak inference functionsStein discrepanciesFisher–Rao manifoldnonformationalpha-stable noise
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Classical information geometry needs densities, scores, and finite Fisher information, so many models—singular continuous laws, parameter-dependent supports, transform-only specifications, and even some dominated mixtures with biased scores—are left outside the Riemannian picture. This paper replaces densities by tempered distributions and extracts information with external instruments (weak regular inference functions or weak Stein representations). Any instrument with full-rank sensitivity and positive-definite variability induces a smooth Godambe metric on the parameter space; the Fisher–Rao metric is recovered exactly when the score itself is admissible, and every Godambe metric sits below the Fisher metric in the Loewner order whenever the latter exists. Five concrete examples, including a Cantor location family that admits no dominating measure and a lattice heat equation driven by alpha-stable noise, show that the construction works where Fisher–Rao cannot even begin. Because no instrument is canonical, a model carries a family of metrics whose members serve different inferential, diagnostic, geometric and computational roles, while weak inferential separation appears as block-diagonality of the metric.

Core claim

Once a parametric family is represented by tempered distributions T_θ, any instrument (weak regular inference function or weak Stein representation) whose sensitivity S is full-rank and whose variability V is positive definite induces the Godambe information G(θ) = S(θ)ᵀ V(θ)⁻¹ S(θ) as a smooth Riemannian metric on the parameter space. The Fisher–Rao manifold is the special case in which the score is itself an admissible instrument; whenever Fisher information exists, every Godambe metric is dominated by it in the Loewner order.

What carries the argument

The Godambe information G = Sᵀ V⁻¹ S built from the weak sensitivity and variability of an instrument acting on tempered distributions; it is the metric tensor that turns the parameter space into a Godambe–Riemannian manifold and recovers Fisher information when the instrument is the score.

Load-bearing premise

The map from parameters to the distributional representation must be weakly differentiable: for every admissible probe the expectation must be continuously differentiable and derivatives must pass inside the pairing.

What would settle it

Exhibit a concrete parametric family of tempered distributions together with an instrument whose sensitivity and variability are smooth, full-rank and positive-definite, yet the resulting G fails to be a Riemannian metric, or show that the claimed Loewner domination G ⪯ I fails when Fisher information exists and the Bartlett interchange holds.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Undominated families such as the Cantor location model, models with parameter-dependent support, and mixtures whose score is biased all become Riemannian manifolds once a suitable instrument is chosen.
  • Quadratic Stein discrepancies built from finite collections of identities induce exactly the same local geometry as the corresponding Godambe metrics.
  • Weak inferential separation (nonformation) is equivalent to block-diagonality of the Godambe metric with respect to the interest–nuisance splitting.
  • In dynamical models such as the lattice stochastic heat equation driven by alpha-stable noise, the Godambe geometry stabilises at a rate controlled by the spectral gap of the discrete Laplacian.
  • Because there is no canonical instrument, a single model carries a family of metrics that can be selected according to inferential, diagnostic, geometric or computational purpose.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same instrument construction could be applied to other singular or heavy-tailed SPDE models to obtain closed-form geometric stability rates without second-moment assumptions.
  • Hierarchy of RKHS Stein geometries suggests a practical route to adaptive metric selection: start with low-frequency or moment instruments and enrich only when diagnostics show near-degeneracy.
  • Automatic block-diagonality for odd/even probes in every symmetric location-scale family (including Cauchy and Cantor) supplies an immediate, likelihood-free method for exact location–scale separation in robust estimation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper argues that parametric models represented by tempered distributions T_θ ∈ S'(R^k) can be equipped with Riemannian metrics far beyond the classical Fisher–Rao setting. An external instrument (weak regular inference function or weak Stein representation) with full-rank sensitivity S and positive-definite variability V induces the Godambe information G = S^T V^{-1} S as a smooth metric on Θ (Theorem 1.1 / Proposition 3.1). Fisher–Rao is recovered when the score is admissible, and G ⪯ I in Loewner order whenever Fisher exists (Proposition 3.2). Five examples outside Fisher–Rao are treated in closed form: uniform scale, shifted exponential, undominated Cantor location, stratified mixture with biased score, and a lattice α-stable heat equation whose Godambe metric stabilises at the spectral-gap rate. Quadratic Stein discrepancies recover the same local geometry; non-canonicity of the instrument and block-diagonality as geometric nonformation are discussed.

Significance. If the construction is accepted, the class of models that can be treated by differential-geometric methods expands substantially: undominated families, parameter-dependent support, moment-free laws, and dominated models with biased scores all become Riemannian. The closed-form metrics (uniform G=3/θ², Cantor G≡8, SPDE spectral-gap rate αθλ_k) and the explicit Loewner comparison are concrete, usable contributions. The paper is honest about non-canonicity (Section 7, Remark 7.1) and about the instrument-relative character of the geometry. Dependence on the author’s 2026 companion series for the full apparatus of weak inference functions and genericity is a presentational limitation rather than a correctness flaw in the derivations as written.

major comments (2)
  1. Assumption 2.3 (weak differentiability of θ ↦ ⟨T_θ, g⟩ with derivatives passing inside the pairing) is the analytic hinge for the weak Bartlett identity (1), for S(θ)=⟨∂_θ T_θ, ψ⟩, and thus for smoothness of G in Theorem 1.1 / Proposition 3.1. It is stated as an assumption and verified only by direct computation in the five examples; no general criterion is given for when a tempered-distribution model map satisfies it. For the central claim to be usable beyond the exhibited cases, the paper should either supply verifiable sufficient conditions (e.g., dominated differentiability of characteristic functions, or continuity of the map in the weak-* topology of S') or state more explicitly that the theorem is conditional on this hypothesis and that existence is left to the companion [25].
  2. Proposition 1.2 and the examples establish existence of instruments case-by-case, but the paper repeatedly speaks of “the class of Godambe–Riemannian models” as if it were characterised. Without a general existence theorem (or a clear statement that none is claimed), the scope of the extension remains an open-ended collection of examples. A short subsection clarifying what is proved versus what is exhibited would prevent over-reading of Theorem 1.1.
minor comments (5)
  1. Heavy self-citation to the 2026 companion series ([21]–[25]) makes the paper hard to read in isolation. A short self-contained appendix restating Definition 2.2 and the admissible classes of Lemma 2.1 would help.
  2. Notation: α and β are used both for the stability index / eigenmode scales (Section 5) and for interest/nuisance parameters (Section 8). A local warning is present but easy to miss; consider distinct symbols.
  3. Remark 4.1 (Student-t / Cauchy) and Remark 4.9 (Cauchy mixture) are useful but sit outside the main example sections; a pointer in the introduction would improve navigation.
  4. Typographical: “F amily of metrics” in the Contents (Section 7 heading) has a stray space; “P´ olya” and similar accented names appear inconsistently.
  5. Section 5.6 claims numerical verification of the closed-form CF and of Proposition 5.1 against Monte Carlo, but no figure or table is supplied. A brief plot of |G_k(t)−G_k(∞)| versus the predicted rate would strengthen the claim.

Circularity Check

1 steps flagged

Minor self-citation to same-author 2026 companions for the distributional apparatus and genericity; central Godambe-metric construction, Loewner comparison, and explicit examples remain independently proved here.

specific steps
  1. self citation load bearing [Section 2 (Preliminaries) and Introduction (references to companions for instruments and genericity)]
    "The systematic development of weak regular inference functions is carried out in the companion paper [23]; here we record only what is needed. ... The companion paper [25] shows that such failures are exceptional in a precise sense: within natural families of instruments, those violating the hypotheses form a negligible set..."

    The foundational notions of weak regular inference functions, observation operators, and the claim that non-degeneracy of (S,V) is generic are justified by citation to same-author arXiv companions rather than derived or machine-checked here. While the paper supplies self-contained definitions and verifies the hypotheses by hand in its examples, the broader apparatus that makes the instrument class well-defined for arbitrary tempered distributions rests on the author-controlled stack; this is a mild load-bearing self-citation but does not force the metric theorems by construction.

full rationale

The paper's core derivation (Proposition 3.1 that G = SᵀV⁻¹S is a smooth Riemannian metric under full-rank S and positive-definite V; Proposition 3.2 Loewner comparison recovering Fisher–Rao for the score; Theorem 1.1) is classical linear algebra plus the weak Bartlett identity under the stated Assumption 2.3, all proved in-place from the definitions given in Section 2. Existence of instruments outside Fisher–Rao is established by five fully explicit closed-form calculations (uniform, shifted exponential, Cantor, stratified mixture, lattice α-stable SPDE) that do not rely on external results. Self-citations to the author's companion series [21–25] supply background on weak moments, observation operators, and transversality/genericity, but these are not load-bearing for the metric claims themselves: the paper recalls the needed definitions, verifies non-degeneracy by direct computation in every example, and never treats a self-cited uniqueness or existence theorem as forcing the present results. No self-definitional loop, no fitted-input-as-prediction, no smuggled ansatz, and no renaming of a known empirical pattern appear. The dependence is therefore ordinary presentational scaffolding rather than circularity of the derivation chain.

Axiom & Free-Parameter Ledger

1 free parameters · 6 axioms · 2 invented entities

The central theorem rests on standard differential geometry and estimating-equation algebra once a distributional model map and an instrument are given. Load-bearing nonstandard inputs are the tempered-distribution representation of laws, the weak (pairing-based) regularity of inference functions, weak differentiability of θ ↦ T_θ, and the existence of instruments with full-rank S and PD V outside the Fisher class—illustrated by examples rather than proved for all models. No numerical free parameters are fitted; instrument frequencies are methodological choices. Invented framing entities (instrument, Godambe–Riemannian manifold) reorganize existing estimating-function and Stein objects rather than postulating new physical degrees of freedom.

free parameters (1)
  • Instrument tuning (e.g. frequency c or t in sinusoidal/transform probes)
    Not fitted to data in the paper, but each choice induces a different metric G_c or G_t; the theory deliberately has no canonical instrument, so geometry depends on user-selected probe parameters.
axioms (6)
  • domain assumption Laws P_θ are represented by tempered distributions T_θ ∈ S'(R^k) via ⟨T_θ,g⟩=E_{P_θ}[g] for g in Schwartz space (and extensions of Lemma 2.1).
    Foundation of the whole framework (§2.1); standard for probability measures with polynomial-growth moments of test functions, but the paper’s reach depends on working in this dual space rather than densities.
  • domain assumption Assumption 2.3: weak C¹ differentiability of θ ↦ ⟨T_θ,g⟩ for admissible probes, with derivatives realized as pairings against ∂_{θ_j} T_θ.
    Used for weak Bartlett identity (1), sensitivity definitions, Stein local expansions, and nonformation geometry; assumed throughout rather than derived from first principles for arbitrary families.
  • domain assumption Existence of at least one weak regular inference function (or Stein-generated instrument) with full column-rank S and positive-definite V (hypotheses of Prop. 3.1 / Thm 1.1).
    The metric exists only when such an instrument is available; Prop. 1.2 and §§4–5 exhibit instruments case-by-case; genericity is deferred to companion [25].
  • standard math Positive-definiteness of a quadratic form aᵀGa when V≻0 and Sa≠0 (standard linear algebra).
    Core of Proposition 3.1’s proof that G is a Riemannian metric.
  • domain assumption Bartlett-type interchange E[ψ u_θᵀ]=S when densities and scores exist (condition (2) in Prop. 3.2).
    Needed for Loewner comparison G ⪯ I; classical under regularity but not automatic for every weak instrument.
  • ad hoc to paper Fréchet differentiability of the embedded model map into H* for RKHS Stein metrics (Assumption 6.2).
    Extra regularity imposed only for the hierarchy beyond finite-dimensional quadratic Stein discrepancies (§6.3).
invented entities (2)
  • Instrument (positive Schwartz kernel / weak regular inference function / weak Stein representation as measurement device external to the law) independent evidence
    purpose: Extracts information from T_θ without being part of the law, enabling G when densities or scores fail.
    Framing of classical estimating functions and Stein operators in distributional language; independent classical evidence exists (Godambe, Stein), but the ‘instrument vs law’ reading is paper-specific.
  • Godambe–Riemannian manifold (Θ,G) no independent evidence
    purpose: Name the parameter space equipped with Godambe information as a Riemannian metric outside Fisher–Rao.
    Definitional packaging of G=SᵀV⁻¹S; not a new physical object, but the paper’s central geometric object.

pith-pipeline@v1.1.0-grok45 · 38384 in / 4014 out tokens · 48065 ms · 2026-07-14T06:01:56.031341+00:00 · methodology

0 comments
read the original abstract

The class of parametric statistical models that can be treated as Riemannian manifolds is considerably larger than the classical Fisher-Rao setting allows, once one works in the space of tempered distributions. A law is represented by a tempered distribution T in S'(R^k), while an instrument - a positive Schwartz kernel, a weak regular inference function, or a weak Stein representation - extracts information from the law without being part of it. Any instrument with full-rank sensitivity and positive-definite variability induces the Godambe information G = S^T V^{-1} S, a Riemannian metric on the parameter space; the Fisher-Rao manifold is recovered exactly when the score is an admissible instrument, and every Godambe metric is dominated by the Fisher metric in the Loewner order whenever the latter exists. Four examples lie outside the Fisher-Rao class for four different reasons: a location model built on the Cantor distribution (an undominated family - no likelihood, no score, and no Fisher information exist at all), the uniform scale model (parameter-dependent support), the shifted exponential model (transform-based inference), and a stratified finite mixture (a provably biased score in a dominated model); a lattice stochastic heat equation driven by alpha-stable noise provides a fifth, dynamical example, whose closed-form weak Godambe information stabilises at a rate governed by the spectral gap of the discrete Laplacian. Quadratic Stein discrepancies induce the same local geometry, and reproducing-kernel constructions generate a hierarchy of geometries. Because there is no canonical instrument, the model carries a family of Godambe metrics; we discuss the inferential, diagnostic, geometric, and computational roles of its members, and show that weak inferential separation (nonformation) appears geometrically as block-diagonality of the Godambe metric.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

28 extracted references · 4 linked inside Pith

  1. [1]

    (1985).Differential-Geometrical Methods in Statistics

    Amari, S. (1985).Differential-Geometrical Methods in Statistics. Lec- ture Notes in Statistics 28, Springer

  2. [2]

    Amari, S. (1998). Natural gradient works efficiently in learning.Neural Computation, 10(2), 251–276

  3. [3]

    and Kawanabe, M

    Amari, S. and Kawanabe, M. (1997). Information geometry of esti- mating functions in semi-parametric statistical models.Bernoulli, 3(1), 29–54

  4. [4]

    and Nagaoka, H

    Amari, S. and Nagaoka, H. (2000).Methods of Information Geometry. Translations of Mathematical Monographs 191, AMS/Oxford

  5. [5]

    and Schwachh¨ ofer, L

    Ay, N., Jost, J., Lˆ e, H.V. and Schwachh¨ ofer, L. (2017).Informa- tion Geometry. Ergebnisse der Mathematik und ihrer Grenzgebiete 64, Springer

  6. [6]

    (1978).Information and Exponential Families in Statistical Theory

    Barndorff-Nielsen, O.E. (1978).Information and Exponential Families in Statistical Theory. Wiley, Chichester

  7. [7]

    and Reid, N

    Barndorff-Nielsen, O.E., Cox, D.R. and Reid, N. (1986). The role of dif- ferential geometry in statistical theory.International Statistical Review, 54(1), 83–96

  8. [8]

    and Mackey, L

    Barp, A., Briol, F.-X., Duncan, A.B., Girolami, M. and Mackey, L. (2019). Minimum Stein discrepancy estimators.Advances in Neural In- formation Processing Systems, 32, 12964–12976

  9. [9]

    and Florens, J.-P

    Carrasco, M. and Florens, J.-P. (2000). Generalization of GMM to a continuum of moment conditions.Econometric Theory, 16(6), 797–834

  10. [10]

    (1982).Statistical Decision Rules and Optimal Inference

    ˇCencov, N.N. (1982).Statistical Decision Rules and Optimal Inference. Translations of Mathematical Monographs 53, American Mathematical Society, Providence. (Russian original, 1972.)

  11. [11]

    and Gretton, A

    Chwialkowski, K., Strathmann, H. and Gretton, A. (2016). A kernel test of goodness of fit.Proceedings of the 33rd International Conference on Machine Learning, 2606–2615. 51

  12. [12]

    and Salmon, M

    Critchley, F., Marriott, P. and Salmon, M. (1993). Preferred point ge- ometry and statistical manifolds.The Annals of Statistics, 21(3), 1197– 1224

  13. [13]

    (2003).Fractal Geometry: Mathematical Foundations and Applications

    Falconer, K. (2003).Fractal Geometry: Mathematical Foundations and Applications. 2nd edition, Wiley, Chichester

  14. [14]

    and McDunnough, P

    Feuerverger, A. and McDunnough, P. (1981). On the efficiency of empir- ical characteristic function procedures.Journal of the Royal Statistical Society, Series B, 43(1), 20–27

  15. [15]

    Godambe, V.P. (1960). An optimum property of regular maximum like- lihood estimation.The Annals of Mathematical Statistics, 31(4), 1208– 1211

  16. [16]

    and Mackey, L

    Gorham, J. and Mackey, L. (2017). Measuring sample quality with Stein’s method.The Annals of Statistics, 45(3), 1069–1109

  17. [17]

    Heathcote, C.R. (1977). The integrated squared error estimation of pa- rameters.Biometrika, 64(2), 255–264

  18. [18]

    and Labouriau, R

    Jørgensen, B. and Labouriau, R. (2012).Exponential Families and Theoretical Inference. Monografias de Matem´ atica 52, Instituto de Matem´ atica Pura e Aplicada (IMPA), Rio de Janeiro

  19. [19]

    (1996).Estimating Functions and Semiparametric Mod- els

    Labouriau, R. (1996).Estimating Functions and Semiparametric Mod- els. Ph.D. thesis, University of Aarhus

  20. [20]

    Labouriau, R. (2023). On the bias of the score function of finite mixture models.Communications in Statistics – Theory and Methods, 52(13). doi:10.1080/03610926.2021.1995429. (arXiv:2002.03307.)

  21. [21]

    Labouriau (2026A).Distributional Statistical Models: Weak Mo- ments, Cumulants, and a Central Limit Theorem, arXiv:2604.20634 [math.PR]

    R. Labouriau (2026A).Distributional Statistical Models: Weak Mo- ments, Cumulants, and a Central Limit Theorem, arXiv:2604.20634 [math.PR]

  22. [22]

    Labouriau (2026B)Weak Moment Methods for Statistical Inference: with an Application to Robust Estimation, arXiv:2604.23619 [stat.ME]

    R. Labouriau (2026B)Weak Moment Methods for Statistical Inference: with an Application to Robust Estimation, arXiv:2604.23619 [stat.ME]

  23. [23]

    Labouriau (2026C).Inference Functionals and Observation Opera- tors for DistributionalStatistical Models

    R. Labouriau (2026C).Inference Functionals and Observation Opera- tors for DistributionalStatistical Models. arXiv:2605.19189 [math.ST]

  24. [24]

    Labouriau (2026D).Weak Stein Discrepancies: Kernel-Regularised Goodness-of-Fit and Minimum Discrepancy Estimation for Heavy- Tailed Models, in preparation, 2026

    R. Labouriau (2026D).Weak Stein Discrepancies: Kernel-Regularised Goodness-of-Fit and Minimum Discrepancy Estimation for Heavy- Tailed Models, in preparation, 2026. 52

  25. [25]

    Labouriau (2026)

    R. Labouriau (2026). Transversality and Geometric Regularisation in Distributional Statistical Models. arXiv:2605.04536 [math.ST]

  26. [26]

    and Jordan, M

    Liu, Q., Lee, J. and Jordan, M. (2016). A kernelized Stein discrepancy for goodness-of-fit tests.Proceedings of the 33rd International Confer- ence on Machine Learning, 276–284

  27. [27]

    (2009).Algebraic Geometry and Statistical Learning The- ory

    Watanabe, S. (2009).Algebraic Geometry and Statistical Learning The- ory. Cambridge University Press

  28. [28]

    and Amari, S

    Watanabe, S. and Amari, S. (2003). Learning coefficients of layered models when the true distribution is not in the model family.Neural Computation, 15(5), 1013–1033. 53