REVIEW 2 major objections 5 minor 28 references
Parametric models without densities, scores, or Fisher information can still carry Riemannian metrics once laws are tempered distributions and instruments extract Godambe information.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 06:01 UTC pith:6I4LOVEG
load-bearing objection Solid, usable extension of information geometry via Godambe instruments on tempered distributions; examples carry the weight and the math is clean under stated premises. the 2 major comments →
Weak Information Geometry: Riemannian Structures from Distributional Inference Functions and Stein Discrepancies
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Once a parametric family is represented by tempered distributions T_θ, any instrument (weak regular inference function or weak Stein representation) whose sensitivity S is full-rank and whose variability V is positive definite induces the Godambe information G(θ) = S(θ)ᵀ V(θ)⁻¹ S(θ) as a smooth Riemannian metric on the parameter space. The Fisher–Rao manifold is the special case in which the score is itself an admissible instrument; whenever Fisher information exists, every Godambe metric is dominated by it in the Loewner order.
What carries the argument
The Godambe information G = Sᵀ V⁻¹ S built from the weak sensitivity and variability of an instrument acting on tempered distributions; it is the metric tensor that turns the parameter space into a Godambe–Riemannian manifold and recovers Fisher information when the instrument is the score.
Load-bearing premise
The map from parameters to the distributional representation must be weakly differentiable: for every admissible probe the expectation must be continuously differentiable and derivatives must pass inside the pairing.
What would settle it
Exhibit a concrete parametric family of tempered distributions together with an instrument whose sensitivity and variability are smooth, full-rank and positive-definite, yet the resulting G fails to be a Riemannian metric, or show that the claimed Loewner domination G ⪯ I fails when Fisher information exists and the Bartlett interchange holds.
If this is right
- Undominated families such as the Cantor location model, models with parameter-dependent support, and mixtures whose score is biased all become Riemannian manifolds once a suitable instrument is chosen.
- Quadratic Stein discrepancies built from finite collections of identities induce exactly the same local geometry as the corresponding Godambe metrics.
- Weak inferential separation (nonformation) is equivalent to block-diagonality of the Godambe metric with respect to the interest–nuisance splitting.
- In dynamical models such as the lattice stochastic heat equation driven by alpha-stable noise, the Godambe geometry stabilises at a rate controlled by the spectral gap of the discrete Laplacian.
- Because there is no canonical instrument, a single model carries a family of metrics that can be selected according to inferential, diagnostic, geometric or computational purpose.
Where Pith is reading between the lines
- The same instrument construction could be applied to other singular or heavy-tailed SPDE models to obtain closed-form geometric stability rates without second-moment assumptions.
- Hierarchy of RKHS Stein geometries suggests a practical route to adaptive metric selection: start with low-frequency or moment instruments and enrich only when diagnostics show near-degeneracy.
- Automatic block-diagonality for odd/even probes in every symmetric location-scale family (including Cauchy and Cantor) supplies an immediate, likelihood-free method for exact location–scale separation in robust estimation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that parametric models represented by tempered distributions T_θ ∈ S'(R^k) can be equipped with Riemannian metrics far beyond the classical Fisher–Rao setting. An external instrument (weak regular inference function or weak Stein representation) with full-rank sensitivity S and positive-definite variability V induces the Godambe information G = S^T V^{-1} S as a smooth metric on Θ (Theorem 1.1 / Proposition 3.1). Fisher–Rao is recovered when the score is admissible, and G ⪯ I in Loewner order whenever Fisher exists (Proposition 3.2). Five examples outside Fisher–Rao are treated in closed form: uniform scale, shifted exponential, undominated Cantor location, stratified mixture with biased score, and a lattice α-stable heat equation whose Godambe metric stabilises at the spectral-gap rate. Quadratic Stein discrepancies recover the same local geometry; non-canonicity of the instrument and block-diagonality as geometric nonformation are discussed.
Significance. If the construction is accepted, the class of models that can be treated by differential-geometric methods expands substantially: undominated families, parameter-dependent support, moment-free laws, and dominated models with biased scores all become Riemannian. The closed-form metrics (uniform G=3/θ², Cantor G≡8, SPDE spectral-gap rate αθλ_k) and the explicit Loewner comparison are concrete, usable contributions. The paper is honest about non-canonicity (Section 7, Remark 7.1) and about the instrument-relative character of the geometry. Dependence on the author’s 2026 companion series for the full apparatus of weak inference functions and genericity is a presentational limitation rather than a correctness flaw in the derivations as written.
major comments (2)
- Assumption 2.3 (weak differentiability of θ ↦ ⟨T_θ, g⟩ with derivatives passing inside the pairing) is the analytic hinge for the weak Bartlett identity (1), for S(θ)=⟨∂_θ T_θ, ψ⟩, and thus for smoothness of G in Theorem 1.1 / Proposition 3.1. It is stated as an assumption and verified only by direct computation in the five examples; no general criterion is given for when a tempered-distribution model map satisfies it. For the central claim to be usable beyond the exhibited cases, the paper should either supply verifiable sufficient conditions (e.g., dominated differentiability of characteristic functions, or continuity of the map in the weak-* topology of S') or state more explicitly that the theorem is conditional on this hypothesis and that existence is left to the companion [25].
- Proposition 1.2 and the examples establish existence of instruments case-by-case, but the paper repeatedly speaks of “the class of Godambe–Riemannian models” as if it were characterised. Without a general existence theorem (or a clear statement that none is claimed), the scope of the extension remains an open-ended collection of examples. A short subsection clarifying what is proved versus what is exhibited would prevent over-reading of Theorem 1.1.
minor comments (5)
- Heavy self-citation to the 2026 companion series ([21]–[25]) makes the paper hard to read in isolation. A short self-contained appendix restating Definition 2.2 and the admissible classes of Lemma 2.1 would help.
- Notation: α and β are used both for the stability index / eigenmode scales (Section 5) and for interest/nuisance parameters (Section 8). A local warning is present but easy to miss; consider distinct symbols.
- Remark 4.1 (Student-t / Cauchy) and Remark 4.9 (Cauchy mixture) are useful but sit outside the main example sections; a pointer in the introduction would improve navigation.
- Typographical: “F amily of metrics” in the Contents (Section 7 heading) has a stray space; “P´ olya” and similar accented names appear inconsistently.
- Section 5.6 claims numerical verification of the closed-form CF and of Proposition 5.1 against Monte Carlo, but no figure or table is supplied. A brief plot of |G_k(t)−G_k(∞)| versus the predicted rate would strengthen the claim.
Circularity Check
Minor self-citation to same-author 2026 companions for the distributional apparatus and genericity; central Godambe-metric construction, Loewner comparison, and explicit examples remain independently proved here.
specific steps
-
self citation load bearing
[Section 2 (Preliminaries) and Introduction (references to companions for instruments and genericity)]
"The systematic development of weak regular inference functions is carried out in the companion paper [23]; here we record only what is needed. ... The companion paper [25] shows that such failures are exceptional in a precise sense: within natural families of instruments, those violating the hypotheses form a negligible set..."
The foundational notions of weak regular inference functions, observation operators, and the claim that non-degeneracy of (S,V) is generic are justified by citation to same-author arXiv companions rather than derived or machine-checked here. While the paper supplies self-contained definitions and verifies the hypotheses by hand in its examples, the broader apparatus that makes the instrument class well-defined for arbitrary tempered distributions rests on the author-controlled stack; this is a mild load-bearing self-citation but does not force the metric theorems by construction.
full rationale
The paper's core derivation (Proposition 3.1 that G = SᵀV⁻¹S is a smooth Riemannian metric under full-rank S and positive-definite V; Proposition 3.2 Loewner comparison recovering Fisher–Rao for the score; Theorem 1.1) is classical linear algebra plus the weak Bartlett identity under the stated Assumption 2.3, all proved in-place from the definitions given in Section 2. Existence of instruments outside Fisher–Rao is established by five fully explicit closed-form calculations (uniform, shifted exponential, Cantor, stratified mixture, lattice α-stable SPDE) that do not rely on external results. Self-citations to the author's companion series [21–25] supply background on weak moments, observation operators, and transversality/genericity, but these are not load-bearing for the metric claims themselves: the paper recalls the needed definitions, verifies non-degeneracy by direct computation in every example, and never treats a self-cited uniqueness or existence theorem as forcing the present results. No self-definitional loop, no fitted-input-as-prediction, no smuggled ansatz, and no renaming of a known empirical pattern appear. The dependence is therefore ordinary presentational scaffolding rather than circularity of the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (1)
- Instrument tuning (e.g. frequency c or t in sinusoidal/transform probes)
axioms (6)
- domain assumption Laws P_θ are represented by tempered distributions T_θ ∈ S'(R^k) via ⟨T_θ,g⟩=E_{P_θ}[g] for g in Schwartz space (and extensions of Lemma 2.1).
- domain assumption Assumption 2.3: weak C¹ differentiability of θ ↦ ⟨T_θ,g⟩ for admissible probes, with derivatives realized as pairings against ∂_{θ_j} T_θ.
- domain assumption Existence of at least one weak regular inference function (or Stein-generated instrument) with full column-rank S and positive-definite V (hypotheses of Prop. 3.1 / Thm 1.1).
- standard math Positive-definiteness of a quadratic form aᵀGa when V≻0 and Sa≠0 (standard linear algebra).
- domain assumption Bartlett-type interchange E[ψ u_θᵀ]=S when densities and scores exist (condition (2) in Prop. 3.2).
- ad hoc to paper Fréchet differentiability of the embedded model map into H* for RKHS Stein metrics (Assumption 6.2).
invented entities (2)
-
Instrument (positive Schwartz kernel / weak regular inference function / weak Stein representation as measurement device external to the law)
independent evidence
-
Godambe–Riemannian manifold (Θ,G)
no independent evidence
read the original abstract
The class of parametric statistical models that can be treated as Riemannian manifolds is considerably larger than the classical Fisher-Rao setting allows, once one works in the space of tempered distributions. A law is represented by a tempered distribution T in S'(R^k), while an instrument - a positive Schwartz kernel, a weak regular inference function, or a weak Stein representation - extracts information from the law without being part of it. Any instrument with full-rank sensitivity and positive-definite variability induces the Godambe information G = S^T V^{-1} S, a Riemannian metric on the parameter space; the Fisher-Rao manifold is recovered exactly when the score is an admissible instrument, and every Godambe metric is dominated by the Fisher metric in the Loewner order whenever the latter exists. Four examples lie outside the Fisher-Rao class for four different reasons: a location model built on the Cantor distribution (an undominated family - no likelihood, no score, and no Fisher information exist at all), the uniform scale model (parameter-dependent support), the shifted exponential model (transform-based inference), and a stratified finite mixture (a provably biased score in a dominated model); a lattice stochastic heat equation driven by alpha-stable noise provides a fifth, dynamical example, whose closed-form weak Godambe information stabilises at a rate governed by the spectral gap of the discrete Laplacian. Quadratic Stein discrepancies induce the same local geometry, and reproducing-kernel constructions generate a hierarchy of geometries. Because there is no canonical instrument, the model carries a family of Godambe metrics; we discuss the inferential, diagnostic, geometric, and computational roles of its members, and show that weak inferential separation (nonformation) appears geometrically as block-diagonality of the Godambe metric.
Reference graph
Works this paper leans on
-
[1]
(1985).Differential-Geometrical Methods in Statistics
Amari, S. (1985).Differential-Geometrical Methods in Statistics. Lec- ture Notes in Statistics 28, Springer
1985
-
[2]
Amari, S. (1998). Natural gradient works efficiently in learning.Neural Computation, 10(2), 251–276
1998
-
[3]
and Kawanabe, M
Amari, S. and Kawanabe, M. (1997). Information geometry of esti- mating functions in semi-parametric statistical models.Bernoulli, 3(1), 29–54
1997
-
[4]
and Nagaoka, H
Amari, S. and Nagaoka, H. (2000).Methods of Information Geometry. Translations of Mathematical Monographs 191, AMS/Oxford
2000
-
[5]
and Schwachh¨ ofer, L
Ay, N., Jost, J., Lˆ e, H.V. and Schwachh¨ ofer, L. (2017).Informa- tion Geometry. Ergebnisse der Mathematik und ihrer Grenzgebiete 64, Springer
2017
-
[6]
(1978).Information and Exponential Families in Statistical Theory
Barndorff-Nielsen, O.E. (1978).Information and Exponential Families in Statistical Theory. Wiley, Chichester
1978
-
[7]
and Reid, N
Barndorff-Nielsen, O.E., Cox, D.R. and Reid, N. (1986). The role of dif- ferential geometry in statistical theory.International Statistical Review, 54(1), 83–96
1986
-
[8]
and Mackey, L
Barp, A., Briol, F.-X., Duncan, A.B., Girolami, M. and Mackey, L. (2019). Minimum Stein discrepancy estimators.Advances in Neural In- formation Processing Systems, 32, 12964–12976
2019
-
[9]
and Florens, J.-P
Carrasco, M. and Florens, J.-P. (2000). Generalization of GMM to a continuum of moment conditions.Econometric Theory, 16(6), 797–834
2000
-
[10]
(1982).Statistical Decision Rules and Optimal Inference
ˇCencov, N.N. (1982).Statistical Decision Rules and Optimal Inference. Translations of Mathematical Monographs 53, American Mathematical Society, Providence. (Russian original, 1972.)
1982
-
[11]
and Gretton, A
Chwialkowski, K., Strathmann, H. and Gretton, A. (2016). A kernel test of goodness of fit.Proceedings of the 33rd International Conference on Machine Learning, 2606–2615. 51
2016
-
[12]
and Salmon, M
Critchley, F., Marriott, P. and Salmon, M. (1993). Preferred point ge- ometry and statistical manifolds.The Annals of Statistics, 21(3), 1197– 1224
1993
-
[13]
(2003).Fractal Geometry: Mathematical Foundations and Applications
Falconer, K. (2003).Fractal Geometry: Mathematical Foundations and Applications. 2nd edition, Wiley, Chichester
2003
-
[14]
and McDunnough, P
Feuerverger, A. and McDunnough, P. (1981). On the efficiency of empir- ical characteristic function procedures.Journal of the Royal Statistical Society, Series B, 43(1), 20–27
1981
-
[15]
Godambe, V.P. (1960). An optimum property of regular maximum like- lihood estimation.The Annals of Mathematical Statistics, 31(4), 1208– 1211
1960
-
[16]
and Mackey, L
Gorham, J. and Mackey, L. (2017). Measuring sample quality with Stein’s method.The Annals of Statistics, 45(3), 1069–1109
2017
-
[17]
Heathcote, C.R. (1977). The integrated squared error estimation of pa- rameters.Biometrika, 64(2), 255–264
1977
-
[18]
and Labouriau, R
Jørgensen, B. and Labouriau, R. (2012).Exponential Families and Theoretical Inference. Monografias de Matem´ atica 52, Instituto de Matem´ atica Pura e Aplicada (IMPA), Rio de Janeiro
2012
-
[19]
(1996).Estimating Functions and Semiparametric Mod- els
Labouriau, R. (1996).Estimating Functions and Semiparametric Mod- els. Ph.D. thesis, University of Aarhus
1996
-
[20]
Labouriau, R. (2023). On the bias of the score function of finite mixture models.Communications in Statistics – Theory and Methods, 52(13). doi:10.1080/03610926.2021.1995429. (arXiv:2002.03307.)
-
[21]
R. Labouriau (2026A).Distributional Statistical Models: Weak Mo- ments, Cumulants, and a Central Limit Theorem, arXiv:2604.20634 [math.PR]
-
[22]
R. Labouriau (2026B)Weak Moment Methods for Statistical Inference: with an Application to Robust Estimation, arXiv:2604.23619 [stat.ME]
-
[23]
R. Labouriau (2026C).Inference Functionals and Observation Opera- tors for DistributionalStatistical Models. arXiv:2605.19189 [math.ST]
-
[24]
Labouriau (2026D).Weak Stein Discrepancies: Kernel-Regularised Goodness-of-Fit and Minimum Discrepancy Estimation for Heavy- Tailed Models, in preparation, 2026
R. Labouriau (2026D).Weak Stein Discrepancies: Kernel-Regularised Goodness-of-Fit and Minimum Discrepancy Estimation for Heavy- Tailed Models, in preparation, 2026. 52
2026
-
[25]
R. Labouriau (2026). Transversality and Geometric Regularisation in Distributional Statistical Models. arXiv:2605.04536 [math.ST]
Pith/arXiv arXiv 2026
-
[26]
and Jordan, M
Liu, Q., Lee, J. and Jordan, M. (2016). A kernelized Stein discrepancy for goodness-of-fit tests.Proceedings of the 33rd International Confer- ence on Machine Learning, 276–284
2016
-
[27]
(2009).Algebraic Geometry and Statistical Learning The- ory
Watanabe, S. (2009).Algebraic Geometry and Statistical Learning The- ory. Cambridge University Press
2009
-
[28]
and Amari, S
Watanabe, S. and Amari, S. (2003). Learning coefficients of layered models when the true distribution is not in the model family.Neural Computation, 15(5), 1013–1033. 53
2003
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.