Pith. sign in

REVIEW 1 major objections 3 minor 49 references

Fisher anisotropy can transfer Gaussian-width complexity between the Fisher and inverse-Fisher geometries, but it cannot make both widths smaller than the Euclidean Gaussian width.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 11:08 UTC pith:A42GSNVX

load-bearing objection A clean and new width product inequality with a careful learning lower bound; the recovery side is plausible but leans on an under-derived citation. the 1 major comments →

arxiv 2607.20578 v1 pith:A42GSNVX submitted 2026-07-22 cs.LG math.STstat.MLstat.TH

Fisher Widths: Local Learning Geometry and Anisotropic Recovery

classification cs.LG math.STstat.MLstat.TH MSC 60D0562B1094A1252A20
keywords Fisher widthGaussian widthFisher informationstatistical dimensionsparse recoveryweighted l1 minimizationdescent coneinformation geometry
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper pairs two complementary geometries carried by the Fisher information matrix of a statistical model: the primal Fisher width w_G(T)=w(G^{1/2}T) and the inverse-Fisher width w_{G^{-1}}(T)=w(G^{-1/2}T), each a Gaussian width of a locally deformed parameter set. It proves a sharp universal relation between them: on every compact coordinate set T and every positive-definite G, w_G(T) w_{G^{-1}}(T) ≥ w(T)^2, so a metric that inflates one width can shrink the other only up to that barrier. On the learning side, it shows that for Fisher-regular losses the uniform empirical-risk fluctuation scale w_G(H_r)/√n is attained on sufficiently small Fisher balls, making the standard upper bound tight rather than loose. On the recovery side, it shows that sparse recovery from Gaussian measurements with covariance G^{-1} is governed by a weighted-ℓ1 descent cone whose statistical dimension has support-dependent bounds, so recovery thresholds depend on where the active coordinates sit in the Fisher spectrum and not merely on sparsity. The net effect is a sharper picture of when local geometry helps and when it hurts: Fisher anisotropy can transfer complexity between learning and recovery, but it cannot erase it.

Core claim

The central claim is a sharp product inequality between the two widths induced by the Fisher metric and its inverse: for every nonempty compact parameter set T and every positive-definite Fisher information matrix G, w_G(T) w_{G^{-1}}(T) ≥ w(T)^2. Equality holds for isotropic G, where both widths are scaled Euclidean widths, and for antipodal two-point sets aligned with an eigenvector of G. The proof uses log-convexity of α ↦ w_{G^α}(T) for commuting powers of the metric, established through a Sudakov–Fernique comparison against a weighted sum of the two widths and a scalar arithmetic–geometric mean step. The same machinery yields a noncommutative geometric-mean bound for arbitrary positive-

What carries the argument

The key objects are the two width functionals obtained by deforming a common compact set T by G^{1/2} and G^{-1/2}. The product inequality follows from log-convexity of the map α ↦ w_{G^α}(T) along commuting powers of G, using Sudakov–Fernique comparison and the arithmetic–geometric mean inequality. On the recovery side, the load-bearing identity rewrites G^{-1/2}D(||·||_1,x*) as the descent cone of the weighted ℓ1 norm f_G(x)=||G^{1/2}x||_1, whose statistical dimension is then controlled by the standard distance-to-subdifferential functional U_G(S) and by U_G(S) minus an additive correction involving the Fisher mass on the support.

Load-bearing premise

The recovery-side lower bound rests on an unproved descent-cone error bound invoked in the proof of the two-sided statistical-dimension estimate; if that bound is not generally valid, the lower half of the recovery theorem and the support-ordering consequences lose their proof, while the width product inequality and learning-side result would still stand.

What would settle it

For G = diag(1,10,10), support S={1,2}, and x* = (1,1,0) in d=3, compute δ(G^{-1/2}D(||·||_1,x*)) by Monte Carlo projection onto the weighted descent cone and compare it with the interval [U_G(S) − 2√(21/11), U_G(S)]. A value outside that interval would refute Theorem 4.6 directly; the paper's own conjecture asks for a uniform lower bound c U_G(S), which can be tested by shrinking γ on the active support and checking whether δ/U_G(S) collapses.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For Fisher-regular losses, empirical-risk fluctuation on a small Fisher ball is, up to universal constants, exactly w_G(H_r)/√n, so the Fisher-geometric upper bound is achieved.
  • In inverse-Fisher sparse recovery, the measurement requirement for basis pursuit is governed by U_G(S), which depends on which coordinates are active; swapping active coordinates toward larger Fisher curvature raises the threshold.
  • Compensating for anisotropy by inverse-square-root weighting or finite-sample column normalization removes profile dependence but can raise the required measurements in profiles where the unweighted decoder already benefits from the geometry.
  • The product inequality implies that any preconditioning of a parameter set by G and by G^{-1} cannot make both widths smaller than the Euclidean Gaussian width; complexity can be redistributed but not destroyed.
  • Nested supports are ordered monotonically: adding active coordinates cannot decrease the recovery upper bound, and equal-cardinality supports can have very different recovery thresholds.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable consequence of the support ordering is that intentionally placing a sparse signal on low-Fisher-curvature coordinates should lower basis-pursuit sample complexity; the paper's own low-support profile shows a several-fold reduction, which could be probed in larger randomized designs.
  • The product inequality suggests a conservation principle for local statistical geometry: the same Fisher matrix governs both parameter sensitivity and estimation noise, so any metric that flattens one geometry must steepen the other; this may inform preconditioning and natural-gradient-style optimization, though the paper does not develop that direction.
  • The two-sided recovery estimate is non-vacuous only when the active support carries a non-negligible fraction of total Fisher mass; proving the paper's Conjecture 4.8 would extend it to extreme sparsity, and that conjecture is a natural target for a direct cone-geometric argument.
  • If the descent-cone error bound used in the recovery proof turns out not to hold in its stated form, the product inequality and the learning-side lower bound would remain intact; the recovery lower half would reduce to a known upper functional plus a conjectured comparison.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 3 minor

Summary. The paper introduces a primal-inverse pair of Fisher widths, w_G(T) = w(G^{1/2}T) and w_{G^{-1}}(T) = w(G^{-1/2}T), and studies their roles in learning and in anisotropic Gaussian recovery. Section 3 proves a nonasymptotic local lower bound showing that the scale w_G(H_r)/√n is attained on small Fisher balls for Fisher-regular losses. Section 4 analyzes inverse-Fisher measurements, reduces unweighted ℓ1 recovery to a weighted-ℓ1 descent cone, and claims a two-sided statistical-dimension estimate in terms of the functional U_G(S). Section 5 establishes the sharp product inequality w_G(T)w_{G^{-1}}(T) ≥ w(T)^2, with a noncommuting geometric-mean extension. The paper closes with controlled numerical experiments on recovery transitions and width redistribution.

Significance. If the claims hold, the most valuable contribution is Theorem 5.3, an elegant and sharp relation showing that Fisher anisotropy cannot reduce both primal and inverse widths below the Euclidean width. Theorem 3.4 is also a solid finite-sample lower bound with explicit constants. The recovery analysis in Section 4 is potentially significant because it predicts support-location effects in sparse recovery under Fisher-induced covariance. The numerical experiments are reproducible in spirit and match the theoretical upper functional closely. However, the two-sided recovery claim rests on an external estimate that is neither derived nor precisely located, and the lower-bound half of Theorem 4.6 is therefore not currently established.

major comments (1)
  1. [§4.3, Eq. (7)] The lower bound in Theorem 4.6 depends on Eq. (7), cited as the 'descent-cone error estimate of Amelunxen et al. [2014]'. The standard ALMT descent-cone upper bound is δ(D(f,x)) ≤ inf_{τ≥0} E dist²(g, τ∂f(x)); the displayed inequality with the additional subtractive term 2 sup_{z∈∂f(x)}‖z‖₂/f(x/‖x‖₂) is not a standard consequence and is not derived. The reader's sanity check with f=‖x‖ does not refute it (the RHS is d−2 while δ is d−1/2), but the estimate remains unsupported. Since the 'two-sided' statistical-dimension estimate is a central claim of the paper, Equation (7) must either be proved in the manuscript or supplied with an exact, verifiable reference, including the hypotheses under which it holds. Without this, Theorem 4.6's lower bound is not established.
minor comments (3)
  1. [§1.1] The introductory subsection contains duplicated paragraphs (the 'Fisher information matrix' paragraph and the 'Throughout this paper' paragraph appear twice).
  2. [Theorem 4.6] The displayed lower bound appears to be missing a division sign. The proof computes the error term as 2√(Tr(G)/Σ_{i∈S}γ_i), but the theorem statement as typeset reads like 2√(Tr(G)·Σ_{i∈S}γ_i), which is dimensionally inconsistent. Please correct the display.
  3. [Eq. (7)] Even if Eq. (7) is a known result, the manuscript should give a precise equation/location in Amelunxen et al. [2014]; the current citation is too vague for a load-bearing estimate.

Circularity Check

0 steps flagged

No significant circularity: central derivations are self-contained and rest on standard external results; only peripheral self-citations appear.

full rationale

The paper's main claims are derived from standard external ingredients rather than from the quantities they predict. Theorem 3.4 uses Paley–Zygmund, isotropy of the whitened score, and the exact identity G^{1/2}H_r = rB_2; it does not assume the desired lower bound. The Section 4 recovery bounds use Gordon's escape theorem and the ALMT conic framework. In particular, U_G(S) is explicitly acknowledged as the standard weighted-l1 expression, and the additive lower bound in Theorem 4.6 is obtained by optimizing the cited ALMT error term over vectors with the same support and sign pattern—a legitimate mathematical optimization, not a fit or a renamed prediction. Theorem 5.3 follows from log-convexity of Gaussian width along commuting powers of G plus the arithmetic–geometric mean inequality; it does not presuppose the product inequality. Self-citations to Ky [2026] are peripheral: they introduce the Fisher-width terminology, quote a perturbation-stability lemma, and provide a motivating example. The flagged concern about Eq. (7) is a potential correctness issue with a cited external estimate, not a circular reduction: the paper invokes Amelunxen et al. as a theorem rather than deriving its conclusion from its own assumptions. No load-bearing step reduces to an input by construction.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

No parameters are fitted to data; the constants (c_k, C) are universal; the Fisher profiles in the experiments are chosen for illustration and the recovery thresholds are computed from the theory, not fitted to the curves. No new physical or probabilistic entities are introduced; the primal–inverse width pair is a definition rather than an explanatory postulate.

axioms (6)
  • standard math Sudakov–Fernique comparison theorem
    Used in Lemma 2.2, Theorem 5.1, and Theorem 5.9 to compare Gaussian processes with ordered increments.
  • standard math Gordon's escape theorem / conic kinematic formula
    Used in Proposition 4.1 and Corollary 4.2 for the sufficient recovery scale.
  • ad hoc to paper ALMT descent-cone error estimate, Eq. (7)
    Invoked in Theorem 4.6 for the lower bound; not proved in the text and appears questionable; a Euclidean-norm example suggests it is not a valid general ALMT statement.
  • domain assumption Fisher-regularity conditions FR1–FR3 (centered gradient with covariance G, L2 Hessian bound, fourth-moment bound)
    Definition 3.1; these are the modeling assumptions under which Theorem 3.4 is proved.
  • domain assumption Correct specification (gradient covariance equals Fisher information)
    Used in FR1 and in Appendix A for logistic/GLM/Gaussian regression.
  • domain assumption Fisher matrix G(θ0) ≻ 0 at reference point
    Assumed throughout (Section 2.1); needed for the square roots and widths to be finite.

pith-pipeline@v1.3.0-alltime-deepseek · 20019 in / 22723 out tokens · 210333 ms · 2026-08-01T11:08:06.514567+00:00 · methodology

0 comments
read the original abstract

We study Gaussian-width complexity on statistical manifolds through a pair of functionals: the primal Fisher width $w_G(T) = w(G^{1/2}T)$, induced by the Fisher metric, and the inverse-Fisher width $w_{G^{-1}}(T) = w(G^{-1/2}T)$, induced by the inverse Fisher metric. The two widths play complementary statistical roles. On the learning side, the Fisher width measures the size of local parameter fluctuations in the geometry induced by the Fisher information. For Fisher-regular losses, we prove that the scale \(w_G(H_r)/\sqrt n\) is attained on sufficiently small Fisher balls. On the recovery side, the inverse-Fisher width captures the effect of anisotropic Gaussian measurements whose covariance is determined by the inverse Fisher information. For sparse recovery, the resulting geometry depends not only on sparsity but also on the position of the active coordinates in the Fisher spectrum. We obtain a two-sided estimate for the corresponding statistical dimension, together with support-sensitive recovery estimates and a natural ordering of supports with different curvature profiles. Finally, we establish a sharp relation between the primal and inverse-Fisher widths. On any common compact coordinate set $T$, they satisfy \[ w_G(T)w_{G^{-1}}(T)\geq w(T)^2. \] Thus, Fisher anisotropy may transfer complexity from one geometry to the other, but cannot reduce both widths relative to the Euclidean scale.

Figures

Figures reproduced from arXiv: 2607.20578 by Vu Khac Ky.

Figure 1
Figure 1. Figure 1: Empirical recovery curves for ordinary basis pursuit under inverse-Fisher [PITH_FULL_IMAGE:figures/full_fig_p028_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of four decoders under the five Fisher profiles, with [PITH_FULL_IMAGE:figures/full_fig_p031_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Primal–inverse width redistribution for Gλ = G0 + λI64 on the subspace ball T = UB10 2 . Widths are estimated using 105 Monte Carlo samples. Along this path, the primal width decreases and the inverse-Fisher width increases, while the normalized product ρλ approaches 1 from above. The figure is an illustration of the redistribution mechanism, not a numerical proof of the product inequality. Acknowledgments… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references

  1. [1]

    2026 , eprint =

    Vu Khac Ky , title =. 2026 , eprint =

  2. [2]

    Amin and Xu, Weiyu and Avestimehr, A

    Khajehnejad, M. Amin and Xu, Weiyu and Avestimehr, A. Salman and Hassibi, Babak , title =. IEEE Transactions on Signal Processing , year =

  3. [3]

    Compressed Sensing of Data with a Known Distribution , journal =

    D. Compressed Sensing of Data with a Known Distribution , journal =. 2018 , volume =

  4. [4]

    , title =

    van der Vaart, Aad W. , title =

  5. [5]

    Bhatia, Rajendra , title =

  6. [6]

    Ledoux, Michel and Talagrand, Michel , title =

  7. [7]

    Concentration Inequalities: A Nonasymptotic Theory of Independence , publisher =

    Boucheron, St. Concentration Inequalities: A Nonasymptotic Theory of Independence , publisher =

  8. [8]

    Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information , journal =

    Cand. Robust uncertainty principles: Exact signal reconstruction from highly incomplete frequency information , journal =

  9. [9]

    , title =

    Donoho, David L. , title =. IEEE Transactions on Information Theory , volume =

  10. [10]

    Foucart, Simon and Rauhut, Holger , title =

  11. [11]

    , title =

    Wainwright, Martin J. , title =. IEEE Transactions on Information Theory , volume =

  12. [12]

    Discrete & Computational Geometry , volume =

    Plan, Yaniv and Vershynin, Roman , title =. Discrete & Computational Geometry , volume =

  13. [13]

    Journal of the Royal Statistical Society: Series B (Methodological) , volume =

    Tibshirani, Robert , title =. Journal of the Royal Statistical Society: Series B (Methodological) , volume =

  14. [14]

    and Ravikumar, Pradeep and Wainwright, Martin J

    Negahban, Sahand N. and Ravikumar, Pradeep and Wainwright, Martin J. and Yu, Bin , title =. Statistical Science , volume =

  15. [15]

    , title =

    Tsybakov, Alexandre B. , title =

  16. [16]

    The Annals of Statistics , volume =

    Spokoiny, Vladimir , title =. The Annals of Statistics , volume =

  17. [17]

    Statistical Decision Rules and Optimal Inference , series =

  18. [18]

    International Conference on Learning Representations , year =

    Pascanu, Razvan and Bengio, Yoshua , title =. International Conference on Learning Representations , year =

  19. [19]

    Entropy , volume =

    Nielsen, Frank , title =. Entropy , volume =

  20. [20]

    Linear Algebra and its Applications , volume =

    Kueng, Richard and Gross, David , title =. Linear Algebra and its Applications , volume =

  21. [21]

    IEEE Transactions on Information Theory , volume =

    Rudelson, Mark and Zhou, Shuheng , title =. IEEE Transactions on Information Theory , volume =

  22. [22]

    Proceedings of the 32nd International Conference on Machine Learning , series =

    Martens, James and Grosse, Roger , title =. Proceedings of the 32nd International Conference on Machine Learning , series =. 2015 , publisher =

  23. [23]

    Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics , series =

    Karakida, Ryo and Akaho, Shotaro and Amari, Shun-ichi , title =. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics , series =. 2019 , publisher =

  24. [24]

    Advances in Neural Information Processing Systems 32 , year =

    Kunstner, Frederik and Balles, Lukas and Hennig, Philipp , title =. Advances in Neural Information Processing Systems 32 , year =

  25. [25]

    Talagrand, Michel , title =

  26. [26]

    and Tanner, Jared , title =

    Donoho, David L. and Tanner, Jared , title =. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences , volume =

  27. [27]

    Rao, C. R. , title =. Bulletin of the Calcutta Mathematical Society , volume =

  28. [28]

    Amari, Shun-ichi and Nagaoka, Hiroshi , title =

  29. [29]

    Mohri, Mehryar and Rostamizadeh, Afshin and Talwalkar, Ameet , title =

  30. [30]

    Advances in Neural Information Processing Systems 30 , year =

    Neyshabur, Behnam and Bhojanapalli, Srinadh and McAllester, David and Srebro, Nathan , title =. Advances in Neural Information Processing Systems 30 , year =

  31. [31]

    and Arsenin, Vasiliy Y

    Tikhonov, Andrey N. and Arsenin, Vasiliy Y. , title =

  32. [32]

    and Foster, Dylan J

    Bartlett, Peter L. and Foster, Dylan J. and Telgarsky, Matus J. , title =. Advances in Neural Information Processing Systems 30 , year =

  33. [33]

    and Solla, Sara A

    LeCun, Yann and Denker, John S. and Solla, Sara A. , title =. Advances in Neural Information Processing Systems 2 , editor =

  34. [34]

    , title =

    Hassibi, Babak and Stork, David G. , title =. Advances in Neural Information Processing Systems 5 , pages =

  35. [35]

    Geometric Aspects of Functional Analysis , series =

    Gordon, Yehoram , title =. Geometric Aspects of Functional Analysis , series =

  36. [36]

    and Tropp, Joel A

    Amelunxen, Dennis and Lotz, Martin and McCoy, Michael B. and Tropp, Joel A. , title =. Information and Inference: A Journal of the IMA , volume =

  37. [37]

    Neural Computation , volume =

    Amari, Shun-ichi , title =. Neural Computation , volume =

  38. [38]

    and Mendelson, Shahar , title =

    Bartlett, Peter L. and Mendelson, Shahar , title =. Journal of Machine Learning Research , volume =

  39. [39]

    Near-Optimal Signal Recovery From Random Projections: Universal Encoding Strategies? , journal =

    Cand. Near-Optimal Signal Recovery From Random Projections: Universal Encoding Strategies? , journal =

  40. [40]

    , title =

    McAllester, David A. , title =. Proceedings of the Eleventh Annual Conference on Computational Learning Theory (COLT) , pages =

  41. [41]

    Catoni, Olivier , title =

  42. [42]

    and Willsky, Alan S

    Chandrasekaran, Venkat and Recht, Benjamin and Parrilo, Pablo A. and Willsky, Alan S. , title =. Foundations of Computational Mathematics , volume =

  43. [43]

    and Casella, George , title =

    Lehmann, Erich L. and Casella, George , title =

  44. [44]

    Vershynin, Roman , title =

  45. [45]

    , title =

    Wainwright, Martin J. , title =

  46. [46]

    Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss , booktitle =

    Chizat, L\'. Implicit Bias of Gradient Descent for Wide Two-layer Neural Networks Trained with the Logistic Loss , booktitle =

  47. [47]

    Flat Minima , journal =

    Hochreiter, Sepp and Schmidhuber, J\". Flat Minima , journal =

  48. [48]

    Surveys in Combinatorics , editor =

    McDiarmid, Colin , title =. Surveys in Combinatorics , editor =

  49. [49]

    SIAM Journal on Mathematical Analysis , volume =

    Jordan, Richard and Kinderlehrer, David and Otto, Felix , title =. SIAM Journal on Mathematical Analysis , volume =