Pith. sign in

REVIEW 1 major objections 4 minor 28 references

A data-based notion of quantiles on Hadamard spaces

T0 review · 1 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper defines data-based quantiles on Hadamard spaces—loss measured at the data point—and proves strong consistency, joint asymptotic normality, exact breakdown points, extreme-quantile divergence, and a bounded gradient.

desk verdict A solid extension of geometric quantiles to Hadamard spaces; the strongest advertised normality result rests on an unverified fix to someone else's theorem. read the letter →

arxiv 2506.12534 v1 pith:T4YDCHGP submitted 2025-06-14 stat.ME math.STstat.TH

classification stat.MEmath.STstat.TH MSC 62R3060F0562G35
keywords geometricquantilesdata-basedHadamardspacesCAT(0)asymptoticnormalitybreakdownpointextremesymmetricpositivedefinitematrices
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces a data-based notion of quantiles for Hadamard spaces—complete metric spaces of non-positive curvature, such as spaces of symmetric positive definite matrices used in diffusion tensor imaging. The loss is $\rho(x,p;\beta,\xi)=d(p,x)-\beta d(p,x)\cos(\angle_x(p,\xi))$, with the angle measured at the data point $x$ rather than at the candidate quantile $p$, and the paper argues this small change unlocks statistical theory that the older parameter-based version does not have. On locally compact Hadamard spaces the sample quantiles are strongly consistent; on Hadamard manifolds several quantiles are jointly asymptotically normal under weaker conditions; every sample quantile function has exact breakdown point $(\lceil(1-\beta)N/2\rceil)/N$; and as $\beta\to 1$ extreme quantiles recede to infinity along the direction $\xi$ while the loss gradient stays bounded between $1-\beta$ and $1+\beta$. If these results hold, quantile-based inference—testing whether two data sets share a distribution, measuring skewness and dispersion, flagging outliers—becomes available, with guarantees, on a broad class of non-Euclidean data.

What carries the argument

The object that carries the argument is the data-based loss function $\rho(x,p;\beta,\xi)=d(p,x)-\beta d(p,x)\cos(\angle_x(p,\xi))$, with the boundary direction $\xi\in\partial M$ serving as the quantile index. The key identity is the gradient formula $\nabla\rho(x,p;\beta,\xi)=-\log_p(x)/d(p,x)-d(\log_x)^\dagger_p(\beta\xi_x)$, which expresses the gradient entirely through the logarithm map and its adjoint differential, so its norm automatically lies in $[1-\beta,1+\beta]$. On locally symmetric Hadamard spaces the same derivative is computed explicitly with Jacobi fields, giving coefficients $(d/dt)g_i(t)|_{t=0}/g_i(1)$ that damp the direction term as curvature becomes more negative. This clean differentiability is what lets the loss enter the asymptotic normality machinery and what makes the cheap approximate gradient used in the descent algorithm accurate in experiments.

What would settle it

Check the published central limit theorem's own proof to see whether the orthogonal rotation factor must appear inside the covariance, as the paper's Remark 3.8 claims. Then simulate a known Hadamard-manifold model (for example, log-normal symmetric positive definite matrices), estimate the covariance of $\sqrt{N}(\hat{q}_N-q)$ for several $(\beta,\xi)$, and compare it with $\Lambda^{-1}\Sigma\Lambda^{-1}$; a systematic mismatch would refute the asymptotic normality claims, and a match only when the rotation factor is included would confirm the correction.

Watch

Extended reading notes

Core claim

The central claim is that the quantile loss should be evaluated from the data point rather than from the parameter. On a Hadamard manifold this loss is $\rho(x,p;\beta,\xi)=\|\log_x(p)\|-\langle\beta\xi_x,\log_x(p)\rangle$, and its gradient at $p\ne x$ is $\nabla\rho(x,p;\beta,\xi)=-\log_p(x)/d(p,x)-d(\log_x)^\dagger_p(\beta\xi_x)$, an expression that requires no derivative of the radial field $p\mapsto \xi_p$—the obstruction that made the parameter-based gradient hard to control. From this identity the paper derives the bounded gradient, an explicit Jacobi-field formula on locally symmetric spaces, and the differentiability needed for the asymptotic theorems. The joint asymptotic normality proof proceeds through an M-estimator theorem and, for the stronger version in dimension at least three, through a smeary central limit theorem whose stated covariance the paper corrects in a remark.

Load-bearing premise

The proof of the stronger joint asymptotic normality theorem depends on a claimed correction to a published central limit theorem: the paper inserts an extra rotation factor into the limiting covariance that the published statement omits. If that correction is wrong, Theorem 3.7 and the covariance formula it asserts are not established.

Editorial extensions

If this is right

  • On every locally compact Hadamard space, not only on manifolds, sample (β, ξ)-quantiles converge almost surely to the population quantile, so tree spaces and Euclidean buildings get quantile consistency.
  • Simultaneous quantiles at several values of (β, ξ) are jointly asymptotically normal with limiting covariance Λ⁻¹ΣΛ⁻¹, under conditions that do not require bounded support or second-order differentiability of the loss.
  • The exact breakdown point of any sample quantile function is (⌈(1−β)N/2⌉)/N, with the stated tie case in Theorem 4.2, and asymptotically it is (1−β)/2; for medians, non-unique selections make the classical ⌈N/2⌉/N formula incomplete.
  • As β→1, quantiles move away to infinity in direction ξ and asymptotically align with ξ from the data's perspective, replicating Euclidean extreme-quantile behavior and supporting outlier-detection uses.
  • The bounded gradient and its cheap approximation allow a simple adaptive descent algorithm to compute quantiles on hyperbolic space and SPD manifolds with small error, agreeing closely with true-gradient results in the paper's experiments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the corrected smeary central limit theorem is valid, it is a transferable tool: other manifold-valued M-estimators whose losses are differentiable on the data side could inherit joint asymptotic normality under the same weakened moment conditions.
  • The gradient's bounded norm suggests stochastic-gradient and online versions of quantile estimation on SPD manifolds may scale to high dimensions, where exact eigendecompositions of curvature operators are too expensive.
  • The extreme-quantile divergence along ξ suggests a directional outlier-detection score on hyperbolic embeddings of hierarchical data: directions whose extreme quantiles pull far from the center label branches of the structure.
  • Because the breakdown point of a median depends on the selection rule when the median set is non-unique, applied work should report how it chose among multiple medians; a norm-minimizing selection can tolerate one extra corrupted observation in the tie case.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This paper introduces a data-based notion of geometric quantiles on Hadamard spaces. The population quantile minimizes G_{β,ξ}(p)=E[d(p,X)-β d(p,X) cos(∠_X(p,ξ))], where the angle is measured at the data point X rather than at the candidate quantile p, in contrast to the parameter-based loss of Shin and Oh (2023). The paper proves strong consistency on locally compact Hadamard spaces (Theorem 3.2), joint asymptotic normality under two sets of conditions (Theorems 3.4 and 3.7), exact worst-case breakdown points for sample quantile functions (Theorem 4.2), divergence of extreme quantiles in the direction ξ (Theorem 5.1), and boundedness of the loss gradient with an explicit formula on locally symmetric spaces (Proposition 6.1 and Theorem 6.3). Simulations in hyperbolic space and real DTI experiments illustrate computation, shape adherence, permutation tests, and distributional characteristic measures.

Significance. If the results hold, the paper provides a substantially more tractable and theoretically cleaner notion of quantiles for non-Euclidean data than parameter-based quantiles. The bounded-gradient result (Proposition 6.1) and the explicit gradient formula on locally symmetric spaces (Theorem 6.3) are genuinely useful and support a simple descent algorithm. The breakdown-point theorem is sharp and carefully addresses selection effects that earlier median breakdown arguments overlooked. The paper is generally well organized, with detailed proofs in the appendices and publicly available code. However, the headline weaker-condition asymptotic normality result, Theorem 3.7, rests on an unverified correction to a published theorem; until that correction is independently established, the strongest advertised advantage is conditional.

major comments (1)
  1. [Section 3.2, Theorem 3.7 and Remark 3.8] The proof of Theorem 3.7 is routed entirely through Theorem 2.11 of Eltzner and Huckemann (2019), but the stated covariance in that theorem is asserted to contain a 'significant typo'. The only support is the sentence 'this is because the proof involves showing that the sequence of random vectors is equal to -(1/r)T^{-1}RG_n + o_p(1)', with no derivation and no independent verification. The algebraic elimination of D and D^T in the proof of Theorem 3.7 (Appendix A.3, near the displayed covariance after applying Theorem 2.11) uses the R factors critically: the covariance is transformed using R R^T = I to obtain Λ^{-1}ΣΛ^{-1}. If the published theorem is correct as printed, or if the correction has a different form, the conclusion of Theorem 3.7 does not follow. Because Theorem 3.7 is presented as the new result that weakens condition (II) for n ≥ 3, this is a load-bearing gap. Please provide a complete derivation or an independent verification of the corrected covariance, or state Theorem 3.7 as conditional on that correction.
minor comments (4)
  1. [Appendix A.2, proof of Lemma 3.1] The text 'the preimage of E_{p,r} := under the map' appears to be missing the definition of E_{p,r}; as written, the set is undefined.
  2. [Appendix A.5, proof of Theorem 5.1(c)] The displayed equation after 'by (39), binomial expansion, and (40)' begins with 'ρ(x, p; 1, ξ)⟩ ≤', which contains an unmatched angle bracket and should presumably read 'ρ(x, p; 1, ξ) ≤'.
  3. [Section 7.2, Tables 3 and 4] The real-data conclusions are based on single point estimates of dispersion, skewness, kurtosis, and spherical asymmetry, with no standard errors, replication, or comparison against the true-gradient solution on P3; the quantiles are computed with the approximate gradient (14) rather than the true gradient of Theorem 6.3.
  4. [Section 7.1.1] The permutation experiment uses 500 permutations on one pair of simulated data sets; a permutation confidence interval or repeated simulation would make the reported p-values more informative.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the data-based quantile theory is derived from its definition using external benchmarks; self-citations are background support, not load-bearing.

full rationale

The paper's central objects are defined independently of the results claimed about them: Definition 2.1/2.2 fixes the (beta,xi)-quantile as the minimizer of G_{beta,xi} with rho(x,p) = d(p,x) - beta d(p,x) cos(angle_x(p,xi)), and every theorem is then proved from this definition against external tools (Huber 1967; Bridson and Haefliger 1999; Eltzner and Huckemann 2019; Koltchinskii 1997; Girard and Stupfler 2017; Lopuhaa and Rousseeuw 1991; Konen and Paindaveine 2024). No parameter is fitted and then relabeled as a prediction: strong consistency (Theorem 3.2), joint asymptotic normality (Theorems 3.4 and 3.7), breakdown points (Theorem 4.2), extreme-quantile divergence (Theorem 5.1), and the gradient bounds (Proposition 6.1) are all derived, with the appendices supplying proofs that do not presuppose the conclusions. The heavy reliance on Shin and Oh (2023) is substantive prior work—geometric background, Lemma B.1, and proof templates—but the cited statements are not the target results, and the new data-based loss is not reduced to parameter-based quantiles by construction. The only substantive caveat is Remark 3.8: Theorem 3.7 imports a corrected covariance for Eltzner and Huckemann's smeary CLT, and the paper's supporting sketch is brief; however, that is an external correctness risk, not a circularity, because the corrected covariance is not assumed as the paper's conclusion. Accordingly, no circular step is exhibited.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper introduces no fitted constants and no ad hoc parameters; beta and xi are user-chosen quantile indices. The assumptions are standard finite-moment and local-compactness conditions, plus reliance on published theorems including a corrected smeary CLT. No new physical or mathematical entities are postulated.

assumptions (6)
  • domain assumption Assumption 2.7: For some beta* in [0,1), xi* in boundary, and p* in M, G_{beta*,xi*}(p*) is finite.
    This finite-moment type condition ensures the population loss is finite and continuous on all of M via Proposition 2.6.
  • standard math M is a complete CAT(0) Hadamard space.
    Used throughout to guarantee unique geodesics, boundary at infinity, and the validity of the Cartan-Hadamard theorem.
  • domain assumption Local compactness of M is assumed for strong consistency and compactness of quantile sets.
    Theorem 3.2 and Proposition 2.6(b) rely on local compactness plus a Hopf-Rinow type theorem for length spaces.
  • domain assumption For asymptotic normality, X has a bounded density in a neighborhood of each true quantile (condition (I) of Theorem 3.4).
    Needed to apply Huber's conditions and the smeary central limit theorem in Theorem 3.7.
  • standard math The corrected form of Theorem 2.11 of Eltzner and Huckemann (2019), stated in Remark 3.8, is correct.
    Theorem 3.7 derives its limiting covariance from this external theorem with the authors' proposed correction. If the correction is wrong, the proof of Theorem 3.7 fails.
  • domain assumption In Theorem 5.1, the support of X is not contained in the image of a single geodesic.
    This excludes the univariate-like case and guarantees that the extreme quantile diverges to infinity along xi.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A data-based notion of quantiles on Hadamard spaces." pith.science (2026). https://pith.science/paper/T4YDCHGP

@misc{pith2026250612534,
  author       = {Pith},
  title        = {Pith review of: A data-based notion of quantiles on Hadamard spaces},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T4YDCHGP}},
  note         = {Machine review of arXiv:2506.12534}
}
read the original abstract

This paper defines an alternative notion, described as data-based, of geometric quantiles on Hadamard spaces, in contrast to the existing methodology, described as parameter-based. In addition to having the same desirable properties as parameter-based quantiles, these data-based quantiles are shown to have several theoretical advantages related to large-sample properties like strong consistency and asymptotic normality, breakdown points, extreme quantiles and the gradient of the loss function. Using simulations, we explore some other advantages of the data-based framework, including simpler computation and better adherence to the shape of the distribution, before performing experiments with real diffusion tensor imaging data lying on a manifold of symmetric positive definite matrices. These experiments illustrate some of the uses of these quantiles by testing the equivalence of the generating distributions of different data sets and measuring distributional characteristics.

Figures

Figures reproduced from arXiv: 2506.12534 by the authors.

Figure 1
Figure 1. The black points are the data points from the first simulated data set, the smaller blue points the computed quantiles, and the blue curves connecting them the isoquantile curves for β ∈ {0.2, 0.4, 0.6, 0.8, 0.98}. The top row shows attempts at computing parameter-based quantiles and the bot￾tom row data-based quantiles. The results in the first column were computed using the true gradients, the second column using … view at source ↗
Figure 2
Figure 2. The black points are the data points from the second simulated data set, the smaller blue points the computed quantiles, and the blue curves connecting them the isoquantile curves for β ∈ {0.2, 0.4, 0.6, 0.8, 0.98}. The top row shows attempted at computing parameter-based quantiles and the bottom row data-based quantiles. The results in the first column were computed using the true gradi￾ents, the second column usin… view at source ↗
Figure 3
Figure 3. Corpus callosum DTI data, visualized as ellipsoids in three-dimensional space. ξI ∈ TIP3 ∼= S3, we use 8 different values of ξ: ξ1 = 1 √ 3       1 0 0 0 1 0 0 0 1       , ξ2 = 1 √ 2       0 1 0 1 0 0 0 0 0       , ξ3 = 1 √ 2       0 0 1 0 0 0 1 0 0       , ξ4 = 1 √ 2       0 0 0 0 0 1 0 1 0       , ξ5 = −ξ1, ξ6 = −ξ2, ξ7 = −ξ3, and ξ8 = −ξ4; ξ1, ξ2, ξ3, and ξ4 are mutu… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Quantiles for the data in [PITH_FULL_IMAGE:figures/full_fig_p029_4.png]
Figure 5
Figure 5. Figure 5: Subset of the original data and quantiles, visualized as ellipsoids in three-dimensional space. respectively. We calculate these quantities, as well as the dispersion measures in (9), with both non-extreme and extreme values of β and β ′ : in the non-extreme case, (β, …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 26 canonical work pages

  1. [1]

    J., Mattiello, J., and LeBihan, D

    Basser, P. J., Mattiello, J., and LeBihan, D. (1994). MR diffusion tensor spectroscopy and imaging. Biophysical Journal , 66:259--267

  2. [2]

    Bridson, M. R. and Haefliger, A. (1999). Metric Spaces of Non-Positive Curvature . Springer Berlin, Heidelberg

  3. [3]

    Chakraborty, B. (2003). On multivariate quantile regression. Journal of Statistical Planning and Inference , 110:109--132

  4. [4]

    and Goga, C

    Chaouch, M. and Goga, C. (2010). Design-based estimation for geometric quantiles with application to outlier detection. Computational Statistics & Data Analysis , 54:2214--2229

  5. [5]

    Chaudhuri, P. (1996). On a geometric notion of quantiles for multivariate data. Journal of the American Statistical Association , 91:862--872

  6. [6]

    and Boumal, N

    Criscitiello, C. and Boumal, N. (2020). An accelerated first-order method for non-convex optimization on manifolds. arXiv:2008.02252v2 [math.OC]

  7. [7]

    Some differential properties of $GL_n(\mathbb{R})$ with the trace metric

    Dolcetti, A. and Pertici, D. (2014). Some differential properties of GL _n( R ) with the trace metric. arXiv:1412.4565v1 [math.DG]

  8. [8]

    and Pertici, D

    Dolcetti, A. and Pertici, D. (2018). Differential properties of spaces of real symmetric matrices. arXiv:1807.01113v1 [math.DG]

Show all 28 references
  1. [9]

    and Huckemann, S

    Eltzner, B. and Huckemann, S. F. (2019). A smeary central limit theorem for manifolds with application to high-dimensional spheres. The Annals of Statistics , 47:3360--3381

  2. [10]

    T., Venkatasubramanian, S., and Joshi, S

    Fletcher, P. T., Venkatasubramanian, S., and Joshi, S. (2009). The geometric median on R iemannian manifolds with application to robust atlas estimation. NeuroImage , 45:143--152

  3. [11]

    and Stupfler, G

    Girard, S. and Stupfler, G. (2017). Intriguing properties of extreme geometric quantiles. REVSTAT--Statistical Journal , 15:107--139

  4. [12]

    Huber, P. J. (1967). The behavior of maximum likelihood estimates under nonstandard conditions. Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability , 1:221--233

  5. [13]

    Koltchinskii, V. I. (1997). M -estimation, convexity and quantiles. The Annals of Statistics , 25:435--477

  6. [14]

    and Paindaveine, D

    Konen, D. and Paindaveine, D. (2024). On the robustness of spatial quantiles. Annales de l'Institut Henri Poincar\'e', Probabilit\'es et Statistiques

  7. [15]

    and Stark, T

    K\"ostenberger, G. and Stark, T. (2023). Robust signal recovery in H adamard spaces. arXiv:2307.06057v1[math.ST]

  8. [16]

    and Lee, J

    Lee, K. and Lee, J. (2018). Optimal B ayesian minimax rates for unconstrained large covariance matrices. Bayesian Analysis , 13:1215--1233

  9. [17]

    Lopuhaa, H. P. and Rousseeuw, P. J. (1991). Breakdown points of affine equivariant estimators of multivariate location and covariance matrices. The Annals of Statistics , 19:229--248

  10. [18]

    and Pierpaoli, C

    Pajevic, S. and Pierpaoli, C. (1999). Color schemes to represent the orientation of anisotropic tissues from diffusion tensor data: application to white matter fiber tract mapping in the human brain. Magnetic Resonance in Medicine , 42:526--540

  11. [19]

    Rousseeuw, P. J. and Croux, C. (1992). Explicit scale estimators with high breakdown point. In Dodge, Y., editor, L1 Statistical Analysis and Related Methods , pages 77--92. North Holland

  12. [20]

    Shin, H.-Y. (2024). Radial fields on the manifolds of symmetric positive definite matrices. arXiv:2405.03297v1 [math.DG]

  13. [21]

    and Oh, H.-S

    Shin, H.-Y. and Oh, H.-S. (2023). Quantiles on global non-positive curvature spaces. arXiv:2312.10870v2 [math.ST]

  14. [22]

    and Oh, H.-S

    Shin, H.-Y. and Oh, H.-S. (2025). Geometric quantile-based measures of multivariate distributional characteristics. Statistics and Proability Letters , 216:110325

  15. [23]

    Sturm, K.-T. (2002). Nonlinear martingale theory for processes with values in metric spaces of nonpositive curvature. The Annals of Probability , 30:1195--1222

  16. [24]

    Sturm, K.-T. (2003). Probability measures on metric spaces of nonpositive curvature. Contemporary Mathematics , 338

  17. [25]

    S., Menon, A., and Kumar, S

    Weber, M., Zaheer, M., Rawat, A. S., Menon, A., and Kumar, S. (2020). Robust large-margin learning in hyperbolic space. 34th Conference on Neural Information Processing Systems (NeurIPS 2020)

  18. [26]

    and Park, B

    Yun, H. and Park, B. U. (2023). Exponential concentration for geometric-median-of-means in non-positive curvature spaces. Bernoulli , 29:2927--2960

  19. [27]

    and Sra, S

    Zhang, H. and Sra, S. (2016). First-order methods for geodesically convex optimization. JMLR: Workshop and Conference Proceedings , 49:1--22

  20. [28]

    G., Li, Y., Hall, C., and Lin, W

    Zhu, H., Chen, Y., Ibrahim, J. G., Li, Y., Hall, C., and Lin, W. (2009). Intrinsic regression models for positive-definite matrices with applications to diffusion tensor imaging. Journal of the American Statistical Association , 104:1203--1212

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.