Pith. sign in

REVIEW 2 major objections 5 minor 22 references

Alpha Procrustes metrics between positive definite operators: a unifying formulation for the Bures-Wasserstein and Log-Euclidean/Log-Hilbert-Schmidt metrics

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A parametrized family of distances, the Alpha Procrustes distances, unifies the Bures-Wasserstein and Log-Euclidean distances on positive definite operators, and extends to infinite-dimensional Hilbert-Schmidt operators.

desk verdict A genuine finite-dimensional unification; the infinite-dimensional claim needs a scope correction before the paper can stand. read the letter →

arxiv 1908.09275 v1 pith:V2DV6AWX submitted 2019-08-25 math.FA

classification math.FA MSC 15B4846E2247B10
keywords AlphaProcrustesdistanceBures-WassersteinLog-EuclideanLog-Hilbert-SchmidtpositivedefiniteoperatorsGaussianmeasuresreproducingkernelHilbertspaceRiemanniansubmersion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces a one-parameter family of distances between positive definite matrices, called Alpha Procrustes distances, defined by aligning matrix powers up to unitary rotations. It proves that at parameter α=1/2 the distance is exactly twice the Bures-Wasserstein (optimal-transport) distance, and that as α→0 it converges to the Log-Euclidean distance. The same construction is carried over to positive definite unitized Hilbert-Schmidt operators on a Hilbert space, where it recovers the Bures-Wasserstein and Log-Hilbert-Schmidt distances as special cases, and to Gaussian measures, where it generalizes the L2-Wasserstein distance. In reproducing kernel Hilbert spaces the distances all have closed forms in terms of kernel Gram matrices. If correct, the paper shows that two previously separate geometries of covariance matrices are the endpoints of one continuous family with explicit formulas.

What carries the argument

The central object is the Alpha Procrustes distance, defined by the unitary alignment problem $d^{\alpha}_{\mathrm{proE}}(A,B)=\min_{U}\|(A^{\alpha}-B^{\alpha}U)/\alpha\|_{F}$ and given in closed form by the trace expression in Eq. (8). Three mechanisms carry the argument. First, the polar decomposition of $B^{\alpha}A^{\alpha}$ solves the Procrustes problem and produces the explicit trace formula. Second, the maps $\pi_{\alpha}(G)=(\alpha^{2}GG^{*})^{1/(2\alpha)}$ from the general linear group to the SPD cone become Riemannian submersions under the metric (29), so geodesics and distances descend from straight lines in $\mathrm{GL}(n)$, which is how the geodesic formula (34) and the metric property arise. Third, in infinite dimensions the algebra of unitized Hilbert-Schmidt operators $\mathrm{HS}_{\mathrm{X}}(H)=\mathrm{HS}(H)+\mathbb{R}I$ and Lemma 1 (polar decomposition stays inside the unitized class) keep the trace expressions finite. The Araki-Lieb-Thirring inequality supplies the comparison with power-Euclidean distances.

What would settle it

Produce an invertible operator $I+A$ with $A$ Hilbert-Schmidt whose polar unitary factor $I+U$ has $U$ not Hilbert-Schmidt; such an example would falsify Lemma 1 and thereby invalidate the trace formula (52) for the infinite-dimensional Alpha Procrustes distance.

Watch

Extended reading notes

Core claim

At the core is the closed-form solution of the Alpha Procrustes optimization problem: $d^{\alpha}_{\mathrm{proE}}(A,B)=\frac{1}{|\alpha|}\min_{U\in\mathrm{U}(n)}\|A^{\alpha}-B^{\alpha}U\|_{F}=\big(\frac{1}{\alpha^{2}}\mathrm{tr}[A^{2\alpha}+B^{2\alpha}-2(A^{\alpha}B^{2\alpha}A^{\alpha})^{1/2}]\big)^{1/2}$ for every nonzero α. The paper proves that this is a metric on the SPD cone, equal to twice the Bures-Wasserstein distance at α=1/2 and to $\|\log A-\log B\|_{F}$ as α→0. It then shows that these distances are the Riemannian distances of explicit Riemannian metrics on the SPD manifold, obtained as Riemannian submersions from the general linear group with the Frobenius metric, with the Wasserstein Riemannian metric and the Log-Euclidean metric as special cases. In infinite dimensions, the family is defined on unitized positive definite Hilbert-Schmidt operators, where the α→0 limit is the Log-Hilbert-Schmidt distance and trace-class operators with α≥1/2 recover the Bures-Wasserstein distance; in the RKHS setting all distances reduce to formulas in kernel Gram matrices.

Load-bearing premise

The infinite-dimensional construction rests on Lemma 1, which asserts that the polar decomposition of any invertible unitized Hilbert-Schmidt operator $A+\gamma I$ has unitary factor $I+R$ with $R$ Hilbert-Schmidt; if the unitary factor could leave the unitized Hilbert-Schmidt class, the explicit Alpha Procrustes distance on $\mathrm{PC}_2(H)$ would not be finite.

Editorial extensions

If this is right

  • There is a one-parameter family of Riemannian metrics on the SPD cone whose geodesic distances interpolate between the Bures-Wasserstein and Log-Euclidean geometries, so computations set up in one geometry can be continuously deformed to the other.
  • The same family defines metrics on positive definite unitized Hilbert-Schmidt operators, with the Log-Hilbert-Schmidt distance as the α→0 limit and the Bures-Wasserstein distance recovered for trace-class operators when $\alpha\geq 1/2$.
  • For Gaussian measures on Euclidean and separable Hilbert spaces, the family yields a parametrized set of metrics containing the squared L2-Wasserstein distance at α=1/2 and a Log-Euclidean-type distance as α→0.
  • In RKHS settings, all members of the family between empirical covariance operators are computable in closed form from the kernel Gram matrices K[X], K[Y], and K[X,Y], without ever constructing the covariance operators explicitly.
  • For any fixed exponent α, the Alpha Procrustes distance is never larger than the power-Euclidean distance with the same α, and the two coincide exactly when the matrices commute.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the family is geodesically complete for every α, it would give a principled way to average covariance matrices while sweeping a trade-off between Wasserstein and Log-Euclidean behavior, which is currently a modeling choice in diffusion-tensor imaging and covariance tracking.
  • The Gram-matrix formulations suggest all RKHS distances in the family can be evaluated from $m\times m$ kernels in $O(m^3)$ time, making the parameter α cheap to select by cross-validation or likelihood.
  • The submersion construction is not tied to the symmetric-positive cone; the same pattern may define analogous Alpha Procrustes metrics on other homogeneous spaces with polar decompositions, such as Grassmannians or flag manifolds.
  • The paper leaves open the case $\gamma\neq\nu$ for two different unitization scales in infinite dimensions; a closed form there would complete the unification and is the natural next step.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces a one-parameter family of "Alpha Procrustes" distances on the SPD matrix cone, defined by minimizing a Frobenius norm between scaled powers A^α and B^α U over unitaries. It proves an explicit trace formula, shows that α=1/2 gives twice the Bures-Wasserstein distance and that α→0 gives the Log-Euclidean distance, compares the family with power-Euclidean distances, and constructs a Riemannian submersion whose distance is the Alpha Procrustes distance. The construction is then extended to positive definite unitized Hilbert-Schmidt operators of the form A+γI, with the scalar part γ fixed, recovering a fiberwise Bures-Wasserstein distance and a fiberwise Log-Hilbert-Schmidt limit. The last sections provide closed-form RKHS Gram-matrix formulas for covariance operators and corresponding distances between Gaussian measures.

Significance. The finite-dimensional part of the paper is coherent and useful: the explicit formula in Theorem 1, the special-case recoveries, the comparison theorem via the Araki-Lieb-Thirring inequality, and the Riemannian submersion interpretation are all substantive and appear to be correct. The paper also supplies full proofs for the main finite-dimensional statements, and the RKHS formulas, if fully justified, would be practically valuable for kernel methods. The main weakness is in the infinite-dimensional part: the claimed unification with the Log-Hilbert-Schmidt distance on the full space PC2(H) is not actually established, because the distance is defined only on fixed scalar fibers. For these reasons the central claim needs substantial revision or extension.

major comments (2)
  1. [§3, Definition 3, Theorem 11, Theorem 13, Remark 2] The claimed infinite-dimensional unification with the Log-Hilbert-Schmidt distance is only proved within a fixed fiber, not on the full set PC2(H). Definition 3, Theorem 11, and Theorem 13 define d^α_proHS only for pairs (A+γI) and (B+γI) with the same γ, i.e. inside PC2(H)(γ) = {A+γI : A ∈ HS(H)}. For γ≠ν, the difference (γ−ν)I is not Hilbert-Schmidt, and Remark 2 explicitly defers this case to future work. The Log-Hilbert-Schmidt distance in Eq. (43), however, is defined on all of PC2(H), including pairs with different scalar parts. Consequently Theorem 12 recovers the Log-Hilbert-Schmidt distance only for operators sharing the same γ, and the abstract's statement that the family generalizes to "the set of positive definite Hilbert-Schmidt operators" and recovers the Log-Hilbert-Schmidt distance as a special case is overbroad. The manuscript should either extend the construction to γ≠ν or explicitly reframe the infinite-dimensional contribution as a fiberwise unification; the same caveat affects the γ-dependent Gaussian-measure distances in Theorems 15 and 16 and Eq. (71).
  2. [§3.4, Proposition 5, Corollary 5, Eqs. (74), (79), (97)] The passage from the repeated-row 3×3 block matrix in Eq. (74) to the limiting block matrix used in the proof of Corollary 5 is not justified. In Eq. (74) the block matrix has identical second and third rows, and this structure is repeated in Eqs. (79) and (97). In the proof of Corollary 5, however, the displayed limit matrix is [[0,0,(A^*A)^{2α}A^*B(B^*B)^{2α−1}],[0,0,0],[0,0,B^*A(A^*A)^{2α−1}A^*B(B^*B)^{2α−1}]], which has a different zero pattern. The proof does not explain how the repeated-row matrix is transformed into this matrix, and Lemma 11 is stated only for the latter block form. Since the closed-form Gram-matrix formulas in Theorem 17 and Corollary 5 depend on this trace identity, the missing spectral argument is load-bearing for the RKHS section. Please either supply the missing equivalence or correct the displayed matrices.
minor comments (5)
  1. [Theorem 2, Eq. (12)] The displayed limit in Eq. (12) has an unmatched bracket in the trace expression; the bracket should be closed before the equality sign.
  2. [Lemma 4, proof] In the proof of Lemma 4, the final displayed estimate says that the bound tends to zero "as α → ∞"; the intended limit is α → 0.
  3. [Proposition 3, proof] In the long displayed computation in the proof of Proposition 3, the term "(I+V)^α(I+V)" appears where "(I+B)^α(I+V)" is evidently intended.
  4. [Lemma 11, Corollary 5] Lemma 11 is stated for finite n×n matrices, but in the proof of Corollary 5 it is applied to block operators on the separable Hilbert space H1; the infinite-dimensional version needed for the proof should be stated and proved.
  5. [Theorem 3, proof] The proof of Theorem 3 is only one sentence for α≠0 and does not explicitly handle the limiting case α=0; the metric property of the Log-Euclidean distance should be cited there.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Alpha Procrustes family is derived from an independent Procrustes optimization, with special cases recovered by explicit limits and the Riemannian distance proven from a submersion construction.

full rationale

I walked the derivation chain. The finite-dimensional distance is defined by the Procrustes minimization Eq. (7), not by the claimed special cases; Theorem 1 derives the trace formula (8) from that minimization, and the α = 1/2 and α → 0 recoveries are computed and proven (Eq. (9), Theorem 2), not assumed. The Riemannian metric in Eq. (29) is constructed by requiring πα to be a Riemannian submersion, and Theorem 8 derives the distance and geodesic from that submersion structure rather than declaring the equality. The infinite-dimensional extension similarly defines d^α_proHS in Eqs. (49)/(55) and proves the explicit formulas in Theorems 10/11 using polar decomposition and trace-class lemmas; the Log-Hilbert-Schmidt limit in Theorem 12 is proven by series expansions in Lemmas 4 and 5. Citations to the author's prior work [12] and [14] supply the Log-Hilbert-Schmidt metric property and the h_α identities used in the RKHS computations, but these are published external results and serve as ingredients, not as the source of the new family's metric property, which is proved for α ≠ 0 in Theorem 13. The main scope limitation—Theorem 11 and Theorem 13 cover only pairs within the same fiber PC2(H)(γ), with Remark 2 deferring the γ ≠ ν case—is an incompleteness of the infinite-dimensional contribution, not a circularity. No load-bearing step reduces by construction to its own input.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

No new physical or mathematical entities are postulated. The unitized Hilbert-Schmidt space HSX(H) and PC2(H) were introduced in prior work by Larotonda and by the author; h_α(E) is from the author's earlier paper [14].

free parameters (2)
  • α (Alpha Procrustes family parameter)
    User-chosen interpolation index; α=1/2 recovers Bures-Wasserstein, α→0 recovers Log-Euclidean. Not fitted to data.
  • γ (unitized operator regularization scale)
    Positive shift added to covariance/trace-class operators so that log and fractional powers stay in the unitized Hilbert-Schmidt class; user-chosen, not fitted.
assumptions (4)
  • standard math Araki-Lieb-Thirring trace inequality
    Used in Theorem 4 to prove d_proE ≤ d_E,α for non-commuting matrices.
  • standard math Riemannian submersion geodesic lifting theorem
    Used in Theorem 7 and 8 to transfer geodesics from GL(n) to Sym++(n).
  • standard math Polar decomposition for matrices and operators
    Used throughout to derive explicit Procrustes distances (Theorems 1, 10).
  • standard math Spectral theorem and functional calculus for positive operators
    Used to define A^α, log(A), h_α(E) and to compute traces of operator powers in the infinite-dimensional setting.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Alpha Procrustes metrics between positive definite operators: a unifying formulation for the Bures-Wasserstein and Log-Euclidean/Log-Hilbert-Schmidt metrics." pith.science (2026). https://pith.science/paper/V2DV6AWX

@misc{pith2026190809275,
  author       = {Pith},
  title        = {Pith review of: Alpha Procrustes metrics between positive definite operators: a unifying formulation for the Bures-Wasserstein and Log-Euclidean/Log-Hilbert-Schmidt metrics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V2DV6AWX}},
  note         = {Machine review of arXiv:1908.09275}
}
read the original abstract

This work presents a parametrized family of distances, namely the Alpha Procrustes distances, on the set of symmetric, positive definite (SPD) matrices. The Alpha Procrustes distances provide a unified formulation encompassing both the Bures-Wasserstein and Log-Euclidean distances between SPD matrices. We show that the Alpha Procrustes distances are the Riemannian distances corresponding to a family of Riemannian metrics on the manifold of SPD matrices, which encompass both the Log-Euclidean and Wasserstein Riemannian metrics. This formulation is then generalized to the set of positive definite Hilbert-Schmidt operators on a Hilbert space, unifying the infinite-dimensional Bures-Wasserstein and Log-Hilbert-Schmidt distances. In the setting of reproducing kernel Hilbert spaces (RKHS) covariance operators, we obtain closed form formulas for all the distances via the corresponding kernel Gram matrices. From a statistical viewpoint, the Alpha Procrustes distances give rise to a parametrized family of distances between Gaussian measures on Euclidean space, in the finite-dimensional case, and separable Hilbert spaces, in the infinite-dimensional case, encompassing the 2-Wasserstein distance, with closed form formulas via Gram matrices in the RKHS setting. The presented formulations are new both in the finite and infinite-dimensional settings.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 22 canonical work pages

  1. [1]

    Letters i n Mathematical Physics 19(2), 167–170 (1990) 20 H` a Quang Minh

    Araki, H.: On an inequality of Lieb and Thirring. Letters i n Mathematical Physics 19(2), 167–170 (1990) 20 H` a Quang Minh

  2. [2]

    Arsigny, V., Fillard, P., Pennec, X., Ayache, N.: Geometr ic means in a novel vector space structure on symmetric positive-definite matrices. S IAM J. on Matrix An. and App. 29(1), 328–347 (2007)

  3. [3]

    Expositiones Mathematicae (2018)

    Bhatia, R., Jain, T., Lim, Y.: On the Bures–Wasserstein di stance between positive definite matrices. Expositiones Mathematicae (2018)

  4. [4]

    Journal of Multivariate Analysis 12(3), 450 – 455 (1982)

    Dowson, D., Landau, B.: The Fr´ echet distance between mul tivariate normal dis- tributions. Journal of Multivariate Analysis 12(3), 450 – 455 (1982)

  5. [5]

    Anna ls of Applied Statistics 3, 1102–1123 (2009)

    Dryden, I., Koloydenko, A., Zhou, D.: Non-Euclidean stat istics for covariance ma- trices, with applications to diffusion tensor imaging. Anna ls of Applied Statistics 3, 1102–1123 (2009)

  6. [6]

    Mathematische Nachrichten 147(1), 185–203 (1990)

    Gelbrich, M.: On a formula for the L2 Wasserstein metric be tween measures on Euclidean and Hilbert spaces. Mathematische Nachrichten 147(1), 185–203 (1990)

  7. [7]

    Michigan Math

    Givens, C.R., Shortt, R.M.: A class of Wasserstein metric s for probability distri- butions. Michigan Math. J. 31(2), 231–240 (1984)

  8. [8]

    Differential Geometry and its Applications 25, 679–700 (2007)

    Larotonda, G.: Nonpositive curvature: A geometrical app roach to Hilbert-Schmidt operators. Differential Geometry and its Applications 25, 679–700 (2007)

Show all 22 references
  1. [9]

    In: Lieb, E., S.B., Wrightman, A

    Lieb, E., Thirring, W.: Inequalities for the moments of th e eigenvalues of the Schr¨ odinger Hamiltonian and their relation to Sobolev ine qualities. In: Lieb, E., S.B., Wrightman, A. (eds.) Studies in Mathematical Physics . Princeton University Press (1976)

  2. [10]

    Information Geometry 1(2), 137–179 (Dec 2018)

    Malag` o, L., Montrucchio, L., Pistone, G.: Wasserstein Riemannian geometry of Gaussian densities. Information Geometry 1(2), 137–179 (Dec 2018)

  3. [11]

    San khya A pp

    Masarotto, V., Panaretos, V., Zemel, Y.: Procrustes met rics on covariance opera- tors and optimal transportation of Gaussian processes. San khya A pp. 1–42 (2018)

  4. [12]

    In: Advances in Ne ural Information Pro- cessing Systems 27 (NIPS 2014), pp

    Minh, H.Q., Biagio, M.S., Murino, V.: Log-Hilbert-Schm idt metric between posi- tive definite operators on Hilbert spaces. In: Advances in Ne ural Information Pro- cessing Systems 27 (NIPS 2014), pp. 388–396 (2014)

  5. [13]

    Linear Algebra and Its Applicat ions 528, 331–383 (2017)

    Minh, H.: Infinite-dimensional Log-Determinant diverg ences between positive defi- nite trace class operators. Linear Algebra and Its Applicat ions 528, 331–383 (2017)

  6. [14]

    In: Information Geo metry and its Appli- cations IV

    Minh, H.: Infinite-dimensional Log-Determinant diverg ences III: Log-Euclidean and Log-Hilbert–Schmidt divergences. In: Information Geo metry and its Appli- cations IV. pp. 209–243. Springer (2018)

  7. [15]

    In: International Conference on Geometric Science of Information

    Minh, H.: A unified formulation for the Bures-Wasserstei n and Log-Euclidean/Log- Hilbert-Schmidt distances between positive definite opera tors. In: International Conference on Geometric Science of Information. Springer ( 2019)

  8. [16]

    Syn- thesis Lectures on Computer Vision 7(4), 1–170 (2017)

    Minh, H., Murino, V.: Covariances in computer vision and machine learning. Syn- thesis Lectures on Computer Vision 7(4), 1–170 (2017)

  9. [17]

    Linear Algebra and its Applications 48, 257 – 263 (1982)

    Olkin, I., Pukelsheim, F.: The distance between two rand om vectors with given dispersion matrices. Linear Algebra and its Applications 48, 257 – 263 (1982)

  10. [18]

    Springer Science & Busi- ness Media (2008)

    Steinwart, I., Christmann, A.: Support vector machines . Springer Science & Busi- ness Media (2008)

  11. [19]

    Osaka Journal of Math- ematics 48(4), 1005–1026 (2011)

    Takatsu, A.: Wasserstein geometry of Gaussian measures . Osaka Journal of Math- ematics 48(4), 1005–1026 (2011)

  12. [20]

    Villani, C.: Optimal transport: old and new, vol. 338. Sp ringer Science & Business Media (2008)

  13. [21]

    Wang, B.Y., Zhang, F.: Trace and eigenvalue inequalitie s for ordinary and Hadamard products of positive semidefinite Hermitian matri ces. SIAM journal on matrix analysis and applications 16(4), 1173–1183 (1995) Unifying Wasserstein and Log-Euclidean/Log-Hilbert-Sch midt metric...

  14. [22]

    For any A0 ∈ Sym++(n), D log(A0) : Sym( n) → Sym(n) and D exp(log(A0)) : Sym(n) → Sym(n) are invertible operators, thus they both have zero null spaces

    ◦ Dg(A0)(X). For any A0 ∈ Sym++(n), D log(A0) : Sym( n) → Sym(n) and D exp(log(A0)) : Sym(n) → Sym(n) are invertible operators, thus they both have zero null spaces. It follows that ker(Dπα (A0)) = ker( Dg(A0)) = {X ∈ M(n) : XA ∗ 0 + A0X ∗ = 0} = {X ∈ M(n) : XA ∗ 0 is skew-sym...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.