{"id":"0134fa49-9584-46ec-835f-fa1e2443c7d7","arxiv_id":"2507.11387","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The authors extend Energy Distances to negative and higher orders, connect them to Fourier-based metrics, and construct scale-invariant divergence measures for high-dimensional machine learning.","lead":"This paper reviews three families of probability divergences rooted in kinetic theory, extends Energy Distances to new parameter ranges, and proposes a whitening step that makes divergences scale invariant. It positions these as tools for comparing high-dimensional data in machine learning, with an illustrative ESG-finance model selection.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3's moment condition is internally inconsistent: for α=3 it allows unequal means, yet the Fourier metric F_{n+3} diverges, so identity (7.74) fails on simple Gaussian inputs.","rationale":"The reader's weakest-assumption analysis correctly flagged Theorem 3 as the load-bearing step, but framed it mainly as an unverified integration-by-parts argument. My reading sharpens this: the stated moment condition in Theorem 3 is not merely unproved but internally inconsistent with the paper's own F_s finiteness criterion in §6.2. For α=3 the theorem permits unequal means, while §6.2 requires equal means for F_{n+3} to be finite; a concrete Gaussian counterexample makes the failure explicit. This is a genuine soft spot in the central claim, not a manufactured objection. I do not move the reader's CONDITIONAL verdict: the paper is a review with useful known material, the application uses α∈(0,2), and the error in Theorem 3 may be correctable by strengthening the moment hypotheses. But the theorem needs to be repaired or qualified before the claimed extension to all α>−n can be accepted. The concrete test is a one-line analytic check that would settle the issue immediately.","tokens_in":36360,"tokens_out":16648,"duration_ms":175959,"concrete_test":"Fix n=1, let μ=N(m1,1), ν=N(m2,1) with m1≠m2, and set α=3. Compute F_4(μ,ν)=∫|μ̂−ν̂|^2/|ξ|^4 dξ: near 0 the integrand behaves like |m1−m2|^2|ξ|^{-2}, so the integral diverges. Evaluate the right-hand side of (7.73) directly; it is finite. This contradicts (7.74). Also check the hypothesis of Theorem 3: the moment condition floor((3−3)/2)=0 is satisfied (only equal mass), while §6.2's criterion for F_4 requires equal means, which fail. Repeating the test for α=2k+β with k=1,2 and Gaussian inputs with matched moments of order ≤k but differing at order k+1 will show exactly where the stated moment condition must be corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 3 (Eq. 7.73) asserts Eα equals CαF_{n+α} whenever μ and ν have equal moments up to order floor((3⌊α/2⌋−α)/2). This condition contradicts the paper's own finiteness criterion for F_s in §6.2. For s=n+α, F_s is finite only if the first l moments agree with l > α/2; for α=3 this requires equal means (moments up to order 1). The theorem's condition for α=3 gives floor((3−3)/2)=0, i.e. only equal total mass, so it permits μ=N(m1,I) and ν=N(m2,I) with m1≠m2. For such inputs, near ξ=0 we have |μ̂−ν̂|^2 ~ |m1−m2|^2|ξ|^2, making the integrand of F_{n+3} behave like |ξ|^{-2} in n dimensions, so F_{n+3} diverges. Meanwhile the right-hand side of (7.73) is finite for these Gaussians. Thus the claimed equality (7.74) cannot hold under the stated hypotheses. This is not merely a missing boundary-term argument in the integration by parts at (7.68)–(7.71); the asserted moment condition is insufficient to guarantee the Fourier representation. Since Theorem 3 underlies Theorems 4, 5, and the metric/sub-additivity/equivalence claims for all α > −n, the central mathematical assertion of the paper is not established as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a review-style contribution that connects divergence measures from kinetic theory (relative entropy, Fisher information, Wasserstein distances, Fourier-based metrics, and energy distances) to modern machine-learning practice. Its main new mathematical claims are: (i) an extension of the Székely–Rizzo energy distance E_α to every α > -n, α not an even integer, with an identity E_α = C_α F_{n+α} to the Fourier-based metric (Definition 12 and Theorem 3); (ii) a series of structural properties for these extended divergences, including sub-additivity by convolution, sub-additivity by convex convolution for α ≥ 4, and unbiased gradients; and (iii) a framework for whitened, scale-invariant divergences based on scale-stable whitening transformations, with claims that metricity, equivalence, sub-additivity, and unbiasedness survive whitening. The paper also contains a heuristic kinetic-theory description of feedforward neural networks and an application to ESG-based financial prediction.","tokens_in":36710,"tokens_out":11039,"duration_ms":138601,"significance":"If the main theorem were correct, the paper would contribute a useful unification: the energy distance, traditionally restricted to 0 < α < 2, would be identified with a Fourier metric for all admissible α > -n, and the whitening construction would yield computable scale-invariant divergences for high-dimensional data. The known parts of the paper are sound: the energy–Fourier equivalence for 0 < α < 2, the Wasserstein–Fourier comparison, and the overall taxonomy of the desiderata in Section 5 are standard and correctly presented. The paper is also commendably explicit about its desiderata and about the moment conditions needed for the Fourier metrics. However, the central new result, Theorem 3, is not established as written and in fact fails under the paper's own stated hypotheses; moreover, the unbiased-gradient theorem for whitened divergences uses an unjustified identification of the whitened empirical measure. These are load-bearing issues for the paper's main claims.","major_comments":[{"comment":"The moment condition in Theorem 3 is internally inconsistent with the paper's own finiteness criterion for F_s in §6.2. For α = 3, Theorem 3 requires only equal moments up to order floor((3k − α)/2) = floor(0) = 0, i.e. equal total mass. But by the criterion in §6.2, the Fourier metric F_{n+3} is finite only when the first l moments agree with l > α/2 = 1.5, hence l ≥ 2. Concretely, take μ = N(m1,I) and ν = N(m2,I) with m1 ≠ m2. These satisfy the theorem's condition for α = 3, but |μ̂(ξ) − ν̂(ξ)|² ≍ |m1 − m2|²|ξ|² near ξ = 0, so the integrand of F²_{n+3} behaves like |m1 − m2|²/|ξ|^{n+1} and the integral diverges. In contrast, the right-hand side of (7.73) is finite for these Gaussian inputs. Thus the asserted identity E_α = C_α F_{n+α} in (7.74) fails under the stated hypotheses. Since Theorem 3 underlies Theorems 4, 5, and the metric, equivalence, and sub-additivity claims for all α > -n, the central mathematical assertion of the paper is not established as written. The moment condition in Definition 12, which requires only l ≥ ⌊α/2⌋, is likewise insufficient.","section":"Theorem 3 / Eq. (7.73)"},{"comment":"The proof of the higher-order energy–Fourier identity is formal. After expanding the polynomial in (7.69), the argument integrates by parts repeatedly to move derivatives from the difference of Fourier transforms onto Ŵ_λ, but it never verifies that the boundary terms at infinity vanish, nor that the repeated derivatives are justified for the measures admitted by the moment conditions. This is not merely a gap in presentation: the counterexample in the previous comment shows that the stated moment hypotheses do not even make the Fourier-side integral finite, so the integration-by-parts step cannot hold under the theorem as stated.","section":"§7.1.2, Eqs. (7.68)–(7.71)"},{"comment":"The proof of Theorem 15 identifies the whitened empirical measure S(μ_N) with the empirical measure of the whitened samples. When the whitening matrix W is estimated from the data, S(μ_N) = W(μ_N) X_N is not the empirical measure of {W(μ_N) X^{(i)}} in the usual sense, because W(μ_N) depends on the same samples. Consequently, the equality E[∇θ D_S(μ_N, νθ)] = E[∇θ D(μ*_N, ν*_θ)] used in the proof is not justified. The unbiased-gradient claim for whitened divergences therefore needs either a substantially different argument, an assumption that the whitening matrix is fixed a priori, or a careful treatment of the estimation error and its effect on the gradient expectation.","section":"Theorem 15 / §8.2.3"},{"comment":"The complexity claim that computing the ZCA-cor or Cholesky whitening matrix takes O(1) operations is incorrect in the empirical setting. If S is applied to an empirical distribution supported on N points, the covariance matrix must first be estimated from the data, which costs O(N n²) operations. The statement in the text that this cost is 'independent of the number of points in the support' is therefore false unless the whitening matrix is assumed to be known from the population covariance, which is not the typical machine-learning scenario described in the paper.","section":"Theorem 16 / §8.2.3"}],"minor_comments":[{"comment":"The sign conventions for E_α^2 in Definition 12, in Theorem 3, and in Eq. (7.72) are confusing and appear mutually inconsistent; for example, Definition 12 uses (−1)^k while Theorem 3 uses (−1)^{⌊α/2⌋+1}. The authors should reconcile these definitions explicitly.","section":"Definition 12 and Theorem 3"},{"comment":"The proof of Theorem 6 assumes the existence of invertible optimal transport maps T and S with T∘S = S∘T = identity. Such invertible maps do not exist for general absolutely continuous measures, e.g. measures with different supports. The inequality being proved is known to be true, but the proof as written needs a different argument.","section":"Theorem 6"},{"comment":"The proof of the scale stability of Cholesky whitening is deferred entirely to the authors' own reference [11]. Since the manuscript relies on this property, a proof sketch, or at least a statement of the exact result in [11], should be included.","section":"Theorem 8"},{"comment":"The name 'Zoloratev' in Definition 6 and later should be 'Zolotarev'. There are also several typos, e.g. 'Unilike' in §7.2.1 and 'aspectes' in §5, which should be corrected.","section":"Throughout"},{"comment":"Table 9.5 reports divergences with very large magnitudes (e.g. 43 million for α = 1.5) but provides no uncertainty quantification, significance tests, or comparison across repeated train-test splits; the conclusion that LINS is consistently best would be more persuasive with such supporting evidence.","section":"Section 9"}],"recommendation":"major_revision","confidential_remarks":"The paper leans heavily on the authors' own recent work (Refs [8, 11, 12]) for key ingredients, including the scale-stability result in Theorem 8. This is not by itself inappropriate, but it means that several load-bearing assertions are being imported from papers that the referees may not have access to. The most serious issue is Theorem 3, whose stated moment condition is demonstrably wrong for α = 3; if the authors can repair the theorem by imposing the correct moment condition (essentially l > α/2) and supply a rigorous boundary-term argument, the paper's core idea may be salvageable, but the current version does not establish its central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper is genuinely useful as a review: it ties together kinetic-theory origins of KL, Fisher, Wasserstein, Fourier metrics, and energy distances, and it packages the whitened-divergence construction from the authors' prior work into a clean set of preservation theorems. Second, the advertised new result — extending the energy distance to α > −n and showing Eα = Cα F_{n+α} — is not proven. The stress-test concern is right: Theorem 3's moment condition is too weak. For α = 3, k = 1, the condition only requires equal total mass, but F_{n+3} needs matching first moments per the paper's own §6.2 criterion. Two Gaussians with different means make the Fourier integral diverge while the proposed Eα remains finite. So equation (7.74) cannot hold under the stated hypotheses. The higher-order extension in §7.1.2 is also formal: the integration by parts at (7.68)–(7.71) skips boundary checks, and the sign convention in (7.72) is asserted without support.\n\nWhat is good: the known 0 < α < 2 case and the Wasserstein–Fourier equivalences are correct, and the paper is honest that some results are re-derivations. The whitening preservation theorems (10–13 and 14 under the stated ideal-divergence assumptions) are straightforward and mostly correct, and they organize useful material from refs [8, 11, 12]. Theorem 15, however, has the second soft spot: it treats the whitened empirical measure as the empirical measure of whitened samples, which fails when the whitening matrix is estimated from the data. That breaks the unbiased-gradient claim in the ML application.\n\nThe application (ESG model selection) is illustrative rather than confirmatory: no code or data, no uncertainty quantification, and the numbers are presented as if the divergence choice is doing work it probably isn't. Minor but real.\n\nBottom line: the review parts deserve readers, and the extension idea is worth pursuing, but the central identity needs either a corrected proof or a corrected statement. As is, the paper is not ready for publication. I'd still send it to a serious referee — the flaws are specific and possibly fixable, and the review portion has value. But I would not cite the new theorems in their present form.","headline":"Interesting review framing and a promising extension idea, but the new theorems — especially Theorem 3 — have a moment condition that contradicts the paper's own finiteness criterion, so the central identity is not established as written.","tokens_in":37250,"tokens_out":5276,"would_cite":false,"duration_ms":58438,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["35B40","35L60","35K55","35Q70","35Q91","35Q92"],"pacs":[],"model":"deepseek-v4-flash","headline":"The energy distance extends to every admissible exponent where it equals a Fourier metric, and whitening makes the family scale-free.","keywords":["energy distance","Fourier-based metrics","scale invariance","whitening","Wasserstein distance","divergence measures","kinetic theory","machine learning"],"falsifier":"Compare the direct double-integral value of $E_\\alpha$ with $C_\\alpha F_{n+\\alpha}$ for two distributions in $\\mathbb{R}^n$ that match moments up to the order Theorem 3 requires but differ at the next order; a mismatch for any $\\alpha>-n$, $\\alpha\\notin 2\\mathbb{Z}$, would refute display 7.74. Separately, on a finite sample, estimate the whitening matrix from the data and test whether $\\mathbb{E}[\\nabla_\\theta D_S(\\mu_N,\\nu_\\theta)]=\\nabla_\\theta D_S(\\mu,\\nu_\\theta)$; if it fails, Theorem 15 does not hold for data-dependent whitening.","tokens_in":36178,"feed_emoji":"📐","tokens_out":8338,"duration_ms":95494,"temperature":0.7,"pith_summary":"The paper's central claim is that the energy distance—a statistical distance styled on Newton's potential—can be extended from its classical exponent range to every exponent $\\alpha > -n$ that is not an even integer, and that in this entire range it coincides, up to a constant, with the Fourier-based metric $F_{n+\\alpha}$. That identity would make the whole family of energy distances topologically equivalent to the Fourier metrics, so a practitioner could select the exponent to tune tail sensitivity without changing the notion of convergence. The paper further claims that applying a scale-stable whitening transformation yields divergences that are scale-invariant while preserving metricity on a quotient space, sub-additivity by convolution, and unbiased gradients. These are precisely the properties needed for unit-invariant losses in machine learning, and the paper demonstrates the choice on an ESG-finance model-selection problem.","feed_headline":"Extended energy distances match Fourier metrics at every order","feed_subtitle":"Whitening makes them scale-free with unbiased gradients, giving machine learning a unit-invariant loss family.","key_machinery":"The workhorse is the Riesz potential kernel $|x-y|^{-\\lambda}$ and its Fourier transform $\\widehat{W}_\\lambda(\\xi)=\\pi^{n/2}2^{n-\\lambda}\\Gamma((n-\\lambda)/2)/\\Gamma(\\lambda/2)\\,|\\xi|^{-(n-\\lambda)}$; repeated integration by parts moves Laplacian powers onto this kernel and turns the double integral defining $E_\\alpha$ into the weighted $L^2$ Fourier integral $C_\\alpha F_{n+\\alpha}^2$. For the scale-invariance half, the machinery is a scale-stable linear whitening map $S(X)=W_\\mu X$ with $W_\\mu^T W_\\mu=\\Sigma_\\mu^{-1}$ and $S(QX)=S(X)$ for diagonal scalings; two such maps (lower-triangular factorization and correlation-based ZCA) are shown to exist, and $S$ induces a quotient $\\mathcal{P}(\\mathbb{R}^n)/\\sim_S$ identified with the identity-covariance measures $\\mathcal{P}_I(\\mathbb{R}^n)$.","core_discovery":"On its own terms, the paper establishes a two-part claim. First, for every $\\alpha > -n$ with $\\alpha \\notin 2\\mathbb{Z}$ and $\\alpha \\neq 0$, the extended energy distance defined by $E_\\alpha(\\mu,\\nu)=(-1)^{\\lfloor\\alpha/2\\rfloor+1}\\int |x-y|^\\alpha\\,d[\\mu(x)-\\nu(x)]\\,d[\\mu(y)-\\nu(y)]$ is well-defined for measures sharing moments up to the order $\\lceil(3k-\\alpha)/2\\rceil$ with $k=\\lfloor\\alpha/2\\rfloor$, and obeys $E_\\alpha(\\mu,\\nu)=C_\\alpha F_{n+\\alpha}(\\mu,\\nu)$ (Theorem 3, display 7.74). Second, given any scale-stable whitening map $S$, the whitened divergence $D_S(X,Y)=D(S(X),S(Y))$ is scale-invariant, agrees with $D$ on the quotient $\\mathcal{P}(\\mathbb{R}^n)/\\sim_S=\\mathcal{P}_I(\\mathbb{R}^n)$, inherits equivalence and sub-additivity by convolution when $D$ is Zolotarev ideal, and inherits unbiased gradients when $D$ has them (Theorems 10–16). The upshot is a computable $O(N^2)$ family of unit-invariant divergences with unbiased gradients, applicable to model comparison across heterogeneous units.","pith_inferences":["The one-dimensional case suggests a full ladder of extensions of Cramér's distance: each exponent $\\alpha$ gives a distance between distribution functions, with negative exponents connecting to fractional Sobolev norms of the difference.","The quotient-space view implies that a scale-invariant loss cannot separate distributions that differ only by a rescaling; for problems where scale is a nuisance, that collapse is a feature, but for problems where scale carries information it would be a hidden modelling choice.","Theorem 15 is the likely failure point in practice; estimating the whitening matrix from the same sample used for the gradient may reintroduce bias, so a held-out or re-estimated whitening step is a natural testable modification.","Because convex sub-additivity kicks in only for $\\alpha \\ge 4$, high-order energy distances may give stronger concentration inequalities for sums of random vectors, at the price of requiring many matched moments."],"forward_implications":["For every admissible $\\alpha$, $E_\\alpha$ and $F_{n+\\alpha}$ induce the same convergence, so the exponent can be chosen for tail behavior without changing the induced topology.","Whitening turns any Zolotarev-ideal divergence into a scale-invariant divergence with sub-additivity by convolution preserved, giving a recipe for unit-invariant comparison functions.","The energy distance has unbiased gradients and $O(N^2)$ sample complexity plus a fixed $O(n^3)$ whitening cost, so it offers a practical alternative to Wasserstein losses whose gradients are biased.","In the reported application, the whitened energy distance selects the same sector-dependent linear model as RMSE, but the energy choice is invariant to the units of the ESG and financial indicators."],"supporting_citations":[{"why":"Defines the energy distance and proves the Fourier representation for the classical range $0<\\alpha<2$, the identity this paper extends.","marker":"[110]"},{"why":"Supplies the Fourier transform of the Riesz potential $|x|^{-\\lambda}$ used in the proof of Theorem 1 and the extension.","marker":"[107]"},{"why":"Gives the Hardy–Littlewood–Sobolev inequality that makes negative-exponent energy distances well-defined.","marker":"[86]"},{"why":"Introduces the Fourier-based metrics $F_s$ that the extended energy distance is shown to equal.","marker":"[67]"},{"why":"Argues that whitening pre-processing enforces scale invariance on divergences, the starting point of Section 8.","marker":"[8]"},{"why":"Proves scale stability of the two whitening processes used to construct $D_S$.","marker":"[11]"},{"why":"Documents the biased-gradient problem of Wasserstein distances that the unbiased-gradient property of energy distances is meant to address.","marker":"[18]"}],"fun_headline_variants":["Energy distances match Fourier metrics at every order","Whitening gives scale-free divergences with unbiased gradients","Kinetic theory divergences become scale-free for AI","Rediscovered: kinetic theory divergences, now scale-free for AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Fourier identity $E_\\alpha=C_\\alpha F_{n+\\alpha}$ follows from a formal repeated integration-by-parts calculation whose boundary terms and moment conditions are not fully checked, and that the whitening map used in the unbiased-gradient theorem can be treated as fixed rather than estimated from finite data.","fun_headline_variants_meta":{"raw":{"variants":["Energy distances match Fourier metrics at every order","Whitening gives scale-free divergences with unbiased gradients","Kinetic theory divergences become scale-free for AI","Rediscovered: kinetic theory divergences, now scale-free for AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000798,"raw_usage":{"total_tokens":3505,"prompt_tokens":932,"completion_tokens":2573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":2509}},"tokens_in":548,"tokens_out":2573,"duration_ms":24762,"temperature":1.0,"reasoning_tokens":2509,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:09:31.599710+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the direct double-integral value of $E_\\alpha$ with $C_\\alpha F_{n+\\alpha}$ for two distributions in $\\mathbb{R}^n$ that match moments up to the order Theorem 3 requires but differ at the next order; a mismatch for any $\\alpha>-n$, $\\alpha\\notin 2\\mathbb{Z}$, would refute display 7.74. Separately, on a finite sample, estimate the whitening matrix from the data and test whether $\\mathbb{E}[\\nabla_\\theta D_S(\\mu_N,\\nu_\\theta)]=\\nabla_\\theta D_S(\\mu,\\nu_\\theta)$; if it fails, Theorem 15 does not hold for data-dependent whitening.","supporting_citations":[{"cited_title":"and Rizzo, M.I., Energy statistics: A class of statistics based on distances.Journal of Statistical Planning and Inference, 143 (2013) 1249–1272","cited_arxiv_id":null,"evidence_quote":"Defines the energy distance and proves the Fourier representation for the classical range $0<\\alpha<2$, the identity this paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Fourier transform of the Riesz potential $|x|^{-\\lambda}$ used in the proof of Theorem 1 and the extension."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Hardy–Littlewood–Sobolev inequality that makes negative-exponent energy distances well-defined."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Fourier-based metrics $F_s$ that the extended energy distance is shown to equal."}],"review_version":1}