REVIEW 4 major objections 4 minor 1 cited by
This paper derives a three-term formula for the squared Wasserstein distance between the cosmic initial density field and an observed galaxy catalog, separating gravitational transport, galaxy bias, and shot noise under a small-fluctuation
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 18:05 UTC pith:WABTV3CU
load-bearing objection A readable, mostly review-level paper whose central three-term Wasserstein decomposition is asserted rather than derived, and whose abstract contradicts the body—worth a look for the known linearized results, not as a new statistic. the 4 major comments →
Wasserstein Distance in Cosmological Structure Formation: An Optimal Transport Perspective
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that, under the small-fluctuation approximation, the expectation value of the squared Wasserstein distance between the initial linear density field ρ_lin and the observed galaxy point process μ_N is approximately ⟨W₂²(ρ_lin, μ_N)⟩ ≃ (1/2π²)∫₀^∞ P_m(k) dk + ∫₀^∞ r ξ_gal(r) dr + 1/(4π^{3/2} n̄ R). The first term is the gravitational transport contribution, derived from the linearized optimal-transport equation ∇²ψ = δ, which yields the matter power spectrum integrated over wavenumber. The second term is the galaxy-bias contribution, arising from the correlation structure of the galaxy point process and weighted by r, the transport displacement scale. The third term
What carries the argument
The central object is the quadratic Wasserstein distance W₂(μ,ν) between probability measures, defined as the minimal expected squared distance over all couplings. Under the small-fluctuation limit, the optimal transport map is linearized to a Poisson equation ∇²ψ = δ for the transport potential ψ, turning the distance into ∫|∇ψ|² d³x, which in Fourier space becomes an integral of the density power spectrum P(k)/k². The generative chain ρ_lin → ρ_NL → λ_gal → μ_N is then combined with the additive decomposition W₂² ≃ W²_grav + W²_bias + W²_samp, where the point-process sampling term is evaluated for a Poisson process and the correlation correction is expressed through the structure factor, y
Load-bearing premise
The load-bearing premise is that the squared Wasserstein distance between the initial density field and the observed galaxy catalog equals the sum of the squared distances of the three separate stages (gravitational evolution, galaxy bias, point-process sampling); the paper asserts this additivity rather than deriving it, and it is not guaranteed by the mathematics of optimal transport.
What would settle it
Compute the exact W₂² between a Gaussian random density field with power spectrum P_m(k) and its Zel'dovich displacement field, then add the independently estimated bias and shot-noise terms from a mock galaxy catalog with known ξ_gal(r), n̄, and smoothing R; if the sum does not match the measured Wasserstein distance between the initial field and the mock catalog within the stated approximation, the three-term decomposition is refuted.
If this is right
- The Wasserstein distance between an initial linear field and an observed galaxy catalog can be computed analytically from the matter power spectrum, the galaxy correlation function, and survey parameters (number density and smoothing scale).
- The transport weighting ∫ r ξ(r) dr measures cumulative matter displacement rather than fluctuation amplitude, giving a different scale sensitivity than the standard ∫ r² ξ(r) dr integral and a direct geometric link to BAO smearing.
- In a finite survey volume, long-wavelength displacement modes (bulk flows) are incompletely sampled, so the measured Wasserstein distance underestimates the true global transport cost.
- Applied to simulations or surveys, the decomposition offers a route to geometrically separate gravitational evolution, galaxy-formation bias, and sampling noise in a single statistic.
Where Pith is reading between the lines
- The stage-by-stage additivity of squared distances (Eq. 117) is a strong assumption: the Wasserstein distance obeys a triangle inequality but not additivity of squares, and the optimal maps for the three stages do not generally compose into a global optimal map. Beyond the small-fluctuation regime, cross-terms between gravity, bias, and sampling are likely to appear.
- The r-weighted integral ∫ r ξ(r) dr may diverge for realistic correlation functions that are steeper than r^{-3} at small scales; a rigorous treatment will need a small-scale cutoff or window, similar to the smoothing scale R used in the shot-noise term.
- The (n̄ R)^{-1} scaling of the shot-noise term suggests a concrete survey-design trade-off: increasing galaxy density or smoothing reduces sampling noise, which could be tested with mock catalogs.
- A direct numerical test is feasible: measure W₂² between initial conditions and mock galaxy catalogs in N-body simulations and compare to the three-term formula; the degree of agreement quantifies the small-fluctuation approximation and the additivity assumption.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes an optimal-transport description of cosmological structure formation. It models the observed galaxy catalog as the endpoint of a generative chain ρ_lin → ρ_NL → λ_gal → μ_N and claims that, in the small-fluctuation approximation, the squared Wasserstein distance W₂²(ρ_lin, μ_N) decomposes additively into gravitational, galaxy-bias, and sampling terms, leading to the abstract formula W₂² ≃ (1/2π²)∫P_m(k)dk + ∫ r ξ_gal(r)dr + 1/(4π^{3/2} \bar n R). Sections II–V review point processes, optimal transport, and the linearized power-spectrum representation. Sections VII–IX introduce the bias and sampling terms, and Sec. X discusses implications. The paper is conceptual: it contains no numerical or simulation validation.
Significance. The motivation—using the Wasserstein distance as a displacement-sensitive cosmological statistic and connecting it to BAO smearing—is potentially interesting, and the paper is clearly written. The linearized relation ⟨W₂²⟩ = (1/2π²)∫P(k)dk for smooth, small-fluctuation fields is a useful reminder of the displacement interpretation. However, the central claim, Eq. (117), is asserted rather than proved, and the shot-noise term is not derived from optimal transport. If a rigorous derivation could be supplied, the result would be a valuable new statistic; as it stands, the paper offers an intuitive analogy rather than a supported quantitative theory.
major comments (4)
- [Sec. IX.B, Eq. (117)] The decomposition W₂²(ρ_lin, μ_N) ≃ W²_grav + W²_bias + W²_samp is the load-bearing step, but it is stated without proof. The Wasserstein distance is a metric and does not satisfy additivity of squares over concatenated maps; such additivity would require intermediate optimal maps to compose optimally and their displacement fields to be mutually orthogonal in L². In particular, the step ρ_NL → λ_gal in Eq. (100) is a local rescaling of the density, not a transport map, so W²_bias in Eq. (120) is not the squared Wasserstein distance of any well-defined pair of probability measures. The caveat in Sec. X that the theory is a 'first approximation' does not supply the missing derivation or an error estimate.
- [Sec. V, Eq. (54); Sec. VIII, Eq. (88)] The formula ⟨W₂²⟩ = (1/2π²)∫P(k)dk is derived for continuous normalized measures with small density contrast δ. It is applied in Eq. (88) to an empirical point-process measure μ_N, whose density field consists of delta spikes and whose contrast is never small. The power spectrum P_pp in Eq. (87) contains the white-noise term 1/\bar n, so the unsmoothed integral in Eq. (88) diverges. The cutoff f_WR in Eq. (121) is introduced without derivation, and the smoothing scale R is a free parameter. This undermines the shot-noise term in the abstract and in Eq. (3).
- [Sec. IX.B, Eqs. (120) and (123)] The bias term is defined as ∫ r[ξ_gal−ξ_m] dr, not as a Wasserstein distance. Under the linear bias model Eq. (103), Eq. (123) reduces to b₁²∫rξ_m dr plus shot noise, so the 'gravitational' and 'bias' terms are not separately identifiable; the claimed three-way separation is not realized. The abstract Eq. (3) writes ∫ r ξ_gal dr instead of the subtracted version. The two forms are algebraically equal because ∫ r ξ_m dr = (1/2π²)∫P_m(k)dk, but the paper does not state this identity, and the notation obscures that the first two terms in Eq. (123) collapse into one.
- [Secs. II.B and VII.B] W₂ is defined for probability measures, but the objects in Eq. (73) are not consistently normalized: ρ_lin is a density with total mass M (Sec. II.B), whereas μ_N is normalized by 1/N. If the total masses are unequal, the Wasserstein distance is not defined as written. The paper needs to specify the normalization or equal-mass condition (e.g., matching total mass or using a partial-transport formulation) before Eq. (116) is meaningful.
minor comments (4)
- [Introduction] The outline states 'Section IV computes the Wasserstein distance for Poisson point processes,' but Sec. IV covers empirical measures and finite-window effects; the Poisson/shot-noise computation appears only later, in Secs. VII–VIII. Please correct the cross-reference.
- [Abstract and Eq. (121)] The smoothing scale R appears in the abstract and in Eq. (3) but is not defined until Eq. (121). Also, f_WR(k) is introduced without an explicit definition; please specify the smoothing window before using it.
- [Sec. VI.A, Eq. (66)] The finite-volume expression Eq. (65) is a sum over discrete k, but Eq. (66) replaces it by an integral from k_min to k_max without a derivation of the truncation or the continuum approximation. A brief justification would improve clarity.
- [General] The paper would be substantially strengthened by a numerical check of the final formula, even in one or two dimensions or with N-body simulations, to show the size of the neglected corrections.
Circularity Check
The three-term decomposition is asserted and the bias term is defined as the residual difference, but the gravitational and shot-noise terms have independent derivations; no data fitting occurs.
specific steps
-
self definitional
[Sec. IX.B, Eqs. (117), (120), (123); cf. abstract Eq. (3)]
"Under the small-fluctuation approximation, one may write W 2 2 (ρlin, µN ) ≃ W 2 grav + W 2 bias + W 2 samp. (117) ... The bias term is given by W 2 bias ∼ ∫ ∞ 0 r[ξ gal(r)−ξ m(r)] dr. (120)"
W²_bias is not computed as the Wasserstein distance of the ρ_NL → λ_gal step; it is defined as the difference between galaxy and matter correlation integrals. Since ∫r ξ_m dr = (1/2π²)∫P_m(k) dk (Eqs. 90–92 with ξ~ = P), adding Eq. 120 to Eq. 119 cancels the matter power-spectrum term, making Eq. 123 reduce to ∫r ξ_gal dr + shot noise. The advertised decomposition is therefore assembled by defining the bias term as the residual needed to produce the final form, rather than derived from optimal transport. The consistency with abstract Eq. (3), which keeps both ∫P_m and ∫r ξ_gal, also changes depending on this definition.
full rationale
The paper does not fit any parameter to data and the main gravitational term is a genuine derivation: linearizing the Monge-Ampère equation gives W²≈∫|∇ψ|², and ensemble-averaging gives (1/2π²)∫P(k)dk. The shot-noise term is a standard empirical-measure/Poisson result, quoted from the point-process literature. The load-bearing weakness is Sec. IX.B's additive decomposition W²(ρ_lin,μ_N)≃W²_grav+W²_bias+W²_samp, which is asserted rather than derived; W² is a metric satisfying a triangle inequality, not an additive quadratic functional of concatenated maps, and the bias step λ_gal=¯n[1+b₁δ] is a local rescaling, not an optimal transport map. The bias term is then defined as the difference integral, which makes the final formula consistent by construction. This is a mild definitional forcing and an omitted proof / correctness risk, but not a fitted-input-called-prediction or a self-citation chain. The paper also flags its own limitations ('first approximation based on the small-fluctuation limit and on a stage-by-stage decomposition'). Overall circularity burden is low; score 2.
Axiom & Free-Parameter Ledger
free parameters (1)
- R (smoothing scale)
axioms (4)
- domain assumption Small-fluctuation linearization: ∇²ψ = δ, so W₂² ≈ ∫|∇ψ|² d³x.
- ad hoc to paper Squared Wasserstein distance decomposes additively over the three-stage generative process: W₂²(ρ_lin,μ_N) ≈ W²_grav + W²_bias + W²_samp.
- domain assumption The linearized formula for W² applies to point-process measures (after smoothing).
- ad hoc to paper Shot-noise formula: ⟨W²_samp⟩ = (1/(2π² \bar n))∫|f_WR(k)|² dk.
read the original abstract
Cosmological large-scale structure is usually characterized by statistics such as the power spectrum and correlation function, which describe fluctuation amplitudes but not the spatial redistribution of matter. We formulate structure formation as a transport problem between mass distributions using the Wasserstein distance of optimal transport theory. The generative process from the initial linear density field to an observed galaxy catalog is represented as a hierarchical mapping from a continuous matter field to a galaxy point process. Under the small-fluctuation approximation, we derive an approximate Wasserstein distance that decomposes into three physically distinct contributions: gravitational mass transport, galaxy formation bias, and shot noise from discrete galaxy sampling. The gravitational term is expressed as an integral of the matter power spectrum, whereas the galaxy formation term is given by a transport-weighted integral of the galaxy correlation function. The sampling term reduces to Poisson shot noise. Unlike conventional volume-weighted clustering statistics, the correlation contribution measures cumulative matter displacements and is naturally connected to baryon acoustic oscillation smearing. This framework provides a transport-geometric description of large-scale structure formation and introduces the Wasserstein distance as a statistical link between continuous density fields and observed galaxy catalogs.(abridged)
Forward citations
Cited by 1 Pith paper
-
A Measure-Theoretic Transport Formulation of Galaxy Evolution on the Galaxy Manifold: Geometric Constraints
Galaxy evolution is cast as a geometrically constrained reaction-transport process on probability measures, using Wasserstein distance and CD(K,∞) conditions to enforce energy dissipation and interaction closure.
Reference graph
Works this paper leans on
-
[1]
P. J. E. Peebles,The large-scale structure of the universe (Princeton University Press, Princeton, 1980)
1980
-
[2]
T. T. Takeuchi,Physics of the Formation and Evolution of Galaxies: Multiwavelength Point of View, Springer Se- ries in Astrophysics and Cosmology (Springer Nature, Singapore, 2025)
2025
-
[3]
Y. B. Zel’dovich, Astronomy & Astrophysics5, 84 (1970)
1970
-
[4]
Villani,Optimal Transport: Old and New, Grundlehren der mathematischen Wissenschaften (Springer, Berlin, 2008)
C. Villani,Optimal Transport: Old and New, Grundlehren der mathematischen Wissenschaften (Springer, Berlin, 2008)
2008
- [5]
-
[6]
F. Nikakhtar, R. K. Sheth, B. L´ evy, and R. Mohayaee, Phys. Rev. Lett.129, 251101 (2022), arXiv:2203.01868 [astro-ph.CO]
Pith/arXiv arXiv 2022
-
[7]
F. Nikakhtar, N. Padmanabhan, B. L´ evy, R. K. Sheth, and R. Mohayaee, Phys. Rev. D108, 083534 (2023), arXiv:2307.03671 [astro-ph.CO]
Pith/arXiv arXiv 2023
-
[8]
F. Nikakhtar, R. K. Sheth, N. Padmanabhan, B. L´ evy, and R. Mohayaee, Phys. Rev. D109, 123512 (2024), arXiv:2403.11951 [astro-ph.CO]
Pith/arXiv arXiv 2024
-
[9]
Davis, E
M. Davis, E. J. Groth, and P. J. E. Peebles, Astrophys. J.212, L107 (1977)
1977
-
[10]
Davis and P
M. Davis and P. J. E. Peebles, Astrophys. J.267, 465 (1983)
1983
-
[11]
Benamou and Y
J.-D. Benamou and Y. Brenier, Numerische Mathematik 84, 375 (2000)
2000
-
[12]
Brenier, Communications on Pure and Applied Math- ematics44, 375 (1991)
Y. Brenier, Communications on Pure and Applied Math- ematics44, 375 (1991)
1991
-
[13]
Fournier and A
N. Fournier and A. Guillin, Probability Theory and Re- lated Fields162, 707 (2015)
2015
-
[14]
In measure-theoretic discussions this notation is commonly used, but in the context of this paper it is essentially equivalent to the ensemble average⟨·⟩
HereE[·] denotes the probabilistic expectation. In measure-theoretic discussions this notation is commonly used, but in the context of this paper it is essentially equivalent to the ensemble average⟨·⟩
-
[15]
Boissard and T
E. Boissard and T. L. Gouic, Annales de l’Institut Henri Poincar´ e, Probabilit´ es et Statistiques50, 539 (2014)
2014
-
[16]
D. J. Daley and D. Vere-Jones,An introduction to the theory of point processes. Volume I: Elementary Theory and Methods(Springer, New York, 2003) elementary the- ory and methods
2003
-
[17]
D. J. Daley and D. Vere-Jones,An Introduction to the Theory of Point Processes. Volume II: General Theory and Structure(Springer, New York, 2008)
2008
-
[18]
Last and M
G. Last and M. Penrose,Lectures on the Poisson Process, Institute of Mathematical Statistics Textbooks (Cam- bridge University Press, Cambridge, 2017)
2017
-
[19]
Baddeley, E
A. Baddeley, E. Rubak, and R. Turner,Spatial Point Pat- terns: Methodology and Applications with R, Chapman & Hall/CRC Interdisciplinary Statistics (Chapman and Hall/CRC, 2015)
2015
-
[20]
J. A. Peacock,Cosmological Physics(Cambridge Univer- sity Press, Cambridge, 1999)
1999
-
[21]
H. A. Feldman, N. Kaiser, and J. A. Peacock, Astrophys. J.426, 23 (1994), arXiv:astro-ph/9304022 [astro-ph]
Pith/arXiv arXiv 1994
-
[22]
F. Bernardeau, S. Colombi, E. Gazta˜ naga, and R. Scoc- cimarro, Physics Reports367, 1 (2002), arXiv:astro- ph/0112551 [astro-ph]
arXiv 2002
-
[23]
M. A. Strauss and J. A. Willick, Physics Reports261, 271 (1995), arXiv:astro-ph/9502079 [astro-ph]
Pith/arXiv arXiv 1995
-
[24]
Kaiser, Astrophys
N. Kaiser, Astrophys. J.284, L9 (1984)
1984
-
[25]
J. N. Fry and E. Gaztanaga, Astrophys. J.413, 447 (1993), arXiv:astro-ph/9302009 [astro-ph]
Pith/arXiv arXiv 1993
-
[26]
V. Desjacques, D. Jeong, and F. Schmidt, Physics Re- ports733, 1 (2018), arXiv:1611.09787 [astro-ph.CO]
Pith/arXiv arXiv 2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.