REVIEW 1 major objections 49 references
The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserstein Space
T0 review · 1 major / 0 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read The Wasserstein manifold turns diffusion models into gradient flows of free energy and flow matching into geodesics.
desk verdict The paper frames diffusion as JKO gradient flow and flow matching as Wasserstein geodesics on the same manifold, but the exact match to DDPM-style updates still needs the extra modeling steps the stress-test flags. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The formal Riemannian structure on the Wasserstein space P_2(R^d) induced by the quadratic Wasserstein distance, which supports both the gradient flows of the free energy and the Benamou-Brenier geodesics.
What would settle it
A calculation showing that the JKO step produces update equations different from those in DDPM or that the trajectories optimized by flow matching deviate from the Benamou-Brenier geodesics would falsify the claimed unification.
Extended reading notes
Core claim
The space P_2(R^d) of probability measures with finite second moment carries the quadratic Wasserstein distance that makes it a complete metric space and a formal Riemannian manifold. The gradient flow of the free energy F(rho) = KL(rho || pi) on this manifold is the Fokker-Planck equation, whose implicit-Euler discretization is the JKO scheme; this scheme recovers DDPM, DDIM, NCSN/SMLD, and energy matching as instances of one discretization. The same manifold supplies a second variational principle whose geodesics are the optimal-transport paths given by the Benamou-Brenier formula, and these are precisely the paths that flow matching learns, yielding deterministic straight-line ODEs.
Load-bearing premise
That identifying the JKO discretization with the concrete update rules of DDPM, DDIM, NCSN and energy matching is enough to derive those models exactly from the manifold geometry without further modeling choices.
Editorial extensions
If this is right
- All listed diffusion models become instances of a single JKO discretization on the manifold rather than separate theories.
- Flow-matching generation reduces to integrating a deterministic ODE along a Wasserstein geodesic, requiring fewer steps than stochastic diffusion paths.
- The two families reach identical target measures but solve different problems: an initial-value energy descent versus a boundary-value geodesic connection.
- Placing both families on one manifold makes their exact relationship visible without additional approximations.
Reading between the lines
- The same manifold geometry could be used to construct hybrid samplers that switch between gradient-flow steps and geodesic segments.
- Other generative techniques might be re-derived by identifying the variational principle they implicitly optimize on P_2(R^d).
- Efficiency gains could be obtained by designing discretizations that better respect the Riemannian metric rather than the current ad-hoc choices.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that diffusion models are instances of the JKO discretization of the gradient flow of the free energy F(ρ) = KL(ρ || π) on the Wasserstein space P₂(ℝᵈ), recovering DDPM, DDIM, NCSN/SMLD and Energy Matching as the same scheme; flow matching instead follows the Benamou-Brenier geodesics on the same manifold, so the two families are related exactly as an initial-value gradient-flow problem versus a boundary-value geodesic problem.
Significance. If the claimed identifications are shown to hold without extra modeling choices, the work would supply a single geometric language for relating score-based diffusion and optimal-transport flow matching, highlighting their shared Wasserstein structure and the distinction between free-energy descent and geodesic interpolation.
major comments (1)
- [Abstract] Abstract (paragraph on JKO and diffusion): the statement that the JKO scheme 'recovers DDPM, DDIM, NCSN/SMLD, and Energy Matching' as one scheme is asserted without an explicit derivation of the proximal map or the error incurred when matching it to the closed-form Gaussian transitions and score-matching approximations used in those algorithms. The skeptic note correctly flags that such a match typically requires a Gaussian ansatz or noise-schedule choices not implied by the formal Riemannian structure alone; this identification is load-bearing for the central unification claim.
Simulated Author's Rebuttal
We thank the referee for the thoughtful review and for highlighting the need for greater rigor in the central identification. We address the major comment below and will revise the manuscript accordingly to strengthen the exposition of the JKO-proximal map connection.
read point-by-point responses
-
Referee: [Abstract] Abstract (paragraph on JKO and diffusion): the statement that the JKO scheme 'recovers DDPM, DDIM, NCSN/SMLD, and Energy Matching' as one scheme is asserted without an explicit derivation of the proximal map or the error incurred when matching it to the closed-form Gaussian transitions and score-matching approximations used in those algorithms. The skeptic note correctly flags that such a match typically requires a Gaussian ansatz or noise-schedule choices not implied by the formal Riemannian structure alone; this identification is load-bearing for the central unification claim.
Authors: We agree that the abstract states the recovery concisely and that an explicit derivation of the proximal map (and the approximation error relative to the Gaussian transitions and score-matching objectives) is not supplied there. The manuscript grounds the claim in the fact that the JKO scheme for the KL free energy is the implicit Euler discretization of the Fokker-Planck equation, whose continuous-time limit is known to underlie the forward noising process; the reverse steps then correspond to the proximal operators realized by the cited algorithms. Nevertheless, we accept that the load-bearing identification would be more convincing with a short derivation or reference to the proximal-operator calculation under a Gaussian ansatz, together with a clear statement of the modeling choices (noise schedule, Gaussian assumption) that are external to the pure Wasserstein geometry. We will therefore expand the relevant section (and, if space permits, the abstract) to include this derivation and to delineate precisely where the formal Riemannian structure ends and the algorithmic approximations begin. This revision will not alter the geometric unification but will make the supporting evidence explicit. revision: yes
Circularity Check
No significant circularity; interpretive unification via standard Wasserstein geometry
full rationale
The paper applies established results from optimal transport (Otto's Riemannian structure on P_2, JKO proximal scheme for gradient flows of KL, Benamou-Brenier geodesic formula) to reinterpret diffusion and flow-matching algorithms. The abstract and provided text present this as a geometric correspondence without fitting parameters to data, without renaming empirical patterns as new derivations, and without load-bearing self-citations that reduce the central claim to an unverified loop. The identification of JKO steps with DDPM/DDIM discretizations is offered as an interpretive lens rather than a closed-form prediction derived solely from the manifold structure; no equation in the given material reduces a claimed result to its own inputs by construction. This is the normal case of a self-contained interpretive paper.
Assumptions & free parameters
assumptions (1)
- domain assumption P_2(R^d) carries a formal Riemannian structure whose gradient flows and geodesics are given by Otto calculus and the Benamou-Brenier formula
Cite this review
Pith. "Pith review of The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserstein Space." pith.science (2026). https://pith.science/paper/IYS5DXR3
@misc{pith2026260624157,
author = {Pith},
title = {Pith review of: The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserstein Space},
year = {2026},
howpublished = {\url{https://pith.science/paper/IYS5DXR3}},
note = {Machine review of arXiv:2606.24157}
}
abstract
The space $\mathcal{P}_2(\mathbb{R}^d$) of probability measures with finite second moment carries a natural geometry: the quadratic Wasserstein distance W_2 makes it a complete metric space and, following Otto, a (formal) Riemannian manifold whose geodesics are the optimal-transport interpolations. On this manifold, the gradient flow of the free energy F(rho) = KL(rho || \pi) is exactly the Fokker-Planck equation, and its implicit-Euler discretization is the JKO scheme. This is the geometry underlying diffusion models: the forward process descends the free energy, and each denoising step realizes one JKO step, which recovers DDPM, DDIM, NCSN/SMLD, and Energy Matching; this is one scheme, not separate theories. The same manifold supports a second variational principle. Its geodesics - the minimum-action curves of the Benamou-Brenier formula - are precisely the optimal-transport paths that Flow Matching learns. Fixing both endpoints and following the geodesic, generation becomes a deterministic ODE along a straight line, hence far fewer sampling steps. Placing both families of models on one manifold makes their relationship exact: diffusion follows a free-energy gradient flow, an initial-value problem; optimal-transport Flow Matching follows a Wasserstein geodesic, a boundary-value problem. The two reach the same endpoints along different paths.
Reference graph
Works this paper leans on
-
[1]
It is the equation of motion in Wasserstein space
Thecontinuity equation∂ tρ+∇ ·(ρv) = 0 describes how probability mass flows. It is the equation of motion in Wasserstein space. It does not require an energy functional; it only requires a velocity field
-
[2]
TheBenamou–Brenier formulareveals thatW 2 = geodesic distance in this Riemannian structure, with kinetic energy R |v|2ρ dxas the metric
-
[3]
˙ρ=−gradW F(ρ)
TheFokker–Planck equationis the Wasserstein gradient flow of the free energy F(ρ) = KL(ρ∥π): “ ˙ρ=−gradW F(ρ)” This is the infinite-dimensional analog of ˙x=−∇f(x). In this special case, the velocity is not arbitrary: it is determined by the free energy
-
[4]
TheJKO schemediscretizes this gradient flow: ρk+1 = arg min ρ F(ρ) + W 2 2 (ρ, ρk) 2τ Asτ→0, it recovers the continuous-time Fokker–Planck equation. One-sentence summary: Wasserstein distance endows probability space with geometry; the continuity equation describes mass-conserving flow; the Fokker–Planck equation is the free-energy-driven special case; Fl...
-
[5]
(1781).M´ emoire sur la th´ eorie des d´ eblais et des remblais.Histoire de l’Acad´ emie Royale des Sciences de Paris, 666–704
Monge, G. (1781).M´ emoire sur la th´ eorie des d´ eblais et des remblais.Histoire de l’Acad´ emie Royale des Sciences de Paris, 666–704. 62
-
[6]
Kantorovich, L. V. (1942).On the translocation of masses.Dokl. Akad. Nauk SSSR37, 199–201
1942
-
[7]
(1991).Polar factorization and monotone rearrangement of vector-valued functions.Comm
Brenier, Y. (1991).Polar factorization and monotone rearrangement of vector-valued functions.Comm. Pure Appl. Math.44(4), 375–417
1991
-
[8]
Wasserstein distance
Vaserstein, L. N. (1969).Markov processes over denumerable products of spaces, describing large systems of automata.Problemy Peredachi Informatsii5(3), 64–72. [The source of the name “Wasserstein distance.”]
1969
Show all 49 references
-
[9]
McCann, R. J. (1997).A convexity principle for interacting gases.Adv. Math.128(1), 153–179. [Displacement interpolation.]
1997
-
[10]
& Brenier, Y
Benamou, J.-D. & Brenier, Y. (2000).A computational fluid mechanics solution to the Monge–Kantorovich mass transfer problem.Numer. Math.84(3), 375–393
2000
-
[11]
(1998).The variational formulation of the Fokker–Planck equation.SIAM J
Jordan, R., Kinderlehrer, D., & Otto, F. (1998).The variational formulation of the Fokker–Planck equation.SIAM J. Math. Anal.29(1), 1–17
1998
-
[12]
(2001).The geometry of dissipative evolution equations: the porous medium equation.Comm
Otto, F. (2001).The geometry of dissipative evolution equations: the porous medium equation.Comm. PDE26(1-2), 101–174
2001
-
[13]
(2008).Gradient Flows in Metric Spaces and in the Space of Probability Measures.Birkh¨ auser
Ambrosio, L., Gigli, N., & Savar´ e, G. (2008).Gradient Flows in Metric Spaces and in the Space of Probability Measures.Birkh¨ auser. [The definitive reference.]
2008
-
[14]
(2003).Topics in Optimal Transportation.AMS
Villani, C. (2003).Topics in Optimal Transportation.AMS. [Excellent introduction.]
2003
-
[15]
(2015).Optimal Transport for Applied Mathematicians.Birkh¨ auser
Santambrogio, F. (2015).Optimal Transport for Applied Mathematicians.Birkh¨ auser. [Very readable.]
2015
-
[16]
& Cuturi, M
Peyr´ e, G. & Cuturi, M. (2019).Computational Optimal Transport.Found. Trends ML 11(5-6), 355–607. [Computational perspective.]
2019
-
[17]
& Santambrogio, F
Lavenant, H. & Santambrogio, F. (2022).The flow map of the Fokker–Planck equation does not provide optimal transport.Appl. Math. Lett.133, 108225
2022
-
[18]
(2021).Large- scale Wasserstein gradient flows.NeurIPS
Mokrov, P., Korotin, A., Li, L., Genevay, A., Solomon, J., & Burnaev, E. (2021).Large- scale Wasserstein gradient flows.NeurIPS
2021
-
[19]
(2023).Normalizing flow neural networks by JKO scheme
Xu, C., Cheng, X., & Xie, Y. (2023).Normalizing flow neural networks by JKO scheme. NeurIPS. [JKO-iFlow.] Generative models
2023
-
[20]
(2005).Estimation of non-normalized statistical models by score matching
Hyv¨ arinen, A. (2005).Estimation of non-normalized statistical models by score matching. J. Mach. Learn. Res.6, 695–709
2005
-
[21]
(2011).A connection between score matching and denoising autoencoders
Vincent, P. (2011).A connection between score matching and denoising autoencoders. Neural Comput.23(7), 1661–1674
2011
-
[22]
& Teh, Y
Welling, M. & Teh, Y. W. (2011).Bayesian learning via stochastic gradient Langevin dynamics.ICML, 681–688
2011
-
[23]
& Ermon, S
Song, Y. & Ermon, S. (2019).Generative modeling by estimating gradients of the data distribution.NeurIPS. [NCSN/SMLD.]
2019
-
[24]
(2020).Denoising diffusion probabilistic models.NeurIPS
Ho, J., Jain, A., & Abbeel, P. (2020).Denoising diffusion probabilistic models.NeurIPS. [DDPM.] 63
2020
-
[25]
(2021).Denoising diffusion implicit models.ICLR
Song, J., Meng, C., & Ermon, S. (2021).Denoising diffusion implicit models.ICLR. [DDIM.]
2021
-
[26]
P., Kumar, A., Ermon, S., & Poole, B
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., & Poole, B. (2021). Score-based generative modeling through stochastic differential equations.ICLR
2021
-
[27]
(2022).Elucidating the design space of diffusion-based generative models.NeurIPS
Karras, T., Aittala, M., Aila, T., & Laine, S. (2022).Elucidating the design space of diffusion-based generative models.NeurIPS. [EDM.]
2022
-
[28]
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., & Le, M. (2023).Flow matching for generative modeling.ICLR
2023
-
[29]
(2025).Energy Matching: unifying flow matching and energy-based models for generative modeling.arXiv:2504.10612 (NeurIPS 2025)
Balcerak, M., et al. (2025).Energy Matching: unifying flow matching and energy-based models for generative modeling.arXiv:2504.10612 (NeurIPS 2025)
2025
-
[30]
(2025).Equilibrium Matching: generative modeling with implicit energy- based models.arXiv:2510.02300
Wang, R., et al. (2025).Equilibrium Matching: generative modeling with implicit energy- based models.arXiv:2510.02300
2025
-
[31]
S., Boffi, N
Albergo, M. S., Boffi, N. M., & Vanden-Eijnden, E. (2025).Stochastic interpolants: a unifying framework for flows and diffusions.J. Mach. Learn. Res.26(209), 1–80
2025
-
[32]
(2024).Flow Matching Guide and Code.arXiv:2412.06264
Lipman, Y., et al. (2024).Flow Matching Guide and Code.arXiv:2412.06264
2024 arXiv
-
[33]
B., McCann, M
Vuong, A. B., McCann, M. T., Santos, J. E., & Lin, Y. T. (2025).Are we really learning the score function? Reinterpreting diffusion models through Wasserstein gradient flow matching.CIKM. (arXiv:2509.00336)
2025
-
[34]
& Erives, E
Holderrieth, P. & Erives, E. (2025).An Introduction to Flow Matching and Diffusion Models.MIT 6.S184 lecture notes. Classical foundations
2025
-
[35]
Fisher, R. A. (1925).Theory of statistical estimation.Proc. Cambridge Philos. Soc.22, 700–725
1925
-
[36]
(1908).Sur la th´ eorie du mouvement brownien.C
Langevin, P. (1908).Sur la th´ eorie du mouvement brownien.C. R. Acad. Sci. Paris146, 530–533
1908
-
[37]
Uhlenbeck, G. E. & Ornstein, L. S. (1930).On the theory of the Brownian motion.Phys. Rev.36(5), 823–841
1930
-
[38]
Kolmogorov, A. N. (1931). ¨Uber die analytischen Methoden in der Wahrscheinlichkeit- srechnung.Math. Ann.104, 415–458. [Kolmogorov forward/backward equations.]
1931
-
[39]
(1951).On stochastic differential equations.Mem
Itˆ o, K. (1951).On stochastic differential equations.Mem. Amer. Math. Soc.4, 1–51
1951
-
[40]
& Leibler, R
Kullback, S. & Leibler, R. A. (1951).On information and sufficiency.Ann. Math. Statist.22(1), 79–86
1951
-
[41]
McKean, H. P. (1966).A class of Markov processes associated with nonlinear parabolic equations.Proc. Natl. Acad. Sci. USA56(6), 1907–1911. [McKean–Vlasov.]
1966
-
[42]
assigning size to sets
Anderson, B. D. O. (1982).Reverse-time diffusion equation models.Stochastic Process. Appl.12(3), 313–326. 64 A Measures and Couplings Before introducing the Wasserstein distance, we need to clarify two foundational concepts: measuresandcouplings. If you are already familiar wi...
1982
-
[43]
the densityρ(x) being large at some point
The measure onR 2 uniformly distributed on the segment{(x,0) :x∈[0,1]}: not abso- lutely continuous× (The segment has Lebesgue measure zero inR 2, yet this measure places all its mass there) 6.µwith densityρ(x) = 1 |x|1/21[−1,1](x) (density tends to∞atx= 0): absolutely continu...
-
[44]
transposing
=−1. ω1 ̸=ω 2!—the samev, after changing basis, “transposing” yields a different covector. So “transpose” is not a map fromVtoV ∗—it is a map from “coordinates ofVin some basis” to “coordinates ofV ∗ in the dual basis.” The moment you say “take coordinates,” you have already m...
-
[45]
The direction of gradfis the direction that maximizesd f(v) (finding the steepest direction in theg-unit ball)
-
[46]
rolling downhill
The direction of gradfis theg-normal direction to the level set{f=c} 3.∥gradf∥ g = max∥v∥g=1 d f(v) Intuition Preview: why are the Wasserstein gradient and theL 2 gradient different? For the same functionalF(ρ) and the same functional derivative δF δρ , but: •L 2 metric⟨u, v⟩ ...
1925
-
[47]
gradient direction
This is what flow matching learns. •Drift:f(x, t), the deterministic part of the SDEdX t =f dt+g dB t. This is given (determined by the forward process). For the SDEdX t =f(X t, t)dt+g(t)dB t, the velocity field of the correspondingproba- bility flow ODEis: vt(x) =f(x, t)− 1 2...
1982
-
[48]
It is the standard tool for proving stability
It is monotonically decreasing along trajectories of the system: d dt V(x(t))≤0 Intuition:A Lyapunov function is like “energy”—if you can find a quantity that only decreases during the system’s evolution, then the system must evolve toward the point where this quantity is zero...
-
[49]
gradient equals zero
and attains its minimum value of zero at the stationary distributionρ=π. H.8 KKT conditions TheKKT conditions(Karush-Kuhn-Tucker conditions) are necessary optimality conditions for constrained optimization problems, generalizing the “gradient equals zero” condition from uncons...
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.