REVIEW 3 major objections 6 minor 2 cited by
Towards Understanding Gradient Dynamics of the Sliced-Wasserstein Distance via Critical Point Analysis
T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper proves that in dimension two, any measure carrying a positive mass on a segment is unstable for the squared sliced-Wasserstein objective, so stable critical points cannot be segment-shaped.
desk verdict Genuinely new segment-instability result for sliced Wasserstein, but the abstract overstates it: the proven perturbation is not the Lagrangian one the paper's own framework needs. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the barycentric vector field $v_\mu(x)=\tfrac1d x-\int_{S^{d-1}}\bar\gamma_\theta(\langle x|\theta\rangle)\,\theta\,d\theta$, where $\bar\gamma_\theta$ is the conditional mean of the optimal 1-D transport plan between $P_\theta\#\mu$ and $P_\theta\#\rho$; criticality is exactly $v_\mu=0$ $\mu$-a.e. For the stability theorem, the carrying computation is the explicit quantile function of the perturbed segment under every projection, combined with the symmetry $x\mapsto 1-x$ and a lower bound on the difference quotient $G_\theta$ of the transported target; these produce the super-quadratic negative term. The same quantile machinery, via the identity $W_2^2=\|F^{-1}_\mu-F^{-1}_\rho\|_{L^2}^2$, is what makes the sliced objective analytically tractable.
What would settle it
In dimension two, choose any target $\rho$ with uniformly bounded one-dimensional projection densities and any segment-supported measure $\mu$, and compute $R(t)=(\mathrm{SW}_2^2(\mu_t,\rho)-\mathrm{SW}_2^2(\mu,\rho))/t^2$ for $\mu_t=\tfrac12(\tau_{-tn}\#\mu+\tau_{tn}\#\mu)$ with $n$ orthogonal to the segment. The paper's Proposition 5.2 entails $R(t)\to-\infty$ as $t\to0$; observing a finite limit, or any well-defined second derivative at $t=0$, would refute the instability claim.
Extended reading notes
Core claim
The central claim is that the landscape of $F(\cdot)=\tfrac12\mathrm{SW}_2^2(\cdot,\rho)$ has a two-sided character: non-global Lagrangian critical points on lower-dimensional sets exist, yet segment-supported ones are unstable in a very strong sense. With the target $\rho$ absolutely continuous and its one-dimensional projections having density bounded by $b$, and with $\mu$ carrying $a\mathcal H^1|_S$ on a segment $S$, the perturbation $\mu_t=\tfrac12(\tau_{-tn}\#\mu+\tau_{tn}\#\mu)$, $n\perp S$, satisfies $F(\mu_t)\le F(\mu)-Ct^2$ on a neighborhood of $t=0$ for every $C>0$ (Proposition 5.2). This is stronger than having a negative second derivative: the function $t\mapsto F(\mu_t)$ is not twice differentiable at $0$, and the decrease rate is arbitrarily steep as $t\to0$. Separately, the paper proves that the variational notion of a Lagrangian critical point is equivalent, under compact-support and atomless assumptions, to the barycentric equation $v_\mu=0$, and that weak limits of barycentric critical points are again barycentric critical.
Load-bearing premise
The proof needs the piece of the target that receives the segment mass to have bounded density in the direction perpendicular to the segment, and it uses the segment being exactly uniform to get closed-form quantiles; without those, the negative quadratic estimate can fail.
Editorial extensions
If this is right
- Every weak limit of discrete barycentric critical points (for example, limits of particle configurations produced by gradient descent) is again a barycentric Lagrangian critical point of the continuous functional, so the discrete-to-continuous limit does not create new critical points.
- There are explicit critical points distinct from the target: a uniform segment in the plane for the sliced-uniform disk target, and a Gaussian line for the standard Gaussian in higher dimensions.
- In dimension two, a stable critical point cannot have any segment piece, regardless of what the rest of the measure does.
- For the particle formulation, step sizes below $Nd$ make the objective strictly decrease, and with bounded projection densities the iterates stay away from particle collisions, so convergent subsequences reach critical points.
Reading between the lines
- Beyond the paper: the arbitrarily steep decrease suggests that sliced-Wasserstein optimization may escape segment traps through non-smooth kink directions, so the practical advantage could come from non-differentiability rather than from curvature; this could be tested by comparing first-order algorithms that deliberately split symmetric particle pairs.
- Beyond the paper: the proof's dependence on the segment being exactly uniform suggests the instability is a resonance between a perfectly flat lower-dimensional piece and a genuinely full-dimensional target; segment-like pieces with curved or non-uniform density profiles might behave differently, which the theorem does not cover.
- Beyond the paper: a natural extension, left open by the authors, is to test hyperplane-supported measures in $d>2$ with multi-directional two-sided perturbations; if the same unbounded negative quadratic appears, low-dimensional collapse should be transient in all dimensions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the squared Sliced-Wasserstein distance SW2^2(·, ρ) as an objective functional over probability measures. It introduces Eulerian, Wasserstein, and Lagrangian notions of critical points, proves a barycentric characterization of Lagrangian critical points (Definitions 4.1 and 4.2, Proposition 4.7), shows that weak limits of discrete barycentric critical points remain barycentric critical (Theorem 4.5), constructs explicit lower-dimensional critical points including a segment for a sliced-uniform target and a Gaussian example (Proposition 5.1), and claims in dimension 2 that stable critical points cannot concentrate on segments (Proposition 5.2). The instability is proven for a specific two-shifted-copy perturbation, not for a Lagrangian pushforward perturbation. Numerical experiments illustrate both the instability and the behavior of gradient descent.
Significance. If the main claim were established within the Lagrangian framework that the paper itself identifies as the relevant one for gradient dynamics, this would be a meaningful step toward understanding the non-convex landscape of SW optimization. The paper contains several useful building blocks: explicit gradient formulas for the semi-discrete SW objective (Proposition 3.1), a descent lemma with a valid step-size range (Proposition 3.2), a non-collapse result for gradient descent (Proposition 3.3), a limit theorem for critical points (Theorem 4.5), explicit critical-point constructions (Proposition 5.1), and a detailed quantitative instability computation (Proposition 5.2). The appendix computations are extensive and the code for the experiments is provided. The proof of Proposition 5.2 appears internally sound under the stated bounded-projection-density hypothesis; in particular, the density bound on the transported part of the segment follows automatically from ρ = (1−λ)ρ_{θ,0} + λρ_{θ,1}.
major comments (3)
- [Abstract and Section 1 vs. Section 5.2 and Conclusion] The abstract and introduction state that stable critical points of SW cannot concentrate on segments, but the only supporting result, Proposition 5.2, uses the perturbation μ_t = (1/2)(τ_{−tn}#μ + τ_{tn}#μ) in Eq. (22), which is a mixture of two translated copies of μ and not a pushforward of the form (Id + tξ)#μ. The paper itself acknowledges this in the remark after Proposition 5.2 and in the Conclusion, yet the unqualified claim appears in the abstract and introduction. Either the authors should prove a Lagrangian-instability analogue, or the advertised claim should be qualified to explicitly state that it concerns the specific mixture perturbation and not the Lagrangian framework.
- [Appendix B.10, Eq. (198) and Eq. (246)] The proof of Proposition 5.2 is labeled 'Sketch of proof' and leaves several load-bearing steps underdeveloped. The decomposition γ̂_θ = (1−λ)γ̂_{θ,0} + λγ̂_{θ,1} and the claim that the disintegrated plans are optimal between their margins are asserted with only a citation; the 'equality at t=0' in Eq. (198) is not immediate because it depends on this optimality. The bound leading to the constant C0 in Eq. (246) is stated as 'there exists some constant C0' without a fully explicit derivation or a discussion of uniformity in θ1 and t. Since Proposition 5.2 is the paper's main instability theorem, a complete proof should be provided in place of the sketch.
- [Section 5.2 remark and Section 6] The paper suggests that a vector field ξ alternating rapidly between n and −n will approximately realize the mixture perturbation of Proposition 5.2, and that SW2^2((Id + tξ)#μ, ρ) will therefore have a local maximum at t=0. This claim is supported only by numerical experiments, not by a theorem. Without a rigorous link between the mixture perturbation and the Lagrangian perturbation family, the paper's conclusion about gradient dynamics — which is the stated motivation of the work — remains unsupported.
minor comments (6)
- [Eq. (7)] The power cell intervals in Eq. (7) are written as [i/N, (i+1)/N], which is shifted by one index; they should presumably be [(i−1)/N, i/N] as used in the proofs in Appendix B.1.
- [Abstract] The abstract contains a typo: 'Sliced-Wassertein' should be 'Sliced-Wasserstein'.
- [Section 4] The sentence 'The goal of this this section' contains a duplicated 'this'.
- [Theorem 4.5] The phrase 'an atomles measure' should read 'an atomless measure'.
- [Section 4.2] The word 'Langrangian' appears in the sentence 'the two definitions of Langrangian critical points'; it should be 'Lagrangian'.
- [Appendix B.10] The main text refers to the proof of Proposition 5.2 as deferred to the appendix, but the appendix labels it 'Sketch of proof'. This inconsistency should be resolved, either by expanding the sketch into a complete proof or by explicitly referring to it as a sketch in the main text.
Circularity Check
No significant circularity: the central critical-point characterization and instability bound are derived from optimal transport first principles, with only minor non-load-bearing self-citations.
full rationale
The paper's central derivation chain is self-contained rather than circular. The Lagrangian and barycentric critical-point notions (Definitions 4.1 and 4.2) are not assumed to be equivalent; Proposition 4.4 proves the equivalence for discrete measures, Proposition 4.7 derives the directional derivative formula from convexity and optimal transport duality, and Corollary 4.8 establishes the equivalence under compact support/atomless assumptions. These results do not presuppose the conclusions of the paper. The lower-dimensional critical points in Proposition 5.1 are verified by explicit quantile and transport-map computations (e.g., T_theta(x) = pi/(4|c_theta|) x for the sliced-uniform example), not by invoking the target instability result. Proposition 5.2's instability inequality is obtained from an explicit computation of the quantile function of the split segment measure and a lower bound on the symmetric quantile difference of the transported target density; the uniform bounded-density hypothesis is used genuinely and is not a renamed version of the conclusion. The Gaussian example's constant alpha_d is a designed value chosen to make the barycentric equation vanish, not a fitted parameter subsequently called a prediction. The paper does contain self-citations, notably to M\'erigot et al. (2021) for semiconcavity of the one-dimensional Wasserstein loss and to Sarrazin (2022) for background on Lagrangian critical points, but these are standard external ingredients or context and are not load-bearing for the main claims. The acknowledged mismatch between Proposition 5.2's perturbation (a mixture of two translated copies) and the Lagrangian framework (pushforwards by Id+t xi) is a scope limitation of the stated instability result, explicitly flagged by the authors in the remark after Proposition 5.2 and in the conclusion; it is a correctness/generalization gap, not a circularity. The abstract's unqualified wording about stable critical points is stronger than what is proven, but this overstatement does not make the derivation equivalent to its inputs. Overall, the derivation is self-contained against external benchmarks, and the only reason the score is not 0 is the presence of minor self-citations that do not affect the validity of the argument.
Assumptions & free parameters
free parameters (1)
- α_d =
d ∫_{S^{d-1}} |⟨θ,e1⟩| dθ
assumptions (3)
- domain assumption The target ρ has bounded one-dimensional projection densities in Proposition 5.2.
- ad hoc to paper The perturbation µ_t concentrates the two shifted copies of the segment with equal weight.
- domain assumption The measure μ contains a positive fraction λ = 2a of a uniform segment S.
Cite this review
Pith. "Pith review of Towards Understanding Gradient Dynamics of the Sliced-Wasserstein Distance via Critical Point Analysis." pith.science (2026). https://pith.science/paper/NW5SQEBA
@misc{pith2026250206525,
author = {Pith},
title = {Pith review of: Towards Understanding Gradient Dynamics of the Sliced-Wasserstein Distance via Critical Point Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/NW5SQEBA}},
note = {Machine review of arXiv:2502.06525}
}
read the original abstract
In this paper, we investigate the properties of the Sliced Wasserstein Distance (SW) when employed as an objective functional. The SW metric has gained significant interest in the optimal transport and machine learning literature, due to its ability to capture intricate geometric properties of probability distributions while remaining computationally tractable, making it a valuable tool for various applications, including generative modeling and domain adaptation. Our study aims to provide a rigorous analysis of the critical points arising from the optimization of the SW objective. By computing explicit perturbations, we establish that stable critical points of SW cannot concentrate on segments. This stability analysis is crucial for understanding the behaviour of optimization algorithms for models trained using the SW objective. Furthermore, we investigate the properties of the SW objective, shedding light on the existence and convergence behavior of critical points. We illustrate our theoretical results through numerical experiments.
Figures
Forward citations
Cited by 2 Pith papers
-
Convergence Rates for Distribution Matching with Sliced Optimal Transport
For Gaussian distributions, slice-matching to an isotropic target with decaying step sizes converges at rate O(k^{-(2α-1)}) in expectation.
-
Convergence of empirical subgradients for optimal transport-based objectives
Under smooth unit costs and models, empirical subdifferentials of parameterized transport objectives converge graphically almost surely to the population subdifferential, so subgradient methods approach population cri...
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Gradient flows: in metric spaces and in the space of probability measures
Ambrosio, L., Gigli, N., and Savar \'e , G. Gradient flows: in metric spaces and in the space of probability measures. Springer Science & Business Media, 2005
2005
-
[3]
W asserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L. W asserstein generative adversarial networks. In Precup, D. and Teh, Y. W. (eds.), Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pp.\ 214--223. PMLR, 06--11 Aug 2017
work page 2017
-
[4]
A., Pantazis, Y., and Rey-Bellet, L
Birrell, J., Dupuis, P., Katsoulakis, M. A., Pantazis, Y., and Rey-Bellet, L. (f, gamma)-divergences: Interpolating between f-divergences and integral probability metrics. Journal of machine learning research, 23 0 (39): 0 1--70, 2022
work page 2022
-
[5]
M., Kucukelbir, A., and McAuliffe, J
Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. Variational Inference: A Review for Statisticians . Journal of the American statistical Association, 112 0 (518): 0 859--877, 2017
work page 2017
-
[6]
Bond-Taylor, S., Leach, A., Long, Y., and Willcocks, C. G. Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models . IEEE transactions on pattern analysis and machine intelligence, 44 0 (11): 0 7327--7347, 2021
work page 2021
-
[7]
Sliced and radon wasserstein barycenters of measures
Bonneel, N., Rabin, J., Peyr \'e , G., and Pfister, H. Sliced and radon wasserstein barycenters of measures. Journal of Mathematical Imaging and Vision, 51: 0 22--45, 2015
work page 2015
-
[8]
Unidimensional and evolution methods for optimal transportation
Bonnotte, N. Unidimensional and evolution methods for optimal transportation. PhD thesis, Universit \'e Paris Sud-Paris XI; Scuola normale superiore (Pise, Italie), 2013
work page 2013
Show all 49 references
-
[9]
P., Kok, P
Bourne, D. P., Kok, P. J., Roper, S. M., and Spanjer, W. D. Laguerre tessellations and polycrystalline microstructures: a fast algorithm for generating grains of given volumes. Philosophical Magazine, 100 0 (21): 0 2677--2707, 2020
2020
-
[10]
EigenVI: score-based variational inference with orthogonal function expansions
Cai, D., Modi, C., Margossian, C., Gower, R., Blei, D., and Saul, L. EigenVI: score-based variational inference with orthogonal function expansions . Advances in Neural Information Processing Systems, 37: 0 132691--132721, 2024 a
2024
-
[11]
C., Gower, R
Cai, D., Modi, C., Pillaud-Vivien, L., Margossian, C. C., Gower, R. M., Blei, D. M., and Saul, L. K. Batch and match: black-box variational inference with a score-based divergence. In International Conference on Machine Learning, 2024 b
2024
-
[12]
and Bach, F
Chizat, L. and Bach, F. On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport . Advances in neural information processing systems, 31, 2018
2018
-
[13]
Scalable wasserstein gradient flow for generative modeling through unbalanced optimal transport
Choi, J., Choi, J., and Kang, M. Scalable wasserstein gradient flow for generative modeling through unbalanced optimal transport. In International Conference on Machine Learning, 2024
2024
-
[14]
and Santambrogio, F
Cozzi, G. and Santambrogio, F. Long-time asymptotics of the sliced-wasserstein flow. SIAM J. Im. Sciences, 2024. URL http://cvgmt.sns.it/paper/6495/. cvgmt preprint
2024
-
[15]
and Seljak, U
Dai, B. and Seljak, U. Sliced iterative normalizing flows. Proceedings of Machine Learning Research, 2021
2021
-
[16]
Deshpande, I., Hu, Y.-T., Sun, R., Pyrros, A., Siddiqui, N., Koyejo, S., Zhao, Z., Forsyth, D., and Schwing, A. G. Max-sliced wasserstein distance and its use for gans. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 10648--10656, 2019
2019
-
[17]
Nonparametric generative modeling with conditional sliced-wasserstein flows
Du, C., Li, T., Pang, T., Yan, S., and Lin, M. Nonparametric generative modeling with conditional sliced-wasserstein flows. In International Conference on Machine Learning (ICML), 2023
2023
-
[18]
Variational wasserstein gradient flow
Fan, J., Zhang, Q., Taghvaei, A., and Chen, Y. Variational wasserstein gradient flow. In International Conference on Machine Learning, 2022
2022
-
[19]
Measure transport with kernel stein discrepancy
Fisher, M., Nolan, T., Graham, M., Prangle, D., and Oates, C. Measure transport with kernel stein discrepancy. In International Conference on Artificial Intelligence and Statistics, pp.\ 1054--1062. PMLR, 2021
2021
-
[20]
Learning generative models with sinkhorn divergences
Genevay, A., Peyr \'e , G., and Cuturi, M. Learning generative models with sinkhorn divergences. In International Conference on Artificial Intelligence and Statistics, pp.\ 1608--1617. PMLR, 2018
2018
-
[21]
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. Communications of the ACM, 63 0 (11): 0 139--144, 2020
2020
-
[22]
Generative sliced MMD flows with R iesz kernels
Hertrich, J., Wald, C., Altekr \"u ger, F., and Hagemann, P. Generative sliced MMD flows with R iesz kernels. In International Conference of Learning Representations, 2024
2024
-
[23]
E., Martin, C
Kolouri, S., Pope, P. E., Martin, C. E., and Rohde, G. K. Sliced wasserstein auto-encoders. In International Conference on Learning Representations, 2018
2018
-
[24]
Kernel stein discrepancy descent
Korba, A., Aubin-Frankowski, P.-C., Majewski, S., and Ablin, P. Kernel stein discrepancy descent. In International Conference on Machine Learning, pp.\ 5719--5730. PMLR, 2021
2021
-
[25]
Mmd gan: Towards deeper understanding of moment matching network
Li, C.-L., Chang, W.-C., Cheng, Y., Yang, Y., and P \'o czos, B. Mmd gan: Towards deeper understanding of moment matching network. Advances in neural information processing systems, 30, 2017
2017
-
[26]
and Moosmueller, C
Li, S. and Moosmueller, C. Measure transfer via stochastic slicing and matching, 2025. URL https://arxiv.org/abs/2307.05705
2025
-
[27]
Sliced-wasserstein flows: Nonparametric generative modeling via optimal transport and diffusions
Liutkus, A., Simsekli, U., Majewski, S., Durmus, A., and St \"o ter, F.-R. Sliced-wasserstein flows: Nonparametric generative modeling via optimal transport and diffusions. In International Conference on Machine Learning, pp.\ 4104--4113. PMLR, 2019
2019
-
[28]
Fast optimal transport through sliced generalized wasserstein geodesics
Mahey, G., Chapel, L., Gasso, G., Bonet, C., and Courty, N. Fast optimal transport through sliced generalized wasserstein geodesics. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[29]
Minimax confidence intervals for the sliced wasserstein distance
Manole, T., Balakrishnan, S., and Wasserman, L. Minimax confidence intervals for the sliced wasserstein distance. Electronic Journal of Statistics, 16 0 (1): 0 2252--2345, 2022
2022
-
[30]
A Mean Field View of the Landscape of Two-Layer Neural Networks
Mei, S., Montanari, A., and Nguyen, P.-M. A Mean Field View of the Landscape of Two-Layer Neural Networks . Proceedings of the National Academy of Sciences, 115 0 (33): 0 E7665--E7671, 2018
2018
-
[31]
Non-asymptotic convergence bounds for wasserstein approximation using point clouds
M \'e rigot, Q., Santambrogio, F., and Sarrazin, C. Non-asymptotic convergence bounds for wasserstein approximation using point clouds. In Neural Information Processing Systems, 2021. URL https://api.semanticscholar.org/CorpusID:235436063
2021
-
[32]
Statistical and topological properties of sliced probability divergences
Nadjahi, K., Durmus, A., Chizat, L., Kolouri, S., Shahrampour, S., and Simsekli, U. Statistical and topological properties of sliced probability divergences. Advances in Neural Information Processing Systems, 33: 0 20802--20812, 2020
2020
-
[33]
and Ho, N
Nguyen, K. and Ho, N. Energy-based sliced wasserstein distance. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[34]
Statistical, robustness, and computational guarantees for sliced wasserstein distances
Nietert, S., Goldfeld, Z., Sadhu, R., and Kato, K. Statistical, robustness, and computational guarantees for sliced wasserstein distances. Advances in Neural Information Processing Systems, 35: 0 28179--28193, 2022
2022
-
[35]
First-order methods almost always avoid saddle points: The case of vanishing step-sizes
Panageas, I., Piliouras, G., and Wang, X. First-order methods almost always avoid saddle points: The case of vanishing step-sizes. In Wallach, H., Larochelle, H., Beygelzimer, A., d Alch\' e -Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing S...
2019
-
[36]
Computational optimal transport: With applications to data science
Peyr \'e , G., Cuturi, M., et al. Computational optimal transport: With applications to data science. Foundations and Trends in Machine Learning , 11 0 (5-6): 0 355--607, 2019
2019
-
[37]
On the sequential convergence of Lloyd's algorithms
Portales, L., Cazelles, E., and Pauwels, E. On the sequential convergence of Lloyd's algorithms . working paper or preprint, May 2024. URL https://hal.science/hal-04593982
2024
-
[38]
Wasserstein barycenter and its application to texture mixing
Rabin, J., Peyr \'e , G., Delon, J., and Bernot, M. Wasserstein barycenter and its application to texture mixing. In Scale Space and Variational Methods in Computer Vision: Third International Conference, SSVM 2011, Ein-Gedi, Israel, May 29--June 2, 2011, Revised Selected Pape...
2011
-
[39]
Rachev, S. T. and R \"u schendorf, L. Mass transportation problems: Volume I: theory, volume 1. Springer Science & Business Media, 1998
1998
-
[40]
Orthogonal estimation of wasserstein distances
Rowland, M., Hron, J., Tang, Y., Choromanski, K., Sarlos, T., and Weller, A. Orthogonal estimation of wasserstein distances. In The 22nd International Conference on Artificial Intelligence and Statistics, pp.\ 186--195. PMLR, 2019
2019
-
[41]
The Wasserstein Proximal Gradient Algorithm
Salim, A., Korba, A., and Luise, G. The Wasserstein Proximal Gradient Algorithm . Advances in Neural Information Processing Systems, 33: 0 12356--12366, 2020
2020
-
[42]
Optimal transport for applied mathematicians
Santambrogio, F. Optimal transport for applied mathematicians. Birk \"a user, NY , 55 0 (58-63): 0 94, 2015
2015
-
[43]
Lagrangian discretization of variational problems in Wasserstein spaces
Sarrazin, C. Lagrangian discretization of variational problems in Wasserstein spaces . Theses, Universit \'e Paris-Saclay , January 2022. URL https://theses.hal.science/tel-03585897
2022
-
[44]
Properties of discrete sliced W asserstein losses
Tanguy, E., Flamary, R., and Delon, J. Properties of discrete sliced W asserstein losses. Mathematics of Computation, June 2024 a . URL https://www.ams.org/journals/mcom/0000-000-00/S0025-5718-2024-03994-7/
2024
-
[45]
Reconstructing discrete measures from projections
Tanguy, E., Flamary, R., and Delon, J. Reconstructing discrete measures from projections. consequences on the empirical sliced W asserstein distance. Comptes Rendus. Math\'ematique, 362: 0 1121--1129, 2024 b . doi:10.5802/crmath.601. URL https://comptes-rendus.academie-science...
2024 doi
-
[46]
Optimal transport -- Old and new, volume 338
Villani, C. Optimal transport -- Old and new, volume 338. Springer Berlin, Heidelberg, 01 2008
2008
-
[47]
Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem
Wibisono, A. Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem . In Conference on Learning Theory, pp.\ 2093--3027. PMLR, 2018
2018
-
[48]
and Liu, S
Yi, M. and Liu, S. Sliced wasserstein variational inference. In Asian Conference on Machine Learning, pp.\ 1213--1228. PMLR, 2023
2023
-
[49]
Monoflow: Rethinking divergence gans via the perspective of wasserstein gradient flows
Yi, M., Zhu, Z., and Liu, S. Monoflow: Rethinking divergence gans via the perspective of wasserstein gradient flows. In International Conference on Machine Learning, pp.\ 39984--40000. PMLR, 2023
2023
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.