Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

Towards Understanding Gradient Dynamics of the Sliced-Wasserstein Distance via Critical Point Analysis

T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper proves that in dimension two, any measure carrying a positive mass on a segment is unstable for the squared sliced-Wasserstein objective, so stable critical points cannot be segment-shaped.

desk verdict Genuinely new segment-instability result for sliced Wasserstein, but the abstract overstates it: the proven perturbation is not the Lagrangian one the paper's own framework needs. read the letter →

arxiv 2502.06525 v2 pith:NW5SQEBA submitted 2025-02-10 stat.ML cs.LG

classification stat.MLcs.LG MSC 49Q22
keywords sliced-Wassersteindistancegradientflowcriticalpointsstabilitybarycentricprojectionoptimaltransportquantilefunctionssemi-discreteoptimization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Optimizing the squared sliced-Wasserstein distance $F(\mu)=\tfrac12\mathrm{SW}_2^2(\mu,\rho)$ by gradient-type particle dynamics is non-convex and can in principle stall at critical points that are not the target $\rho$. This paper shows that such bad critical points exist—including measures supported on lower-dimensional sets—but that the most dangerous ones are unstable. In dimension two, whenever a measure contains a positive fraction of mass on a segment and the target's one-dimensional projections have bounded densities, splitting that segment in two and pushing the halves in opposite directions lowers $F$ by more than any quadratic in the perturbation size $t$, for $t$ small enough. Consequently, stable critical points of the sliced-Wasserstein objective cannot concentrate on segments. The paper also provides a characterization of critical points by a barycentric fixed-point equation and proves that this characterization passes to weak limits, so limits of discrete particle critical points are critical in the continuous problem.

What carries the argument

The central object is the barycentric vector field $v_\mu(x)=\tfrac1d x-\int_{S^{d-1}}\bar\gamma_\theta(\langle x|\theta\rangle)\,\theta\,d\theta$, where $\bar\gamma_\theta$ is the conditional mean of the optimal 1-D transport plan between $P_\theta\#\mu$ and $P_\theta\#\rho$; criticality is exactly $v_\mu=0$ $\mu$-a.e. For the stability theorem, the carrying computation is the explicit quantile function of the perturbed segment under every projection, combined with the symmetry $x\mapsto 1-x$ and a lower bound on the difference quotient $G_\theta$ of the transported target; these produce the super-quadratic negative term. The same quantile machinery, via the identity $W_2^2=\|F^{-1}_\mu-F^{-1}_\rho\|_{L^2}^2$, is what makes the sliced objective analytically tractable.

What would settle it

In dimension two, choose any target $\rho$ with uniformly bounded one-dimensional projection densities and any segment-supported measure $\mu$, and compute $R(t)=(\mathrm{SW}_2^2(\mu_t,\rho)-\mathrm{SW}_2^2(\mu,\rho))/t^2$ for $\mu_t=\tfrac12(\tau_{-tn}\#\mu+\tau_{tn}\#\mu)$ with $n$ orthogonal to the segment. The paper's Proposition 5.2 entails $R(t)\to-\infty$ as $t\to0$; observing a finite limit, or any well-defined second derivative at $t=0$, would refute the instability claim.

Watch

Extended reading notes

Core claim

The central claim is that the landscape of $F(\cdot)=\tfrac12\mathrm{SW}_2^2(\cdot,\rho)$ has a two-sided character: non-global Lagrangian critical points on lower-dimensional sets exist, yet segment-supported ones are unstable in a very strong sense. With the target $\rho$ absolutely continuous and its one-dimensional projections having density bounded by $b$, and with $\mu$ carrying $a\mathcal H^1|_S$ on a segment $S$, the perturbation $\mu_t=\tfrac12(\tau_{-tn}\#\mu+\tau_{tn}\#\mu)$, $n\perp S$, satisfies $F(\mu_t)\le F(\mu)-Ct^2$ on a neighborhood of $t=0$ for every $C>0$ (Proposition 5.2). This is stronger than having a negative second derivative: the function $t\mapsto F(\mu_t)$ is not twice differentiable at $0$, and the decrease rate is arbitrarily steep as $t\to0$. Separately, the paper proves that the variational notion of a Lagrangian critical point is equivalent, under compact-support and atomless assumptions, to the barycentric equation $v_\mu=0$, and that weak limits of barycentric critical points are again barycentric critical.

Load-bearing premise

The proof needs the piece of the target that receives the segment mass to have bounded density in the direction perpendicular to the segment, and it uses the segment being exactly uniform to get closed-form quantiles; without those, the negative quadratic estimate can fail.

Editorial extensions

If this is right

  • Every weak limit of discrete barycentric critical points (for example, limits of particle configurations produced by gradient descent) is again a barycentric Lagrangian critical point of the continuous functional, so the discrete-to-continuous limit does not create new critical points.
  • There are explicit critical points distinct from the target: a uniform segment in the plane for the sliced-uniform disk target, and a Gaussian line for the standard Gaussian in higher dimensions.
  • In dimension two, a stable critical point cannot have any segment piece, regardless of what the rest of the measure does.
  • For the particle formulation, step sizes below $Nd$ make the objective strictly decrease, and with bounded projection densities the iterates stay away from particle collisions, so convergent subsequences reach critical points.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the arbitrarily steep decrease suggests that sliced-Wasserstein optimization may escape segment traps through non-smooth kink directions, so the practical advantage could come from non-differentiability rather than from curvature; this could be tested by comparing first-order algorithms that deliberately split symmetric particle pairs.
  • Beyond the paper: the proof's dependence on the segment being exactly uniform suggests the instability is a resonance between a perfectly flat lower-dimensional piece and a genuinely full-dimensional target; segment-like pieces with curved or non-uniform density profiles might behave differently, which the theorem does not cover.
  • Beyond the paper: a natural extension, left open by the authors, is to test hyperplane-supported measures in $d>2$ with multi-directional two-sided perturbations; if the same unbounded negative quadratic appears, low-dimensional collapse should be transient in all dimensions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper studies the squared Sliced-Wasserstein distance SW2^2(·, ρ) as an objective functional over probability measures. It introduces Eulerian, Wasserstein, and Lagrangian notions of critical points, proves a barycentric characterization of Lagrangian critical points (Definitions 4.1 and 4.2, Proposition 4.7), shows that weak limits of discrete barycentric critical points remain barycentric critical (Theorem 4.5), constructs explicit lower-dimensional critical points including a segment for a sliced-uniform target and a Gaussian example (Proposition 5.1), and claims in dimension 2 that stable critical points cannot concentrate on segments (Proposition 5.2). The instability is proven for a specific two-shifted-copy perturbation, not for a Lagrangian pushforward perturbation. Numerical experiments illustrate both the instability and the behavior of gradient descent.

Significance. If the main claim were established within the Lagrangian framework that the paper itself identifies as the relevant one for gradient dynamics, this would be a meaningful step toward understanding the non-convex landscape of SW optimization. The paper contains several useful building blocks: explicit gradient formulas for the semi-discrete SW objective (Proposition 3.1), a descent lemma with a valid step-size range (Proposition 3.2), a non-collapse result for gradient descent (Proposition 3.3), a limit theorem for critical points (Theorem 4.5), explicit critical-point constructions (Proposition 5.1), and a detailed quantitative instability computation (Proposition 5.2). The appendix computations are extensive and the code for the experiments is provided. The proof of Proposition 5.2 appears internally sound under the stated bounded-projection-density hypothesis; in particular, the density bound on the transported part of the segment follows automatically from ρ = (1−λ)ρ_{θ,0} + λρ_{θ,1}.

major comments (3)
  1. [Abstract and Section 1 vs. Section 5.2 and Conclusion] The abstract and introduction state that stable critical points of SW cannot concentrate on segments, but the only supporting result, Proposition 5.2, uses the perturbation μ_t = (1/2)(τ_{−tn}#μ + τ_{tn}#μ) in Eq. (22), which is a mixture of two translated copies of μ and not a pushforward of the form (Id + tξ)#μ. The paper itself acknowledges this in the remark after Proposition 5.2 and in the Conclusion, yet the unqualified claim appears in the abstract and introduction. Either the authors should prove a Lagrangian-instability analogue, or the advertised claim should be qualified to explicitly state that it concerns the specific mixture perturbation and not the Lagrangian framework.
  2. [Appendix B.10, Eq. (198) and Eq. (246)] The proof of Proposition 5.2 is labeled 'Sketch of proof' and leaves several load-bearing steps underdeveloped. The decomposition γ̂_θ = (1−λ)γ̂_{θ,0} + λγ̂_{θ,1} and the claim that the disintegrated plans are optimal between their margins are asserted with only a citation; the 'equality at t=0' in Eq. (198) is not immediate because it depends on this optimality. The bound leading to the constant C0 in Eq. (246) is stated as 'there exists some constant C0' without a fully explicit derivation or a discussion of uniformity in θ1 and t. Since Proposition 5.2 is the paper's main instability theorem, a complete proof should be provided in place of the sketch.
  3. [Section 5.2 remark and Section 6] The paper suggests that a vector field ξ alternating rapidly between n and −n will approximately realize the mixture perturbation of Proposition 5.2, and that SW2^2((Id + tξ)#μ, ρ) will therefore have a local maximum at t=0. This claim is supported only by numerical experiments, not by a theorem. Without a rigorous link between the mixture perturbation and the Lagrangian perturbation family, the paper's conclusion about gradient dynamics — which is the stated motivation of the work — remains unsupported.
minor comments (6)
  1. [Eq. (7)] The power cell intervals in Eq. (7) are written as [i/N, (i+1)/N], which is shifted by one index; they should presumably be [(i−1)/N, i/N] as used in the proofs in Appendix B.1.
  2. [Abstract] The abstract contains a typo: 'Sliced-Wassertein' should be 'Sliced-Wasserstein'.
  3. [Section 4] The sentence 'The goal of this this section' contains a duplicated 'this'.
  4. [Theorem 4.5] The phrase 'an atomles measure' should read 'an atomless measure'.
  5. [Section 4.2] The word 'Langrangian' appears in the sentence 'the two definitions of Langrangian critical points'; it should be 'Lagrangian'.
  6. [Appendix B.10] The main text refers to the proof of Proposition 5.2 as deferred to the appendix, but the appendix labels it 'Sketch of proof'. This inconsistency should be resolved, either by expanding the sketch into a complete proof or by explicitly referring to it as a sketch in the main text.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the central critical-point characterization and instability bound are derived from optimal transport first principles, with only minor non-load-bearing self-citations.

full rationale

The paper's central derivation chain is self-contained rather than circular. The Lagrangian and barycentric critical-point notions (Definitions 4.1 and 4.2) are not assumed to be equivalent; Proposition 4.4 proves the equivalence for discrete measures, Proposition 4.7 derives the directional derivative formula from convexity and optimal transport duality, and Corollary 4.8 establishes the equivalence under compact support/atomless assumptions. These results do not presuppose the conclusions of the paper. The lower-dimensional critical points in Proposition 5.1 are verified by explicit quantile and transport-map computations (e.g., T_theta(x) = pi/(4|c_theta|) x for the sliced-uniform example), not by invoking the target instability result. Proposition 5.2's instability inequality is obtained from an explicit computation of the quantile function of the split segment measure and a lower bound on the symmetric quantile difference of the transported target density; the uniform bounded-density hypothesis is used genuinely and is not a renamed version of the conclusion. The Gaussian example's constant alpha_d is a designed value chosen to make the barycentric equation vanish, not a fitted parameter subsequently called a prediction. The paper does contain self-citations, notably to M\'erigot et al. (2021) for semiconcavity of the one-dimensional Wasserstein loss and to Sarrazin (2022) for background on Lagrangian critical points, but these are standard external ingredients or context and are not load-bearing for the main claims. The acknowledged mismatch between Proposition 5.2's perturbation (a mixture of two translated copies) and the Lagrangian framework (pushforwards by Id+t xi) is a scope limitation of the stated instability result, explicitly flagged by the authors in the remark after Proposition 5.2 and in the conclusion; it is a correctness/generalization gap, not a circularity. The abstract's unqualified wording about stable critical points is stronger than what is proven, but this overstatement does not make the derivation equivalent to its inputs. Overall, the derivation is self-contained against external benchmarks, and the only reason the score is not 0 is the presence of minor self-citations that do not affect the validity of the argument.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The main new objects are the barycentric critical-point equation v_μ = 0 and the specific perturbation µ_t. These are not unexplained entities: they are derived from optimal transport plans. The only tuning-like constant is α_d in the Gaussian example.

free parameters (1)
  • α_d = d ∫_{S^{d-1}} |⟨θ,e1⟩| dθ
    Chosen constant in Proposition 5.1(b) to make the Gaussian critical point satisfy the barycentric equation. It is defined by the requirement that the identity holds.
assumptions (3)
  • domain assumption The target ρ has bounded one-dimensional projection densities in Proposition 5.2.
    Used to derive the lower bound on G_θ via the density bound b, a key step in the instability estimate.
  • ad hoc to paper The perturbation µ_t concentrates the two shifted copies of the segment with equal weight.
    This specific perturbation is constructed to make the quantile computations tractable; it is not the standard Lagrangian perturbation (Id + tξ)#µ.
  • domain assumption The measure μ contains a positive fraction λ = 2a of a uniform segment S.
    The whole proof decomposes μ into a segment component μ_1 and a remainder μ_0; the explicit quantile functions are only available for a uniform segment.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Understanding Gradient Dynamics of the Sliced-Wasserstein Distance via Critical Point Analysis." pith.science (2026). https://pith.science/paper/NW5SQEBA

@misc{pith2026250206525,
  author       = {Pith},
  title        = {Pith review of: Towards Understanding Gradient Dynamics of the Sliced-Wasserstein Distance via Critical Point Analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NW5SQEBA}},
  note         = {Machine review of arXiv:2502.06525}
}
read the original abstract

In this paper, we investigate the properties of the Sliced Wasserstein Distance (SW) when employed as an objective functional. The SW metric has gained significant interest in the optimal transport and machine learning literature, due to its ability to capture intricate geometric properties of probability distributions while remaining computationally tractable, making it a valuable tool for various applications, including generative modeling and domain adaptation. Our study aims to provide a rigorous analysis of the critical points arising from the optimization of the SW objective. By computing explicit perturbations, we establish that stable critical points of SW cannot concentrate on segments. This stability analysis is crucial for understanding the behaviour of optimization algorithms for models trained using the SW objective. Furthermore, we investigate the properties of the SW objective, shedding light on the existence and convergence behavior of critical points. We illustrate our theoretical results through numerical experiments.

Figures

Figures reproduced from arXiv: 2502.06525 by the authors.

Figure 1
Figure 1. Instability of measures containing an horizontal segment. On the top line are plotted the value SW2 2(µ t , ρ) for different measures µ, ρ and perturbations ξ. On the bottom line are depictions of the different µ (black points), ρ (approximated by the blue points) and ξ (red arrows). Columns (a) and (b): µ is a point cloud of N = 100 points uniformly distributed on the segment [−4/π, 4/π] × {0}, ξ alternates between… view at source ↗
Figure 2
Figure 2. Gradient descent of SW2 2. On a point cloud of N = 1000 points for different choices of step-size and ρ. Left : convergence speed of gradient descent, where ρ is the normal distribution, for different step-sizes (given in multiples of N in the legend). Center left : Initial point cloud (in green), sampled uniformly in [−1, 1]2 , and final point cloud (in red) after 200 iterations with step-size λ = 2N. Center right … view at source ↗
Figure 3
Figure 3. Behavior of FL for different sets of test directions. Depicts the value of FL(X + tξ), where X is a point cloud of N = 100 points uniformly distributed on the segment [−4/π, 4/π] × {0}, ξ alternates between e2 and −e2, and ρ is the sliced-uniform distribution. Each column corresponds to a different number L ∈ {10, 20, 40, 100} of fixed test directions ; on the top line e2 is included in the test directions while on … view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Convergence Rates for Distribution Matching with Sliced Optimal Transport

    stat.ML 2026-02 conditional novelty 7.0 of 10

    For Gaussian distributions, slice-matching to an isotropic target with decaying step sizes converges at rate O(k^{-(2α-1)}) in expectation.

  2. Convergence of empirical subgradients for optimal transport-based objectives

    math.OC 2026-05 unverdicted novelty 6.0 of 10

    Under smooth unit costs and models, empirical subdifferentials of parameterized transport objectives converge graphically almost surely to the population subdifferential, so subgradient methods approach population cri...

Reference graph

Works this paper leans on

49 extracted references · 41 canonical work pages · cited by 2 Pith papers

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Gradient flows: in metric spaces and in the space of probability measures

    Ambrosio, L., Gigli, N., and Savar \'e , G. Gradient flows: in metric spaces and in the space of probability measures. Springer Science & Business Media, 2005

  3. [3]

    W asserstein generative adversarial networks

    Arjovsky, M., Chintala, S., and Bottou, L. W asserstein generative adversarial networks. In Precup, D. and Teh, Y. W. (eds.), Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pp.\ 214--223. PMLR, 06--11 Aug 2017

  4. [4]

    A., Pantazis, Y., and Rey-Bellet, L

    Birrell, J., Dupuis, P., Katsoulakis, M. A., Pantazis, Y., and Rey-Bellet, L. (f, gamma)-divergences: Interpolating between f-divergences and integral probability metrics. Journal of machine learning research, 23 0 (39): 0 1--70, 2022

  5. [5]

    M., Kucukelbir, A., and McAuliffe, J

    Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. Variational Inference: A Review for Statisticians . Journal of the American statistical Association, 112 0 (518): 0 859--877, 2017

  6. [6]

    Bond-Taylor, S., Leach, A., Long, Y., and Willcocks, C. G. Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models . IEEE transactions on pattern analysis and machine intelligence, 44 0 (11): 0 7327--7347, 2021

  7. [7]

    Sliced and radon wasserstein barycenters of measures

    Bonneel, N., Rabin, J., Peyr \'e , G., and Pfister, H. Sliced and radon wasserstein barycenters of measures. Journal of Mathematical Imaging and Vision, 51: 0 22--45, 2015

  8. [8]

    Unidimensional and evolution methods for optimal transportation

    Bonnotte, N. Unidimensional and evolution methods for optimal transportation. PhD thesis, Universit \'e Paris Sud-Paris XI; Scuola normale superiore (Pise, Italie), 2013

Show all 49 references
  1. [9]

    P., Kok, P

    Bourne, D. P., Kok, P. J., Roper, S. M., and Spanjer, W. D. Laguerre tessellations and polycrystalline microstructures: a fast algorithm for generating grains of given volumes. Philosophical Magazine, 100 0 (21): 0 2677--2707, 2020

  2. [10]

    EigenVI: score-based variational inference with orthogonal function expansions

    Cai, D., Modi, C., Margossian, C., Gower, R., Blei, D., and Saul, L. EigenVI: score-based variational inference with orthogonal function expansions . Advances in Neural Information Processing Systems, 37: 0 132691--132721, 2024 a

  3. [11]

    C., Gower, R

    Cai, D., Modi, C., Pillaud-Vivien, L., Margossian, C. C., Gower, R. M., Blei, D. M., and Saul, L. K. Batch and match: black-box variational inference with a score-based divergence. In International Conference on Machine Learning, 2024 b

  4. [12]

    and Bach, F

    Chizat, L. and Bach, F. On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport . Advances in neural information processing systems, 31, 2018

  5. [13]

    Scalable wasserstein gradient flow for generative modeling through unbalanced optimal transport

    Choi, J., Choi, J., and Kang, M. Scalable wasserstein gradient flow for generative modeling through unbalanced optimal transport. In International Conference on Machine Learning, 2024

  6. [14]

    and Santambrogio, F

    Cozzi, G. and Santambrogio, F. Long-time asymptotics of the sliced-wasserstein flow. SIAM J. Im. Sciences, 2024. URL http://cvgmt.sns.it/paper/6495/. cvgmt preprint

  7. [15]

    and Seljak, U

    Dai, B. and Seljak, U. Sliced iterative normalizing flows. Proceedings of Machine Learning Research, 2021

  8. [16]

    Deshpande, I., Hu, Y.-T., Sun, R., Pyrros, A., Siddiqui, N., Koyejo, S., Zhao, Z., Forsyth, D., and Schwing, A. G. Max-sliced wasserstein distance and its use for gans. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 10648--10656, 2019

  9. [17]

    Nonparametric generative modeling with conditional sliced-wasserstein flows

    Du, C., Li, T., Pang, T., Yan, S., and Lin, M. Nonparametric generative modeling with conditional sliced-wasserstein flows. In International Conference on Machine Learning (ICML), 2023

  10. [18]

    Variational wasserstein gradient flow

    Fan, J., Zhang, Q., Taghvaei, A., and Chen, Y. Variational wasserstein gradient flow. In International Conference on Machine Learning, 2022

  11. [19]

    Measure transport with kernel stein discrepancy

    Fisher, M., Nolan, T., Graham, M., Prangle, D., and Oates, C. Measure transport with kernel stein discrepancy. In International Conference on Artificial Intelligence and Statistics, pp.\ 1054--1062. PMLR, 2021

  12. [20]

    Learning generative models with sinkhorn divergences

    Genevay, A., Peyr \'e , G., and Cuturi, M. Learning generative models with sinkhorn divergences. In International Conference on Artificial Intelligence and Statistics, pp.\ 1608--1617. PMLR, 2018

  13. [21]

    Generative adversarial networks

    Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. Communications of the ACM, 63 0 (11): 0 139--144, 2020

  14. [22]

    Generative sliced MMD flows with R iesz kernels

    Hertrich, J., Wald, C., Altekr \"u ger, F., and Hagemann, P. Generative sliced MMD flows with R iesz kernels. In International Conference of Learning Representations, 2024

  15. [23]

    E., Martin, C

    Kolouri, S., Pope, P. E., Martin, C. E., and Rohde, G. K. Sliced wasserstein auto-encoders. In International Conference on Learning Representations, 2018

  16. [24]

    Kernel stein discrepancy descent

    Korba, A., Aubin-Frankowski, P.-C., Majewski, S., and Ablin, P. Kernel stein discrepancy descent. In International Conference on Machine Learning, pp.\ 5719--5730. PMLR, 2021

  17. [25]

    Mmd gan: Towards deeper understanding of moment matching network

    Li, C.-L., Chang, W.-C., Cheng, Y., Yang, Y., and P \'o czos, B. Mmd gan: Towards deeper understanding of moment matching network. Advances in neural information processing systems, 30, 2017

  18. [26]

    and Moosmueller, C

    Li, S. and Moosmueller, C. Measure transfer via stochastic slicing and matching, 2025. URL https://arxiv.org/abs/2307.05705

  19. [27]

    Sliced-wasserstein flows: Nonparametric generative modeling via optimal transport and diffusions

    Liutkus, A., Simsekli, U., Majewski, S., Durmus, A., and St \"o ter, F.-R. Sliced-wasserstein flows: Nonparametric generative modeling via optimal transport and diffusions. In International Conference on Machine Learning, pp.\ 4104--4113. PMLR, 2019

  20. [28]

    Fast optimal transport through sliced generalized wasserstein geodesics

    Mahey, G., Chapel, L., Gasso, G., Bonet, C., and Courty, N. Fast optimal transport through sliced generalized wasserstein geodesics. Advances in Neural Information Processing Systems, 36, 2024

  21. [29]

    Minimax confidence intervals for the sliced wasserstein distance

    Manole, T., Balakrishnan, S., and Wasserman, L. Minimax confidence intervals for the sliced wasserstein distance. Electronic Journal of Statistics, 16 0 (1): 0 2252--2345, 2022

  22. [30]

    A Mean Field View of the Landscape of Two-Layer Neural Networks

    Mei, S., Montanari, A., and Nguyen, P.-M. A Mean Field View of the Landscape of Two-Layer Neural Networks . Proceedings of the National Academy of Sciences, 115 0 (33): 0 E7665--E7671, 2018

  23. [31]

    Non-asymptotic convergence bounds for wasserstein approximation using point clouds

    M \'e rigot, Q., Santambrogio, F., and Sarrazin, C. Non-asymptotic convergence bounds for wasserstein approximation using point clouds. In Neural Information Processing Systems, 2021. URL https://api.semanticscholar.org/CorpusID:235436063

  24. [32]

    Statistical and topological properties of sliced probability divergences

    Nadjahi, K., Durmus, A., Chizat, L., Kolouri, S., Shahrampour, S., and Simsekli, U. Statistical and topological properties of sliced probability divergences. Advances in Neural Information Processing Systems, 33: 0 20802--20812, 2020

  25. [33]

    and Ho, N

    Nguyen, K. and Ho, N. Energy-based sliced wasserstein distance. Advances in Neural Information Processing Systems, 36, 2024

  26. [34]

    Statistical, robustness, and computational guarantees for sliced wasserstein distances

    Nietert, S., Goldfeld, Z., Sadhu, R., and Kato, K. Statistical, robustness, and computational guarantees for sliced wasserstein distances. Advances in Neural Information Processing Systems, 35: 0 28179--28193, 2022

  27. [35]

    First-order methods almost always avoid saddle points: The case of vanishing step-sizes

    Panageas, I., Piliouras, G., and Wang, X. First-order methods almost always avoid saddle points: The case of vanishing step-sizes. In Wallach, H., Larochelle, H., Beygelzimer, A., d Alch\' e -Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing S...

  28. [36]

    Computational optimal transport: With applications to data science

    Peyr \'e , G., Cuturi, M., et al. Computational optimal transport: With applications to data science. Foundations and Trends in Machine Learning , 11 0 (5-6): 0 355--607, 2019

  29. [37]

    On the sequential convergence of Lloyd's algorithms

    Portales, L., Cazelles, E., and Pauwels, E. On the sequential convergence of Lloyd's algorithms . working paper or preprint, May 2024. URL https://hal.science/hal-04593982

  30. [38]

    Wasserstein barycenter and its application to texture mixing

    Rabin, J., Peyr \'e , G., Delon, J., and Bernot, M. Wasserstein barycenter and its application to texture mixing. In Scale Space and Variational Methods in Computer Vision: Third International Conference, SSVM 2011, Ein-Gedi, Israel, May 29--June 2, 2011, Revised Selected Pape...

  31. [39]

    Rachev, S. T. and R \"u schendorf, L. Mass transportation problems: Volume I: theory, volume 1. Springer Science & Business Media, 1998

  32. [40]

    Orthogonal estimation of wasserstein distances

    Rowland, M., Hron, J., Tang, Y., Choromanski, K., Sarlos, T., and Weller, A. Orthogonal estimation of wasserstein distances. In The 22nd International Conference on Artificial Intelligence and Statistics, pp.\ 186--195. PMLR, 2019

  33. [41]

    The Wasserstein Proximal Gradient Algorithm

    Salim, A., Korba, A., and Luise, G. The Wasserstein Proximal Gradient Algorithm . Advances in Neural Information Processing Systems, 33: 0 12356--12366, 2020

  34. [42]

    Optimal transport for applied mathematicians

    Santambrogio, F. Optimal transport for applied mathematicians. Birk \"a user, NY , 55 0 (58-63): 0 94, 2015

  35. [43]

    Lagrangian discretization of variational problems in Wasserstein spaces

    Sarrazin, C. Lagrangian discretization of variational problems in Wasserstein spaces . Theses, Universit \'e Paris-Saclay , January 2022. URL https://theses.hal.science/tel-03585897

  36. [44]

    Properties of discrete sliced W asserstein losses

    Tanguy, E., Flamary, R., and Delon, J. Properties of discrete sliced W asserstein losses. Mathematics of Computation, June 2024 a . URL https://www.ams.org/journals/mcom/0000-000-00/S0025-5718-2024-03994-7/

  37. [45]

    Reconstructing discrete measures from projections

    Tanguy, E., Flamary, R., and Delon, J. Reconstructing discrete measures from projections. consequences on the empirical sliced W asserstein distance. Comptes Rendus. Math\'ematique, 362: 0 1121--1129, 2024 b . doi:10.5802/crmath.601. URL https://comptes-rendus.academie-science...

  38. [46]

    Optimal transport -- Old and new, volume 338

    Villani, C. Optimal transport -- Old and new, volume 338. Springer Berlin, Heidelberg, 01 2008

  39. [47]

    Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem

    Wibisono, A. Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem . In Conference on Learning Theory, pp.\ 2093--3027. PMLR, 2018

  40. [48]

    and Liu, S

    Yi, M. and Liu, S. Sliced wasserstein variational inference. In Asian Conference on Machine Learning, pp.\ 1213--1228. PMLR, 2023

  41. [49]

    Monoflow: Rethinking divergence gans via the perspective of wasserstein gradient flows

    Yi, M., Zhu, Z., and Liu, S. Monoflow: Rethinking divergence gans via the perspective of wasserstein gradient flows. In International Conference on Machine Learning, pp.\ 39984--40000. PMLR, 2023

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.