Pith. sign in

REVIEW 2 cited by

Learning with minibatch Wasserstein : asymptotic and gradient properties

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04091 v4 pith:REZNNERJ submitted 2019-10-09 stat.ML cs.LG

Learning with minibatch Wasserstein : asymptotic and gradient properties

classification stat.ML cs.LG
keywords analysisdistancesgradientlearningoptimalpropertiestransportalgorithmic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Optimal transport distances are powerful tools to compare probability distributions and have found many applications in machine learning. Yet their algorithmic complexity prevents their direct use on large scale datasets. To overcome this challenge, practitioners compute these distances on minibatches {\em i.e.} they average the outcome of several smaller optimal transport problems. We propose in this paper an analysis of this practice, which effects are not well understood so far. We notably argue that it is equivalent to an implicit regularization of the original problem, with appealing properties such as unbiased estimators, gradients and a concentration bound around the expectation, but also with defects such as loss of distance property. Along with this theoretical analysis, we also conduct empirical experiments on gradient flows, GANs or color transfer that highlight the practical interest of this strategy.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Multivariate Distributional Reinforcement Learning Using Sliced Divergences

    cs.LG 2026-05 unverdicted novelty 7.0

    SDRL applies sliced projections of one-dimensional divergences (Wasserstein, Cramér, MMD) to multivariate return distributions in RL, with Bellman contraction proofs for scalar and general matrix discounting.

  2. Wasserstein normalized autoencoder for anomaly detection

    hep-ex 2025-10 conditional novelty 6.0

    A Wasserstein-distance-trained normalized autoencoder detects semivisible jets in simulated LHC events with AUCs around 0.69–0.77, outperforming standard and normalized autoencoders on a ttbar background.