REVIEW 3 cited by
Hierarchical Sliced Wasserstein Distance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Sliced Wasserstein (SW) distance has been widely used in different application scenarios since it can be scaled to a large number of supports without suffering from the curse of dimensionality. The value of sliced Wasserstein distance is the average of transportation cost between one-dimensional representations (projections) of original measures that are obtained by Radon Transform (RT). Despite its efficiency in the number of supports, estimating the sliced Wasserstein requires a relatively large number of projections in high-dimensional settings. Therefore, for applications where the number of supports is relatively small compared with the dimension, e.g., several deep learning applications where the mini-batch approaches are utilized, the complexities from matrix multiplication of Radon Transform become the main computational bottleneck. To address this issue, we propose to derive projections by linearly and randomly combining a smaller number of projections which are named bottleneck projections. We explain the usage of these projections by introducing Hierarchical Radon Transform (HRT) which is constructed by applying Radon Transform variants recursively. We then formulate the approach into a new metric between measures, named Hierarchical Sliced Wasserstein (HSW) distance. By proving the injectivity of HRT, we derive the metricity of HSW. Moreover, we investigate the theoretical properties of HSW including its connection to SW variants and its computational and sample complexities. Finally, we compare the computational cost and generative quality of HSW with the conventional SW on the task of deep generative modeling using various benchmark datasets including CIFAR10, CelebA, and Tiny ImageNet.
Forward citations
Cited by 3 Pith papers
-
Post-Training Augmentation Invariance
Develops post-training augmentation invariance via augmented encoders and two Wasserstein-based losses to train lightweight MLP adapters that boost robustness on DINOv2 features for STL10 without fine-tuning.
-
ESPFormer: Doubly-Stochastic Attention with Expected Sliced Transport Plans
A new Transformer attention mechanism that builds doubly-stochastic attention from sliced optimal transport plans with soft sorting, yielding modest gains over baselines.
-
Mutual Regression Distance
Mutual Regression Distance is a new pseudometric between sample sets based on mutual linear regression, with simplified variants and empirical gains in clustering, GANs, and domain adaptation.
Discussion (0). Continue with ORCID to comment.