Pith. sign in

REVIEW 3 cited by

Dataset Distillation via the Wasserstein Metric

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.18531 v3 pith:FHO3VRPL submitted 2023-11-30 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords datasetwassersteindistillationdistributionwmddbarycenterdatamatching
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Dataset Distillation (DD) aims to generate a compact synthetic dataset that enables models to achieve performance comparable to training on the full large dataset, significantly reducing computational costs. Drawing from optimal transport theory, we introduce WMDD (Wasserstein Metric-based Dataset Distillation), a straightforward yet powerful method that employs the Wasserstein metric to enhance distribution matching. We compute the Wasserstein barycenter of features from a pretrained classifier to capture essential characteristics of the original data distribution. By optimizing synthetic data to align with this barycenter in feature space and leveraging per-class BatchNorm statistics to preserve intra-class variations, WMDD maintains the efficiency of distribution matching approaches while achieving state-of-the-art results across various high-resolution datasets. Our extensive experiments demonstrate WMDD's effectiveness and adaptability, highlighting its potential for advancing machine learning applications at scale.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Noise Efficiency in Privacy-preserving Dataset Distillation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Dosser improves differentially private dataset distillation by decoupling sampling from optimization and projecting signals into a learned subspace.

  2. Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation

    stat.ML 2025-10 conditional novelty 5.0 of 10

    A mini-batch Wasserstein gradient-flow algorithm computes scalable and label-aware Wasserstein barycenters, with empirical gains on domain adaptation.

  3. A Discrepancy-Based Perspective on Dataset Condensation

    cs.LG 2025-09 conditional novelty 4.0 of 10

    Dataset condensation is reframed as minimizing distribution discrepancies, and existing methods are sorted into a taxonomy; no new algorithm or experiments are provided.

Pith tools