Pith. sign in

REVIEW 20 cited by

Dataset Condensation with Gradient Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.05929 v3 pith:CMBUID3P submitted 2020-06-10 cs.CV cs.LG

Dataset Condensation with Gradient Matching

classification cs.CV cs.LG
keywords datasetlearningneuraltrainingcondensationdatasetsdeepgradient
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As the state-of-the-art machine learning methods in many fields rely on larger datasets, storing datasets and training models on them become significantly more expensive. This paper proposes a training set synthesis technique for data-efficient learning, called Dataset Condensation, that learns to condense large dataset into a small set of informative synthetic samples for training deep neural networks from scratch. We formulate this goal as a gradient matching problem between the gradients of deep neural network weights that are trained on the original and our synthetic data. We rigorously evaluate its performance in several computer vision benchmarks and demonstrate that it significantly outperforms the state-of-the-art methods. Finally we explore the use of our method in continual learning and neural architecture search and report promising gains when limited memory and computations are available.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Condensing Large-Scale Datasets Directly with Minimal Information Loss

    cs.CV 2026-07 unverdicted novelty 7.0

    CIM directly aligns data distributions to condense large-scale datasets with minimal information loss, achieving new SOTA results on ImageNet-1K distillation at IPC=10.

  2. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

    cs.CV 2026-05 unverdicted novelty 7.0

    DMGD achieves better performance than fine-tuned SOTA methods in dataset distillation on ImageNet subsets by using semantic matching through conditional likelihood optimization and OT-based distribution matching in a ...

  3. Direct Discrepancy Replay: Distribution-Discrepancy Condensation and Manifold-Consistent Replay for Continual Face Forgery Detection

    cs.CV 2026-04 unverdicted novelty 7.0

    A replay method for continual face forgery detection condenses real-fake distribution discrepancies into compact maps and synthesizes compatible samples from current real faces to reduce forgetting under tight memory ...

  4. Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets

    cs.CV 2026-06 unverdicted novelty 6.0

    Introduces SGR and TIAT for robust dataset distillation that suppresses noise while preserving knowledge under noisy supervision.

  5. Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

    cs.CV 2026-06 unverdicted novelty 6.0

    RAHA applies rank-aware hyperbolic alignment to vision-language dataset distillation by enforcing geodesic alignment in the shared low-rank range and regularizing the residual subspace for improved transfer.

  6. Pool-Select-Refine for Allocation-Aware Generative Dataset Distillation

    cs.CV 2026-06 unverdicted novelty 6.0

    Pool-Select-Refine decouples generation, selection, and refinement in diffusion-based dataset distillation to allocate a fixed synthetic sample budget more effectively, delivering consistent gains over prior baselines...

  7. Pool-Select-Refine for Allocation-Aware Generative Dataset Distillation

    cs.CV 2026-06 unverdicted novelty 6.0

    A two-stage framework that decouples generation, selection, and refinement to improve budget use in diffusion-based dataset distillation.

  8. Multimodal Distribution Matching for Vision-Language Dataset Distillation

    cs.CV 2026-05 unverdicted novelty 6.0

    MDM distills vision-language datasets via joint embedding clustering, weight-space model interpolation, and geometry-aware distribution matching on the unit hypersphere.

  9. DIVER:Diving Deeper into Distilled Data via Expressive Semantic Recovery

    cs.CV 2026-05 unverdicted novelty 6.0

    DIVER is a dual-stage distillation method using diffusion models to enhance semantic preservation and cross-architecture generalization in dataset distillation.

  10. Fair Dataset Distillation via Cross-Group Barycenter Alignment

    cs.LG 2026-04 unverdicted novelty 6.0

    Dataset distillation introduces fairness gaps from subgroup pattern mismatches rather than just imbalance; distilling to a group-agnostic barycenter of predictive information reduces these gaps.

  11. Omnimodal Dataset Distillation via High-order Proxy Alignment

    cs.CV 2026-04 unverdicted novelty 6.0

    HoPA captures high-order cross-modal alignments via a shared proxy to enable scalable omnimodal dataset distillation with better performance-compression trade-offs.

  12. ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks

    cs.CR 2026-03 unverdicted novelty 6.0

    ROAST selectively trains anomaly detectors on less vulnerable patient data with targeted outlier exposure, boosting recall by 16.2% in black-box settings and reducing training time by 88.3%.

  13. Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic Drift

    cs.CV 2025-12 conditional novelty 6.0

    A soft-hard-soft training schedule uses hard labels as an intermediate anchor to correct local semantic drift and improves accuracy under 100x-reduced soft-label storage.

  14. Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation

    cs.CV 2026-06 unverdicted novelty 5.0

    DO-ALL uses dataset distillation to create synthetic source anchors that enable stable long-term continual test-time adaptation without storing original source data.

  15. Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation

    cs.CV 2026-06 unverdicted novelty 5.0

    DO-ALL applies dataset distillation to generate synthetic source anchors that stabilize continual test-time adaptation under evolving domains without storing original source data.

  16. An Efficient and Scalable Graph Condensation with Structure-Preserving

    cs.LG 2026-05 unverdicted novelty 5.0

    SP-ESGC decouples graph condensation into heat-kernel node condensation and pre-trained edge prediction for structure, claiming high efficiency and cross-GNN generalization on real-world datasets.

  17. DIVER:Diving Deeper into Distilled Data via Expressive Semantic Recovery

    cs.CV 2026-05 unverdicted novelty 5.0

    DIVER applies a pre-trained diffusion model in a dual-stage process of semantic inheritance, guidance, and fusion to improve semantic expression and cross-architecture generalization in dataset distillation.

  18. CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection

    cs.CV 2026-05 unverdicted novelty 5.0

    CAST selects better multimodal coresets by fusing collapse-aware topologies across modalities and matching distributions at multiple scales in the diffusion wavelet domain.

  19. Data-Efficient Ensemble Weather Forecasting with Diffusion Models

    cs.LG 2025-09 conditional novelty 4.0

    Training an autoregressive diffusion weather forecaster on 20% of ERA5 data selected uniformly by calendar month matches full-data CRPS/RMSE and improves the spread-skill ratio on the 2018 test year.

  20. Domain-Generalization to Improve Learning in Meta-Learning Algorithms

    cs.LG 2025-08 reject novelty 4.0

    DGS-MAML layers gradient matching onto SharpMAML and claims O(1/T) convergence and tighter PAC-Bayes bounds, but the displayed theorems give O(1/sqrt T) under the paper's own parameter choices.