Pith. sign in

REVIEW 12 cited by

Dataset Condensation with Gradient Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.05929 v3 pith:CMBUID3P submitted 2020-06-10 cs.CV cs.LG

classification cs.CVcs.LG
keywords datasetlearningneuraltrainingcondensationdatasetsdeepgradient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the state-of-the-art machine learning methods in many fields rely on larger datasets, storing datasets and training models on them become significantly more expensive. This paper proposes a training set synthesis technique for data-efficient learning, called Dataset Condensation, that learns to condense large dataset into a small set of informative synthetic samples for training deep neural networks from scratch. We formulate this goal as a gradient matching problem between the gradients of deep neural network weights that are trained on the original and our synthetic data. We rigorously evaluate its performance in several computer vision benchmarks and demonstrate that it significantly outperforms the state-of-the-art methods. Finally we explore the use of our method in continual learning and neural architecture search and report promising gains when limited memory and computations are available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation

    cs.CR 2025-06 conditional novelty 7.0 of 10

    A systematic survey and benchmark showing that diffusion-based synthetic data can achieve better utility-privacy tradeoffs than DP-SGD on real data for some image classifiers, with the best release strategy depending ...

  2. Hard Labels In! Rethinking the Role of Hard Labels in Mitigating Local Semantic Drift

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A soft-hard-soft training schedule uses hard labels as an intermediate anchor to correct local semantic drift and improves accuracy under 100x-reduced soft-label storage.

  3. A computational fluid dynamics model for the simulation of flashboiling flow inside pressurized metered dose inhalers

    physics.flu-dyn 2025-08 unverdicted novelty 6.0 of 10

    The abstract claims a first-of-kind open-source CFD model, combining Volume-of-Fluid and cavitation modeling, that quantitatively predicts flashboiling flow in pressurized metered dose inhalers, but the supplied text ...

  4. GVD: Guiding Video Diffusion Model for Scalable Video Distillation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GVD guides a pre-trained video diffusion model with clustering-derived features to distill video datasets, outperforming prior methods on MiniUCF and HMDB51 while retaining over 70% of full-data accuracy using under 4...

  5. Approximating Language Model Training Data from Weights

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A gradient-based greedy selection method (SELECT) recovers effective substitute fine-tuning data from two language model checkpoints, approaching the original model's performance on classification and SFT tasks.

  6. Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Post-training an LLM on economic reasoning problems improves accuracy on economic benchmarks and, without game-specific training, raises its Nash equilibrium frequency and win rates in strategic games.

  7. GLOBE: Trajectory-Aligned Gradient Matching with Structured SparseOptimization for Coreset Selection

    cs.LG 2026-08 conditional novelty 5.0 of 10

    GLOBE selects compact training subsets by matching multi-checkpoint gradient trajectories and their second-order statistics under structured sparsity, outperforming prior coreset methods on six image benchmarks.

  8. NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds

    cs.LG 2025-08 reject novelty 5.0 of 10

    A t-SNE-based sampling algorithm with differential evolution and near-memory hardware is claimed to speed up edge DNN training and reduce memory energy.

  9. Data-Distill-Net: A Data Distillation Approach Tailored for Reply-based Continual Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A plug-in module that generates learned soft labels for memory buffer samples improves accuracy and reduces forgetting across several replay-based continual learning baselines.

  10. Data-Efficient Ensemble Weather Forecasting with Diffusion Models

    cs.LG 2025-09 conditional novelty 4.0 of 10

    Training an autoregressive diffusion weather forecaster on 20% of ERA5 data selected uniformly by calendar month matches full-data CRPS/RMSE and improves the spread-skill ratio on the 2018 test year.

  11. Domain-Generalization to Improve Learning in Meta-Learning Algorithms

    cs.LG 2025-08 reject novelty 4.0 of 10

    DGS-MAML layers gradient matching onto SharpMAML and claims O(1/T) convergence and tighter PAC-Bayes bounds, but the displayed theorems give O(1/sqrt T) under the paper's own parameter choices.

  12. Leveraging Distribution Matching to Make Approximate Machine Unlearning Faster

    cs.LG 2025-07 reject novelty 4.0 of 10

    A dual data and loss-centric method claims to speed up machine unlearning, but its MIA regularizer cancels itself and the test set is leaked into training.

Pith tools