REVIEW 4 cited by
Vision-Language Dataset Distillation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Dataset distillation methods reduce large-scale datasets to smaller sets of synthetic data, preserving sufficient information to quickly train a new model from scratch. However, prior work on dataset distillation has focused exclusively on image classification datasets, whereas modern large-scale datasets are primarily vision-language datasets. In this work, we design the first vision-language dataset distillation method, building on the idea of trajectory matching. A key challenge is that vision-language datasets do not have a set of discrete classes. To overcome this, our proposed method jointly distills image-text pairs in a contrastive formulation. Further, we leverage Low-Rank Adaptation (LoRA) matching to enable more efficient and effective trajectory matching in complex modern vision-language models. Since there are no existing baselines, we compare our distillation approach with three adapted vision-language coreset selection methods. We demonstrate significant improvements on the challenging Flickr30K and COCO retrieval benchmarks: for example, on Flickr30K, the best coreset selection method selecting 1000 image-text pairs for training achieves only 5.6% image-to-text retrieval accuracy (i.e., recall@1); in contrast, our dataset distillation almost doubles that to 9.9% with just 100 training pairs, an order of magnitude fewer.
Forward citations
Cited by 4 Pith papers
-
Multi-Modal Dataset Distillation in the Wild
MDW distills noisy image-text data into small clean synthetic sets using learnable soft matching probabilities, Grad-CAM guided pixel weighting, and a noise-tolerant negative match loss.
-
Dataset Distillation by Influence Matching
Inf-Match distills datasets by matching estimated parameter influence of real and synthetic data, reporting SOTA classification and retrieval, but with an unsupported theoretical core.
-
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
Grouping instruction-tuning datasets by redundancy, uniqueness, or synergy of text-image interaction improves vision-language model accuracy over single-task and unselective multi-task tuning.
-
Video Set Distillation: Information Diversification and Temporal Densification
IDTD distills a video dataset into a small set of synthetic videos by jointly reducing redundancy between videos and inside each video, and it reports accuracy gains over prior video dataset distillation baselines on ...
Discussion (0). Continue with ORCID to comment.