Pith. sign in

REVIEW 4 cited by

DeepCore: A Comprehensive Library for Coreset Selection in Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.08499 v3 pith:GJTRVXII submitted 2022-04-18 cs.LG cs.CV

classification cs.LGcs.CV
keywords learningselectioncoresetmethodsdatasetsdeepcifar10comprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Coreset selection, which aims to select a subset of the most informative training samples, is a long-standing learning problem that can benefit many downstream tasks such as data-efficient learning, continual learning, neural architecture search, active learning, etc. However, many existing coreset selection methods are not designed for deep learning, which may have high complexity and poor generalization performance. In addition, the recently proposed methods are evaluated on models, datasets, and settings of different complexities. To advance the research of coreset selection in deep learning, we contribute a comprehensive code library, namely DeepCore, and provide an empirical study on popular coreset selection methods on CIFAR10 and ImageNet datasets. Extensive experiments on CIFAR10 and ImageNet datasets verify that, although various methods have advantages in certain experiment settings, random selection is still a strong baseline.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection

    cs.CV 2025-06 conditional novelty 6.0 of 10

    RAM-APL combines distance rankings and pseudo-class label accuracy from two foundation models to select training subsets, outperforming twelve baselines on fine-grained image datasets.

  2. Info-Coevolution: An Efficient Framework for Data Model Coevolution

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A data-model coevolution framework that fuses model and nearest-neighbor predictions to select labels, reaching ImageNet-1K accuracy with 68% of annotations and 50% under semi-supervised training.

  3. Learning Faster without Deeper Networks: A*-Inspired Batch Selection for Efficient CNN Training

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A*-inspired mini-batch selection with a loss-based difficulty score and reuse penalty improves CNN accuracy on all twelve MedMNIST-2D tasks and outperforms ResNet baselines on six.

  4. X-Factor: Quality Is a Dataset-Intrinsic Property

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Across 2,500 class-balanced MNIST subsets and 10 model architectures, test-error Z-scores correlate strongly across models (mean R2=0.82 excluding GNB), supporting dataset quality as an intrinsic property.

Pith tools