Pith. sign in

REVIEW 2 cited by

Coresets for Data-efficient Training of Machine Learning Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.01827 v3 pith:LDW4QRJI submitted 2019-06-05 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords trainingsubsetcraigdata-efficientgradientlearningmachinemethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Incremental gradient (IG) methods, such as stochastic gradient descent and its variants are commonly used for large scale optimization in machine learning. Despite the sustained effort to make IG methods more data-efficient, it remains an open question how to select a training data subset that can theoretically and practically perform on par with the full dataset. Here we develop CRAIG, a method to select a weighted subset (or coreset) of training data that closely estimates the full gradient by maximizing a submodular function. We prove that applying IG to this subset is guaranteed to converge to the (near)optimal solution with the same convergence rate as that of IG for convex optimization. As a result, CRAIG achieves a speedup that is inversely proportional to the size of the subset. To our knowledge, this is the first rigorous method for data-efficient training of general machine learning models. Our extensive set of experiments show that CRAIG, while achieving practically the same solution, speeds up various IG methods by up to 6x for logistic regression and 3x for training deep neural networks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

    cs.AI 2026-07 accept novelty 6.0 of 10

    Facility location on semantic prompt embeddings selects evaluation-unsupervised prompt coresets that preserve LLM scores and rankings better than twelve baselines across 35 benchmarks.

  2. GRID: Scaling Task-Agnostic Inference in Continual Prompt Tuning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    GRID combines output-space constrained decoding with gradient-guided prompt compression for task-agnostic, bounded-memory continual prompt tuning.

Pith tools