Pith. sign in

REVIEW 2 cited by

BERT on a Data Diet: Finding Important Examples by Gradient-Based Pruning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.05610 v2 pith:BR5ILJPE submitted 2022-11-10 cs.CL

classification cs.CL
keywords examplesel2ngrandimportantmetricsfindinggradient-basedperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current pre-trained language models rely on large datasets for achieving state-of-the-art performance. However, past research has shown that not all examples in a dataset are equally important during training. In fact, it is sometimes possible to prune a considerable fraction of the training set while maintaining the test performance. Established on standard vision benchmarks, two gradient-based scoring metrics for finding important examples are GraNd and its estimated version, EL2N. In this work, we employ these two metrics for the first time in NLP. We demonstrate that these metrics need to be computed after at least one epoch of fine-tuning and they are not reliable in early steps. Furthermore, we show that by pruning a small portion of the examples with the highest GraNd/EL2N scores, we can not only preserve the test accuracy, but also surpass it. This paper details adjustments and implementation choices which enable GraNd and EL2N to be applied to NLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Position: Stop Preaching and Start Practising Data Frugality for Responsible Development of AI

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Data frugality is practical: pruning 25% of ImageNet-like datasets can cut training energy by roughly 29–33% with negligible accuracy loss, while dataset-level carbon costs are substantial.

  2. A Survey of LLM $\times$ DATA

    cs.DB 2025-05 conditional novelty 5.0 of 10

    A comprehensive survey of the bidirectional links between LLMs and data management, organized as DATA4LLM and LLM4DATA with a new 'IaaS' data-quality framework.

Pith tools