Pith. sign in

Automatic Document Selection for Efficient Encoder Pretraining

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Building pretrained language models is considered expensive and data-intensive, but must we increase dataset size to achieve better performance? We propose an alternative to larger training sets by automatically identifying smaller yet domain-representative subsets. We extend Cynical Data Selection, a statistical sentence scoring method that conditions on a representative target domain corpus. As an example, we treat the OntoNotes corpus as a target domain and pretrain a RoBERTa-like encoder from a cynically selected subset of the Pile. On both perplexity and across several downstream tasks in the target domain, it consistently outperforms random selection with 20x less data, 3x fewer training iterations, and 2x less estimated cloud compute cost, validating the recipe of automatic document selection for LM pretraining.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Approximating Language Model Training Data from Weights

cs.CL · 2025-06-18 · conditional · novelty 6.0

A gradient-based greedy selection method (SELECT) recovers effective substitute fine-tuning data from two language model checkpoints, approaching the original model's performance on classification and SFT tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Approximating Language Model Training Data from Weights cs.CL · 2025-06-18 · conditional · none · ref 15 · internal anchor

    A gradient-based greedy selection method (SELECT) recovers effective substitute fine-tuning data from two language model checkpoints, approaching the original model's performance on classification and SFT tasks.