REVIEW 12 cited by
Does your data spark joy? Performance gains from domain upsampling at the end of training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Pretraining datasets for large language models (LLMs) have grown to trillions of tokens composed of large amounts of CommonCrawl (CC) web scrape along with smaller, domain-specific datasets. It is expensive to understand the impact of these domain-specific datasets on model capabilities as training at large FLOP scales is required to reveal significant changes to difficult and emergent benchmarks. Given the increasing cost of experimenting with pretraining data, how does one determine the optimal balance between the diversity in general web scrapes and the information density of domain specific data? In this work, we show how to leverage the smaller domain specific datasets by upsampling them relative to CC at the end of training to drive performance improvements on difficult benchmarks. This simple technique allows us to improve up to 6.90 pp on MMLU, 8.26 pp on GSM8K, and 6.17 pp on HumanEval relative to the base data mix for a 7B model trained for 1 trillion (T) tokens, thus rivaling Llama-2 (7B)$\unicode{x2014}$a model trained for twice as long. We experiment with ablating the duration of domain upsampling from 5% to 30% of training and find that 10% to 20% percent is optimal for navigating the tradeoff between general language modeling capabilities and targeted benchmarks. We also use domain upsampling to characterize at scale the utility of individual datasets for improving various benchmarks by removing them during this final phase of training. This tool opens up the ability to experiment with the impact of different pretraining datasets at scale, but at an order of magnitude lower cost compared to full pretraining runs.
Forward citations
Cited by 12 Pith papers
-
Predicting Emergent Capabilities by Finetuning
Finetuning small models shifts the point where capability emerges, and extrapolating this shift to the low-data limit predicts few-shot emergence up to about 4x the compute in advance.
-
Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
Benchmark signal-to-noise ratio, measured as score spread across models divided by checkpoint-to-checkpoint variability, predicts small-to-large model decision accuracy and can be improved by subtask filtering, checkp...
-
Language Models Improve When Pretraining Data Matches Target Tasks
Ranking pretraining documents by similarity to benchmark training examples (BETR) yields consistent benchmark gains and a 2.1x compute multiplier over DCLM-Baseline.
-
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
CLIMB automatically discovers pre-training data mixtures by clustering text embeddings and iteratively refining mixture weights with a predictor, improving 1B-model reasoning accuracy over standard baselines.
-
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
Two-phase pretraining, wide web data first and high-quality data second, improves average downstream accuracy by 3.4% over random ordering and 17% over natural token distribution, and the best 1T-scale blend transfers...
-
The Zamba2 Suite: Technical Report
This paper introduces Zamba2, a suite of 1.2B, 2.7B, and 7.4B hybrid Mamba2-transformer models that claims state-of-the-art small-model quality and 30-50% lower time-to-first-token, with open weights and a 5T-token pr...
-
ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
ZUNA1.1, an open-source 380M EEG diffusion autoencoder, reconstructs variable-length, flexibly masked EEG at least as well as its predecessor and far better than spherical spline interpolation.
-
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
Running multiple short annealing runs at different token scales can reveal per-source utility scaling curves that change data-source rankings compared with single point estimates.
-
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
A new benchmark called GEOHALUBENCH measures how often LLMs invent, omit, or confuse real-world places and relations, and a dynamic-beta KTO method reduces these errors on the benchmark.
-
Trillion 7B Technical Report
Trillion-7B pairs Korean documents with English documents during pretraining and lets them attend to each other, claiming competitive Korean performance with only about 10% multilingual tokens.
-
Sparse Upcycling: Inference Inefficient Finetuning
Sparse upcycling beats continued pretraining on quality by up to roughly 20 percent at matched compute, but cut serving throughput by 34 to 44 percent in vLLM benchmarks.
-
Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
A continued-pretrained 8B cybersecurity LLM claims to match GPT-4o-mini and Llama 3.1-70B on certain cyber threat intelligence benchmarks, but the decisive benchmark overlaps with its training corpus.
Discussion (0). Continue with ORCID to comment.