WARP recovers training domain mixtures from fine-tuned model weights using weight-space interpolation via model merging to generate pseudo-checkpoints and geometric features mapped to proportions.
Autoscale: Scale-aware data mixing for pre-training llms.arXiv preprint arXiv:2407.20177
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
Facility location — a classic submodular coverage objective — predicts a training subset's held-out accuracy far better than the Vendi score, which becomes misleading at high values.
CAMEL is a scaling law capturing nonlinear model-size and mixture interactions to extrapolate optimal data mixtures for large LLMs from small-model experiments, reducing optimization cost by 50% and improving benchmarks by up to 3%.
HDS uses Soft Actor-Critic RL with a multi-objective reward (data quality, inter-domain loss influence, weight norms) for online data mixing in LLM pre-training, reaching target perplexity with 44% fewer iterations and 7.2% MMLU gain on The Pile.
TANDEM solves bi-level data mixture optimization for LLMs via twin proxy and reference networks that measure domain efficacy by model difference and up-weight beneficial domains, with claimed theoretical guarantees and gains in restricted-data and SFT settings.
citing papers explorer
-
WARP: Weight-Space Analysis for Recovering Training Data Portfolios
WARP recovers training domain mixtures from fine-tuned model weights using weight-space interpolation via model merging to generate pseudo-checkpoints and geometric features mapped to proportions.
-
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Facility location — a classic submodular coverage objective — predicts a training subset's held-out accuracy far better than the Vendi score, which becomes misleading at high values.
-
Capacity-Aware Mixture Law Enables Efficient LLM Data Optimization
CAMEL is a scaling law capturing nonlinear model-size and mixture interactions to extrapolate optimal data mixtures for large LLMs from small-model experiments, reducing optimization cost by 50% and improving benchmarks by up to 3%.
-
Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning
HDS uses Soft Actor-Critic RL with a multi-objective reward (data quality, inter-domain loss influence, weight norms) for online data mixing in LLM pre-training, reaching target perplexity with 44% fewer iterations and 7.2% MMLU gain on The Pile.
-
TANDEM: Bi-Level Data Mixture Optimization with Twin Networks
TANDEM solves bi-level data mixture optimization for LLMs via twin proxy and reference networks that measure domain efficacy by model difference and up-weight beneficial domains, with claimed theoretical guarantees and gains in restricted-data and SFT settings.
- DataComp-VLM: Improved Open Datasets for Vision-Language Models