PODS is a plug-and-play oscillatory data-volume scheduler that alternates low-ratio regularization phases with high-ratio recovery phases to improve data selection efficiency across training tasks.
S., Daruwalla, K., and Lipasti, M
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6verdicts
UNVERDICTED 6representative citing papers
BLS approximates per-sample loss importance via EMA of batch losses, enabling simple and effective dynamic pruning of 20-50% samples losslessly across many datasets and models.
Dataset pruning is unified as a maximum-weight clique problem on a sample graph, solved greedily with an approximation guarantee and reported >40% ImageNet training-time cut without accuracy loss.
AlignPrune uses a Dynamic Alignment Score from loss trajectories to identify noisy samples more accurately than per-sample loss, improving pruning accuracy by up to 6.3% on noisy benchmarks.
Data Agent learns a co-evolving sample selection policy end-to-end that accelerates training by over 50% on ImageNet-1k and MMLU with no performance loss.
RCAP introduces class-aware probabilistic pruning that uses closed-form per-class fractions updated by loss and high-loss sampling to preserve worst-group accuracy at high pruning rates.
citing papers explorer
-
Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training
PODS is a plug-and-play oscillatory data-volume scheduler that alternates low-ratio regularization phases with high-ratio recovery phases to improve data selection efficiency across training tasks.
-
Batch Loss Score for Dynamic Data Pruning
BLS approximates per-sample loss importance via EMA of batch losses, enabling simple and effective dynamic pruning of 20-50% samples losslessly across many datasets and models.
-
Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration
Dataset pruning is unified as a maximum-weight clique problem on a sample graph, solved greedily with an approximation guarantee and reported >40% ImageNet training-time cut without accuracy loss.
-
Beyond Loss Values: Robust Dynamic Pruning via Loss Trajectory Alignment
AlignPrune uses a Dynamic Alignment Score from loss trajectories to identify noisy samples more accurately than per-sample loss, improving pruning accuracy by up to 6.3% on noisy benchmarks.
-
Data Agent: Learning to Select Data via End-to-End Dynamic Optimization
Data Agent learns a co-evolving sample selection policy end-to-end that accelerates training by over 50% on ImageNet-1k and MMLU with no performance loss.
-
RCAP: Robust, Class-Aware, Probabilistic Dynamic Dataset Pruning
RCAP introduces class-aware probabilistic pruning that uses closed-form per-class fractions updated by loss and high-loss sampling to preserve worst-group accuracy at high pruning rates.