Bergson is an open-source library providing scalable data attribution tools and the first open implementations of MAGIC, SOURCE, and TrackStar for large language models.
Optimizing ml training with metagradient descent.arXiv preprint arXiv:2503.13751
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 6roles
method 1polarities
use method 1representative citing papers
Under a stability assumption, deep-learning outputs after deleting training subsets can be predicted to error ε with failure δ using only Õ(log(1/δ)/ε²) extra models and matching slowdown factors.
New analysis without global strong convexity yields tight scaling laws: NS error ~Θ(kd/n²) and NS-IF difference ~Θ((k+d)√(kd)/n²) for well-behaved logistic regressions.
NoiseRater meta-learns instance-level importance scores for noise in diffusion training via bilevel optimization, then uses a two-stage pipeline to improve efficiency and generation quality on FFHQ and ImageNet.
Kernel surrogate models with first-order gradient approximation achieve 25% higher correlation to leave-one-out ground truth for task attribution and 40% better downstream data selection than linear surrogates.
LGD reaches Bayes optimality at optimal hyperparameters and admits an O(dh) pseudo-dimension bound for meta-learning hyperparameters on convex regression tasks.
citing papers explorer
-
Bergson: An Open Source Library for Data Attribution
Bergson is an open-source library providing scalable data attribution tools and the first open implementations of MAGIC, SOURCE, and TrackStar for large language models.
-
How to sketch a learning algorithm
Under a stability assumption, deep-learning outputs after deleting training subsets can be predicted to error ε with failure δ using only Õ(log(1/δ)/ε²) extra models and matching slowdown factors.
-
On the Accuracy of Newton Step and Influence Function Data Attributions
New analysis without global strong convexity yields tight scaling laws: NS error ~Θ(kd/n²) and NS-IF difference ~Θ((k+d)√(kd)/n²) for well-behaved logistic regressions.
-
NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
NoiseRater meta-learns instance-level importance scores for noise in diffusion training via bilevel optimization, then uses a two-stage pipeline to improve efficiency and generation quality on FFHQ and ImageNet.
-
Efficient Estimation of Kernel Surrogate Models for Task Attribution
Kernel surrogate models with first-order gradient approximation achieve 25% higher correlation to leave-one-out ground truth for task attribution and 40% better downstream data selection than linear surrogates.
-
Generalization Guarantees on Data-Driven Tuning of Gradient Descent with Langevin Updates
LGD reaches Bayes optimality at optimal hyperparameters and admits an O(dh) pseudo-dimension bound for meta-learning hyperparameters on convex regression tasks.