The paper decomposes errors in trajectory-based data attribution into config, algorithm, and system levels, proposes AdamW-influence to fix optimizer mismatch, derives an error proxy for Taylor approximation, and unifies data selection under a K-step look-ahead framework.
Andrew Ilyas and Logan Engstrom
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
125 coordinated Wikipedia animal-welfare edits dominate attribution and counterfactual influence for animal-welfare queries on Llama models, with no spillover to general queries about the same entities.
Under a stability assumption, deep-learning outputs after deleting training subsets can be predicted to error ε with failure δ using only Õ(log(1/δ)/ε²) extra models and matching slowdown factors.
PRISM forms predictions as sparse mixtures of learned prototypes trained with clustering objectives, matching dense model accuracy while enabling ~500x faster data attribution and behavior editing without finetuning.
Kernel surrogate models with first-order gradient approximation achieve 25% higher correlation to leave-one-out ground truth for task attribution and 40% better downstream data selection than linear surrogates.
A science of AI requires theories of training dynamics to predict outcomes from early signals, intervene on trajectories, and design procedures that reliably produce desired capabilities, biases, robustness, and safety properties.
citing papers explorer
-
How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines
The paper decomposes errors in trajectory-based data attribution into config, algorithm, and system levels, proposes AdamW-influence to fix optimizer mismatch, derives an error proxy for Taylor approximation, and unifies data selection under a K-step look-ahead framework.
-
Small edits, large models: How Wikipedia advocacy shapes LLM values
125 coordinated Wikipedia animal-welfare edits dominate attribution and counterfactual influence for animal-welfare queries on Llama models, with no spillover to general queries about the same entities.
-
How to sketch a learning algorithm
Under a stability assumption, deep-learning outputs after deleting training subsets can be predicted to error ε with failure δ using only Õ(log(1/δ)/ε²) extra models and matching slowdown factors.
-
Prototype Language Models
PRISM forms predictions as sparse mixtures of learned prototypes trained with clustering objectives, matching dense model accuracy while enabling ~500x faster data attribution and behavior editing without finetuning.
-
Efficient Estimation of Kernel Surrogate Models for Task Attribution
Kernel surrogate models with first-order gradient approximation achieve 25% higher correlation to leave-one-out ground truth for task attribution and 40% better downstream data selection than linear surrogates.
-
Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
A science of AI requires theories of training dynamics to predict outcomes from early signals, intervene on trajectories, and design procedures that reliably produce desired capabilities, biases, robustness, and safety properties.