Pith. sign in

REVIEW 13 cited by

Understanding Black-box Predictions via Influence Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.04730 v3 pith:Z3FGY6C4 submitted 2017-03-14 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG
keywords functionsinfluencemodelmodelsblack-boxevenlearningprediction
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

How can we explain the predictions of a black-box model? In this paper, we use influence functions -- a classic technique from robust statistics -- to trace a model's prediction through the learning algorithm and back to its training data, thereby identifying training points most responsible for a given prediction. To scale up influence functions to modern machine learning settings, we develop a simple, efficient implementation that requires only oracle access to gradients and Hessian-vector products. We show that even on non-convex and non-differentiable models where the theory breaks down, approximations to influence functions can still provide valuable information. On linear models and convolutional neural networks, we demonstrate that influence functions are useful for multiple purposes: understanding model behavior, debugging models, detecting dataset errors, and even creating visually-indistinguishable training-set attacks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Influence Dynamics and Stagewise Data Attribution

    cs.LG 2025-10 conditional novelty 7.0 of 10

    Using Bayesian influence functions and singular learning theory, the authors show that a sample's influence on a model varies non-monotonically over training, peaking and flipping sign at phase transitions.

  2. Implicit Reasoning Steering via Concept Chaining

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Reinforcement-learning-optimized concept-chain paragraphs covertly steer language-model multiple-choice preferences after continued pretraining, with far lower detectability than direct paraphrases.

  3. Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    RISE applies CountSketch to dual lexical and semantic channels derived from output-layer gradient outer products, cutting data attribution storage by up to 112x and enabling retrospective and prospective influence ana...

  4. GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning

    cs.LG 2026-01 reject novelty 6.0 of 10

    GUDA approximates leave-one-group-out counterfactual models with unlearning and ranks group influence by ELBO differences.

  5. Evaluating the Dynamics of Membership Privacy in Deep Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Per-sample membership vulnerability is established early in training, especially for hard-to-learn examples, and can be tracked on an FPR-TPR plane.

  6. How to Protect Models against Adversarial Unlearning?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A 'healing' procedure that uses similar real examples as surrogates for unlearned data during fine-tuning mitigates accuracy loss from machine unlearning in several classification benchmarks.

  7. Position: The Most Expensive Part of an LLM should be its Training Data

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Even at conservative wages, recreating LLM training data from scratch would cost 10 to 1000 times more than the compute and energy used to train the models.

  8. AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution

    cs.LG 2024-11 conditional novelty 6.0 of 10

    AttriBoT combines caching, hierarchical pruning, and smaller proxy models to approximate leave-one-out context attribution with a >300x speedup and little loss in faithfulness.

  9. Optimizing for Interpretability in Deep Neural Networks with Tree Regularization

    cs.LG 2019-08 conditional novelty 6.0 of 10

    Tree regularization, which penalizes the decision path length of a tree fitted to a deep network's predictions, produces deep models with higher accuracy at low complexity than L1 or L2 penalties.

  10. Regional Tree Regularization for Interpretability in Black Box Models

    cs.LG 2019-08 conditional novelty 6.0 of 10

    Regional tree regularization applies an L0-style penalty on the average decision path length of region-specific decision trees, using SparseMax to make optimization practical.

  11. When unlearning is free: leveraging low influence points to reduce computational costs

    cs.LG 2025-12 conditional novelty 5.0 of 10

    Low-influence training points can be dropped from forget/retain sets before unlearning, cutting runtime up to ~50% with little measured loss in accuracy or MIA-based privacy.

  12. Methods to Assess the UK Government's Current Role as a Data Provider for AI

    cs.CY 2024-11 reject novelty 4.0 of 10

    Using unlearning-based ablation and information-leakage tests, the paper finds UK government websites matter for LLM performance on welfare queries while data.gov.uk datasets are not recalled.

  13. New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing

    cs.CL 2024-11 conditional novelty 4.0 of 10

    The thesis shows that randomly masking input tokens during fine-tuning makes post-hoc explanations of NLP models consistently faithful under an erasure-based faithfulness metric.

Pith tools