Pith. sign in

REVIEW 1 cited by

Generalized Group Data Attribution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09940 v2 pith:Q6I5AHCH submitted 2024-10-13 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords attributiondatamethodsggdaapplicationsappliedcomputationallydemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data Attribution (DA) methods quantify the influence of individual training data points on model outputs and have broad applications such as explainability, data selection, and noisy label identification. However, existing DA methods are often computationally intensive, limiting their applicability to large-scale machine learning models. To address this challenge, we introduce the Generalized Group Data Attribution (GGDA) framework, which computationally simplifies DA by attributing to groups of training points instead of individual ones. GGDA is a general framework that subsumes existing attribution methods and can be applied to new DA techniques as they emerge. It allows users to optimize the trade-off between efficiency and fidelity based on their needs. Our empirical results demonstrate that GGDA applied to popular DA methods such as Influence Functions, TracIn, and TRAK results in upto 10x-50x speedups over standard DA methods while gracefully trading off attribution fidelity. For downstream applications such as dataset pruning and noisy label identification, we demonstrate that GGDA significantly improves computational efficiency and maintains effectiveness, enabling practical applications in large-scale machine learning scenarios that were previously infeasible.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A Taylor-based estimator traces a model's final behavior to individual training stages, quantifying what would change if a stage had been skipped.

Pith tools