Pith. sign in

REVIEW 3 cited by

Most Influential Subset Selection: Challenges, Promises, and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18153 v2 pith:EEVXY4RG submitted 2024-09-25 cs.LG stat.ML

classification cs.LGstat.ML
keywords influencesamplescollectivemisssubsetanalysiscapturecomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How can we attribute the behaviors of machine learning models to their training data? While the classic influence function sheds light on the impact of individual samples, it often fails to capture the more complex and pronounced collective influence of a set of samples. To tackle this challenge, we study the Most Influential Subset Selection (MISS) problem, which aims to identify a subset of training samples with the greatest collective influence. We conduct a comprehensive analysis of the prevailing approaches in MISS, elucidating their strengths and weaknesses. Our findings reveal that influence-based greedy heuristics, a dominant class of algorithms in MISS, can provably fail even in linear regression. We delineate the failure modes, including the errors of influence function and the non-additive structure of the collective influence. Conversely, we demonstrate that an adaptive version of these heuristics which applies them iteratively, can effectively capture the interactions among samples and thus partially address the issues. Experiments on real-world datasets corroborate these theoretical findings and further demonstrate that the merit of adaptivity can extend to more complex scenarios such as classification tasks and non-linear neural networks. We conclude our analysis by emphasizing the inherent trade-off between performance and computational efficiency, questioning the use of additive metrics such as the Linear Datamodeling Score, and offering a range of discussions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Finding Most Influential Sets

    stat.ML 2026-06 unverdicted novelty 7.0 of 10

    For linear-fractional leave-set-out estimands, most influential set selection reduces to a one-parameter sequence of top-k problems solved efficiently by Dinkelbach's algorithm with global optimality for fixed residuals.

  2. Testing Most Influential Sets

    stat.ML 2025-10 reject novelty 6.0 of 10

    Maximum influence of the most influential k-point subset in OLS follows a Fréchet distribution (heavy tails, fixed k) or Gumbel distribution (light tails or growing k), enabling tests of excessive influence.

  3. Better Training Data Attribution via Better Inverse Hessian-Vector Products

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ASTRA, an EKFAC-preconditioned Neumann series iteration, computes more accurate inverse Hessian-vector products and improves training data attribution scores over EKFAC baselines.

Pith tools