Pith. sign in

REVIEW 1 cited by

Towards Practical Robustness Auditing for Linear Regression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.16315 v1 pith:T35IOPQD submitted 2023-07-30 stat.ME cs.LGecon.EMstat.ML

classification stat.MEcs.LGecon.EMstat.ML
keywords regressionalgorithmicproblemschallengedatasetexistencelinearmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We investigate practical algorithms to find or disprove the existence of small subsets of a dataset which, when removed, reverse the sign of a coefficient in an ordinary least squares regression involving that dataset. We empirically study the performance of well-established algorithmic techniques for this task -- mixed integer quadratically constrained optimization for general linear regression problems and exact greedy methods for special cases. We show that these methods largely outperform the state of the art and provide a useful robustness check for regression problems in a few dimensions. However, significant computational bottlenecks remain, especially for the important task of disproving the existence of such small sets of influential samples for regression problems of dimension $3$ or greater. We make some headway on this challenge via a spectral algorithm using ideas drawn from recent innovations in algorithmic robust statistics. We summarize the limitations of known techniques in several challenge datasets to encourage further algorithmic innovation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Testing Most Influential Sets

    stat.ML 2025-10 reject novelty 6.0 of 10

    Maximum influence of the most influential k-point subset in OLS follows a Fréchet distribution (heavy tails, fixed k) or Gumbel distribution (light tails or growing k), enabling tests of excessive influence.

Pith tools