Pith. sign in

REVIEW 1 cited by

Provably Auditing Ordinary Least Squares in Low Dimensions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14284 v2 pith:I42C34UN submitted 2022-05-28 stat.ML cs.DScs.LGecon.EM

classification stat.MLcs.DScs.LGecon.EM
keywords stabilityheuristicmetricnumberonlysamplesalgorithmsanalyses
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Measuring the stability of conclusions derived from Ordinary Least Squares linear regression is critically important, but most metrics either only measure local stability (i.e. against infinitesimal changes in the data), or are only interpretable under statistical assumptions. Recent work proposes a simple, global, finite-sample stability metric: the minimum number of samples that need to be removed so that rerunning the analysis overturns the conclusion, specifically meaning that the sign of a particular coefficient of the estimated regressor changes. However, besides the trivial exponential-time algorithm, the only approach for computing this metric is a greedy heuristic that lacks provable guarantees under reasonable, verifiable assumptions; the heuristic provides a loose upper bound on the stability and also cannot certify lower bounds on it. We show that in the low-dimensional regime where the number of covariates is a constant but the number of samples is large, there are efficient algorithms for provably estimating (a fractional version of) this metric. Applying our algorithms to the Boston Housing dataset, we exhibit regression analyses where we can estimate the stability up to a factor of $3$ better than the greedy heuristic, and analyses where we can certify stability to dropping even a majority of the samples.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Testing Most Influential Sets

    stat.ML 2025-10 reject novelty 6.0 of 10

    Maximum influence of the most influential k-point subset in OLS follows a Fréchet distribution (heavy tails, fixed k) or Gumbel distribution (light tails or growing k), enabling tests of excessive influence.

Pith tools