Pith. sign in

REVIEW 1 cited by

Approximations to worst-case data dropping: unmasking failure modes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.09008 v5 pith:4JL2YXT6 submitted 2024-08-16 stat.ME stat.CO

classification stat.MEstat.CO
keywords dataapproximationsnon-robustnesssimpleworkalgorithmdetectdropping
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A data analyst might worry about generalization if dropping a very small fraction of data points from a study could change its substantive conclusions. Checking this non-robustness directly poses a combinatorial optimization problem and is intractable even for simple models and moderate data sizes. Recently various authors have proposed a diverse set of approximations to detect this non-robustness. In the present work, we show that, even in a setting as simple as ordinary least squares (OLS) linear regression, many of these approximations can fail to detect (true) non-robustness in realistic data arrangements. We focus on OLS in the present work due its widespread use and since some approximations work only for OLS. Across our synthetic and real-world data sets, we find that a simple recursive greedy algorithm is the sole algorithm that does not fail any of our tests and also that it can be orders of magnitude faster to run than some competitors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive sequential Monte Carlo for structured cross validation in Bayesian hierarchical models

    stat.CO 2025-01 conditional novelty 6.0 of 10

    Adaptive sequential Monte Carlo with automatically constructed intermediate posteriors approximates structured leave-group, leave-subset, and leave-end-out cross-validation without full MCMC reruns.

Pith tools