Pith. sign in

REVIEW 1 cited by

Sensitivity of MCMC-based analyses to small-data removal

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.07240 v2 pith:JPJH2ZY4 submitted 2024-08-14 stat.ME stat.CO

classification stat.MEstat.CO
keywords dataapproximationmcmcmodelsworkaccurateanalysisbayesian
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

If the conclusion of a data analysis is sensitive to dropping very few data points, that conclusion might hinge on the particular data at hand rather than representing a more broadly applicable truth. How could we check whether this sensitivity holds? One idea is to consider every small subset of data, drop it from the dataset, and re-run our analysis. But running MCMC to approximate a Bayesian posterior is already very expensive; running multiple times is prohibitive, and the number of re-runs needed here is combinatorially large. Recent work proposes a fast and accurate approximation to find the worst-case dropped data subset, but that work was developed for problems based on estimating equations -- and does not directly handle Bayesian posterior approximations using MCMC. We make two principal contributions in the present work. We adapt the existing data-dropping approximation to estimators computed via MCMC. Observing that Monte Carlo errors induce variability in the approximation, we use a variant of the bootstrap to quantify this uncertainty. We demonstrate how to use our approximation in practice to determine whether there is non-robustness in a problem. Empirically, our method is accurate in simple models, such as linear regression. In models with complicated structure, such as hierarchical models, the performance of our method is mixed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive sequential Monte Carlo for structured cross validation in Bayesian hierarchical models

    stat.CO 2025-01 conditional novelty 6.0 of 10

    Adaptive sequential Monte Carlo with automatically constructed intermediate posteriors approximates structured leave-group, leave-subset, and leave-end-out cross-validation without full MCMC reruns.

Pith tools