Pith. sign in

REVIEW 2 cited by

Batch Predictive Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.13990 v5 pith:QC2TGX5W submitted 2024-09-21 stat.ME

classification stat.ME
keywords batchinferencepredictionpredictivetestdatasetscalibration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Constructing prediction sets with coverage guarantees for unobserved outcomes is a core problem in modern statistics. Methods for predictive inference have been developed for a wide range of settings, but usually only consider test data points one at a time. Here we study the problem of distribution-free predictive inference for a batch of multiple test points, aiming to construct prediction sets for functions -- such as the mean or median -- of any number of unobserved test datapoints. This setting includes constructing simultaneous prediction sets with a high probability of coverage, and selecting datapoints satisfying a specified condition while controlling the number of false claims. For the general task of predictive inference on a function of a batch of test points, we introduce a methodology called batch predictive inference (batch PI), and provide a distribution-free coverage guarantee under exchangeability of the calibration and test data. Batch PI requires the quantiles of a rank ordering function defined on certain subsets of ranks. While computing these quantiles is NP-hard in general, we show that it can be done efficiently in many cases of interest, most notably for batch score functions with a compositional structure -- which includes examples of interest such as the mean -- via a dynamic programming algorithm that we develop. Batch PI has advantages over naive approaches (such as partitioning the calibration data or directly extending conformal prediction) in many settings, as it can deliver informative prediction sets even using small calibration sample sizes. We illustrate that our procedures provide informative inference across the use cases mentioned above, through experiments on both simulated data and a drug-target interaction dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. General Synthetic-Powered Inference

    stat.ME 2025-09 conditional novelty 6.0 of 10

    GESPI combines real and synthetic data by aggregating three runs of a base inference method and guarantees an error rate of at most alpha+epsilon without any assumptions on the synthetic distribution.

  2. Semi-Supervised Risk Control via Prediction-Powered Inference

    cs.LG 2024-12 accept novelty 6.0 of 10

    Semi-supervised risk-controlling prediction sets that use unlabeled data via prediction-powered inference, with finite-sample guarantees and reduced conservatism when imputations are accurate.

Pith tools