Pith. sign in

REVIEW 10 cited by

Prediction-Powered Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.09633 v4 pith:LVEEXPWF submitted 2023-01-23 stat.ML cs.AIcs.LGq-bio.QMstat.ME

Prediction-Powered Inference

classification stat.ML cs.AIcs.LGq-bio.QMstat.ME
keywords inferenceprediction-poweredpredictionsvalidconfidenceframeworkintervalsmachine-learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Prediction-powered inference is a framework for performing valid statistical inference when an experimental dataset is supplemented with predictions from a machine-learning system. The framework yields simple algorithms for computing provably valid confidence intervals for quantities such as means, quantiles, and linear and logistic regression coefficients, without making any assumptions on the machine-learning algorithm that supplies the predictions. Furthermore, more accurate predictions translate to smaller confidence intervals. Prediction-powered inference could enable researchers to draw valid and more data-efficient conclusions using machine learning. The benefits of prediction-powered inference are demonstrated with datasets from proteomics, astronomy, genomics, remote sensing, census analysis, and ecology.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Statistical Cost of Adaptation in Multi-Source Transfer Learning

    math.ST 2026-05 unverdicted novelty 8.0

    Multi-source transfer learning incurs an intrinsic adaptation cost that can exceed one, with phase transitions separating regimes where bias-agnostic estimators match oracle performance from those where they cannot.

  2. NaiAD: Initiate Data-Driven Research for LLM Advertising

    cs.LG 2026-05 unverdicted novelty 7.0

    NaiAD is a new dataset and framework for LLM-native advertising that uses decoupled generation and calibrated scoring to identify four semantic strategies for balancing user and commercial utilities.

  3. Reasoning Gets Harder for LLMs Inside A Dialogue

    cs.CL 2026-03 unverdicted novelty 7.0

    LLMs show a consistent performance drop on arithmetic, spatial, and temporal reasoning tasks when framed in multi-turn dialogues versus isolated settings, demonstrated by the new BOULDER benchmark across eight travel-...

  4. Power-Optimal Covariate Adjustment for Switchback Experiments

    stat.ME 2026-07 conditional novelty 6.0

    Training the CUPAC covariate and its regression coefficient with a between-cell-weighted loss improves switchback estimator power, with gains concentrated in within-noise-dominated regimes.

  5. Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

    cs.LG 2026-05 unverdicted novelty 6.0

    Counterfactual metrics on semi-simulated benchmarks fail to identify the treatment effect estimators preferred by observable metrics on real datasets, with simple meta-learners outperforming specialized causal models.

  6. Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation

    cs.LG 2026-05 unverdicted novelty 6.0

    Large-scale study finds that counterfactual metrics on semi-simulated data do not select the same estimators as observable metrics on real data, and benchmark rankings fail to transfer.

  7. Response Time Enhances Alignment with Heterogeneous Preferences

    cs.LG 2026-05 unverdicted novelty 6.0

    Response times modeled as drift-diffusion processes enable consistent estimation of population-average preferences from heterogeneous anonymous binary choices.

  8. PPI++: Efficient Prediction-Powered Inference

    stat.ML 2023-11 unverdicted novelty 6.0

    PPI++ yields easy-to-compute confidence sets for any-dimensional parameters that always improve on classical intervals from labeled data alone by leveraging abundant ML predictions.

  9. Efficient Sequential Evaluation of Large Language Models

    stat.ML 2026-07 conditional novelty 5.0

    A confidence-sequence framework for sequentially estimating an LLM's average benchmark accuracy under adaptive question selection, with growth-oriented sampling rules that in practice often lose to uniform sampling.

  10. Towards Auditing AI Systems in the Wild

    cs.CY 2026-06 unverdicted novelty 4.0

    Proposes framing auditing of deployed AI systems as continuous statistical monitoring of risk-controlled constraints like fairness and safety under uncertainty.