REVIEW 3 cited by
Bayesian Prediction-Powered Inference
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Prediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. Specifically, PPI methods provide tighter confidence intervals by combining small amounts of human-labeled data with larger amounts of data labeled by a reasonably accurate, but potentially biased, automatic system. We propose a framework for PPI based on Bayesian inference that allows researchers to develop new task-appropriate PPI methods easily. Exploiting the ease with which we can design new metrics, we propose improved PPI methods for several importantcases, such as autoraters that give discrete responses (e.g., prompted LLM ``judges'') and autoraters with scores that have a non-linear relationship to human scores.
Forward citations
Cited by 3 Pith papers
-
QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model Predictions
QuEst gives point estimates and asymptotic confidence intervals for quantile-based distributional measures by optimally combining scarce observed data with abundant model-imputed data.
-
FAB-PPI: Frequentist, Assisted by Bayes, Prediction-Powered Inference
FAB-PPI applies frequentist-assisted-by-Bayes confidence regions to the PPI rectifier, yielding shorter asymptotic confidence intervals when predictions are accurate and reverting to standard PPI under a horseshoe pri...
-
Predictions as Surrogates: Revisiting Surrogate Outcomes in the Age of AI
Recalibrated prediction-powered inference estimates the optimal imputed loss by cross-fitted machine learning, always improving on label-only inference and matching the best possible variance among PPI estimators when...
Discussion (0). Continue with ORCID to comment.