Pith. sign in

REVIEW 6 cited by

Assumption-Lean and Data-Adaptive Post-Prediction Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.14220 v4 pith:6G34BHD2 submitted 2023-11-23 stat.ME cs.LGstat.ML

Assumption-Lean and Data-Adaptive Post-Prediction Inference

classification stat.ME cs.LGstat.ML
keywords inferencepredictionstatisticalassumption-leandatadata-adaptivegold-standardguarantees
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

A primary challenge facing modern scientific research is the limited availability of gold-standard data which can be costly, labor-intensive, or invasive to obtain. With the rapid development of machine learning (ML), scientists can now employ ML algorithms to predict gold-standard outcomes with variables that are easier to obtain. However, these predicted outcomes are often used directly in subsequent statistical analyses, ignoring imprecision and heterogeneity introduced by the prediction procedure. This will likely result in false positive findings and invalid scientific conclusions. In this work, we introduce PoSt-Prediction Adaptive inference (PSPA) that allows valid and powerful inference based on ML-predicted data. Its "assumption-lean" property guarantees reliable statistical inference without assumptions on the ML prediction. Its "data-adaptive" feature guarantees an efficiency gain over existing methods, regardless of the accuracy of ML prediction. We demonstrate the statistical superiority and broad applicability of our method through simulations and real-data applications.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Active Statistical Inference

    stat.ML 2024-03 unverdicted novelty 7.0

    Active inference adapts label collection via ML uncertainty to deliver valid statistical inference with substantially fewer samples than standard non-adaptive methods across any data distribution.

  2. Towards Efficient Inference under Nonmonotone Missingness with General Imputation

    stat.ME 2025-09 conditional novelty 6.0

    The RAY decomposition approximates the semiparametrically efficient estimator under blockwise missingness, yielding AI-powered estimators that are unbiased, asymptotically normal, and adaptively no worse than complete...

  3. Revisiting Active Sequential Prediction-Powered Mean Estimation

    stat.ML 2026-04 unverdicted novelty 5.0

    Non-asymptotic analysis of prediction-powered mean estimation shows that no-regret learning for query probabilities converges to the maximum allowed constant value, independent of covariates.

  4. Semiparametric semi-supervised learning for general targets under distribution shift and decaying overlap

    math.ST 2025-05 unverdicted novelty 5.0

    Introduces D2S3 semiparametric framework that extends AIPW estimators to semi-supervised settings with MAR labeling, distribution shift, and decaying overlap, supplying corrected asymptotic rates instead of root-n con...

  5. High-Dimensional Statistics: Reflections on Progress and Open Problems

    math.ST 2026-05 unverdicted novelty 2.0

    A survey synthesizing representative advances, common themes, and open problems in high-dimensional statistics while pointing to key entry-point works.

  6. High-Dimensional Statistics: Reflections on Progress and Open Problems

    math.ST 2026-05 unverdicted novelty 2.0

    This review synthesizes representative advances in high-dimensional statistics, highlights common themes and open problems, and points to key entry works.