REVIEW 17 cited by
A Unified Framework for Semiparametrically Efficient Semi-Supervised Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
A Unified Framework for Semiparametrically Efficient Semi-Supervised Learning
read the original abstract
We consider statistical inference under a semi-supervised setting where we have access to both a labeled dataset consisting of pairs $\{X_i, Y_i \}_{i=1}^n$ and an unlabeled dataset $\{ X_i \}_{i=n+1}^{n+N}$. We ask the question: under what circumstances, and by how much, can incorporating the unlabeled dataset improve upon inference using the labeled data? To answer this question, we investigate semi-supervised learning through the lens of semiparametric efficiency theory. We characterize the efficiency lower bound under the semi-supervised setting for an arbitrary inferential problem, and show that incorporating unlabeled data can potentially improve efficiency if the parameter is not well-specified. We then propose two types of semi-supervised estimators: a safe estimator that imposes minimal assumptions, is simple to compute, and is guaranteed to be at least as efficient as the initial supervised estimator; and an efficient estimator, which -- under stronger assumptions -- achieves the semiparametric efficiency bound. Our findings unify existing semiparametric efficiency results for particular special cases, and extend these results to a much more general class of problems. Moreover, we show that our estimators can flexibly incorporate predicted outcomes arising from ``black-box" machine learning models, and thereby achieve the same goal as prediction-powered inference (PPI), but with superior theoretical guarantees. We also provide a complete understanding of the theoretical basis for the existing set of PPI methods. Finally, we apply the theoretical framework developed to derive and analyze efficient semi-supervised estimators in a number of settings, including M-estimation, U-statistics, and average treatment effect estimation, and demonstrate the performance of the proposed estimators via simulations.
Forward citations
Cited by 17 Pith papers
-
Semiparametric Mediation Analysis with Separately Observed Mediator and Outcome under Unmeasured Confounding
A semiparametric data fusion framework restores identification of mediation effects from separately observed mediator and outcome data using shared IVs under unmeasured confounding and no-interaction plus latent align...
-
Prediction-Powered Linear Regression: A Balance Between Interpretation and Prediction
PUMA uses model averaging to jointly handle uncertainties from model misspecification, tuning, and ML choice, delivering asymptotic in-sample and out-of-sample prediction optimality plus estimation consistency.
-
Prediction-powered Inference by Mixture of Experts
An MOE-powered PPI framework adaptively blends multiple predictors to achieve minimal variance and a best-expert guarantee for semi-supervised mean estimation, linear regression, quantile estimation, and M-estimation,...
-
Calibeating Prediction-Powered Inference
Post-hoc calibration of miscalibrated black-box predictions on a labeled sample improves efficiency of prediction-powered inference for semisupervised mean estimation.
-
Statistical Optimality of Prediction-Powered Inference
PPI is consistent and asymptotically normal, attaining the semiparametric efficiency lower bound under score-calibration of the predictor, with theory for cross-fitting and variance correction.
-
Optimized Labeling Resource Allocation for Prediction-Assisted Inference via OPAL
OPAL learns optimal smooth labeling policies from ML uncertainty scores to enable low-variance prediction-assisted inference with finite-sample coverage guarantees.
-
Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety
Introduces a regularized estimator achieving optimal MSE rates under a new relative balancedness condition while providing safety guarantees that match independent learning when tasks are unrelated.
-
Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety
Matrix-weighted regularization for robust multi-task regression achieves optimal MSE under weaker spectral assumptions and performs no worse than independent learning when balancedness is poor.
-
Transporting treatment effects by calibrating large-scale observational outcomes
Proposes a calibration-based estimator for transported average treatment effects that is consistent under correct specification and achieves semiparametric efficiency with large observational data.
-
Transporting treatment effects by calibrating large-scale observational outcomes
A calibration procedure yields a weighted transported average treatment effect with asymptotically valid and efficient inference when experimental data grows slower than observational data, even without positivity or ...
-
A Functional-Class Meta-Analytic Framework for Quantifying Surrogate Resilience
A meta-analytic framework estimates the resilience probability of a surrogate marker to the surrogate paradox in a new study by modeling deviations from functional relationships observed in completed trials.
-
Revisiting Active Sequential Prediction-Powered Mean Estimation
Non-asymptotic analysis of prediction-powered mean estimation shows that no-regret learning for query probabilities converges to the maximum allowed constant value, independent of covariates.
-
Semiparametric semi-supervised learning for general targets under distribution shift and decaying overlap
Introduces D2S3 semiparametric framework that extends AIPW estimators to semi-supervised settings with MAR labeling, distribution shift, and decaying overlap, supplying corrected asymptotic rates instead of root-n con...
-
Robust Estimation and Inference with Selective Borrowing in Hybrid Controlled Trials: A Tutorial with SelectiveIntegrative and intFRT
Tutorial on a statistical roadmap and R packages for selective borrowing in hybrid controlled trials, demonstrated on synthetic lung cancer data.
-
High-Dimensional Statistics: Reflections on Progress and Open Problems
This review synthesizes representative advances in high-dimensional statistics, highlights common themes and open problems, and points to key entry works.
-
High-Dimensional Statistics: Reflections on Progress and Open Problems
A survey synthesizing representative advances, common themes, and open problems in high-dimensional statistics while pointing to key entry-point works.
-
Externally Controlled Trials: A Review of Design and Borrowing Through a Causal Lens
A review organizes externally controlled trial methodology through causal estimands and identifiability assumptions for single-arm and hybrid designs with borrowing strategies.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.