Pith. sign in

REVIEW 3 cited by

Adaptive and Efficient Learning with Blockwise Missing and Semi-Supervised Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.18722 v4 pith:3YQ3QBOA submitted 2024-05-29 stat.ME

classification stat.ME
keywords datasourcesacrossdata-adaptivedefuseestimationfusionsemi-supervised
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Data fusion enables powerful and generalizable analyses across multiple sources. However, different data collection capacities across different sources lead to blockwise missingness (BM), which poses challenges in practice. Meanwhile, the high cost of obtaining gold-standard labels leaves the majority of samples unlabeled, known as the semi-supervised (SS) problem. In this paper, we propose a novel Data-adaptive Estimation approach for data FUsion in the SEmi-supervised setting (DEFUSE) that handles both BM and SS issues in the presence of distributional shifts across data sources under a missing at random (MAR) mechanism}. DEFUSE starts with a complete-data-only estimator derived from the primary data source, and uses data-adaptive and distributional-shift-adjusted procedures to successively incorporate the data with BM covariates and the large unlabeled sample to effectively reduce the estimation variance without incurring bias. To further avoid bias due to fusion of misaligned data violating of the MAR assumption, a screening method is developed to identify and exclude data sources that are not aligned with the primary source. Compared to existing approaches, DEFUSE offers two main improvements. First, it offers a new data-adaptive control variate approach to handle BM, which achieves intrinsic efficiency and robustness against distributional shifts. Second, it reveals a more essential role for the unlabeled sample in the BM regression problem, leading to improved estimation. These advantages are theoretically guaranteed and empirically supported by simulation studies and two real-world biomedical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal domain adaptation under label shift and blockwise missing modalities

    stat.ME 2026-08 conditional novelty 6.0 of 10

    Prediction transfers across sources with different modality sets by first reweighting each source to the target outcome distribution, then aligning all modalities to a target CCA anchor via ridge maps.

  2. Towards Efficient Inference under Nonmonotone Missingness with General Imputation

    stat.ME 2025-09 conditional novelty 6.0 of 10

    The RAY decomposition approximates the semiparametrically efficient estimator under blockwise missingness, yielding AI-powered estimators that are unbiased, asymptotically normal, and adaptively no worse than complete...

  3. Hierarchical Projection for Adaptive Knowledge Transfer

    cs.LG 2026-06 unverdicted novelty 4.0 of 10

    ProjectionTL is a hierarchical Bayesian plus adaptive projection framework that performs simultaneous source selection and feature selection to mitigate negative transfer in cross-domain learning.

Pith tools