Pith. sign in

REVIEW

Merging versus Ensembling in Multi-Study Prediction: Theoretical Insight from Random Effects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.07382 v4 pith:LAWCFQMW submitted 2019-05-17 stat.ML cs.LG

classification stat.MLcs.LG
keywords ensemblingmergingpointstudieslearnermulti-studypredictiontraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A critical decision point when training predictors using multiple studies is whether studies should be combined or treated separately. We compare two multi-study prediction approaches in the presence of potential heterogeneity in predictor-outcome relationships across datasets: 1) merging all of the datasets and training a single learner, and 2) multi-study ensembling, which involves training a separate learner on each dataset and combining the predictions resulting from each learner. For ridge regression, we show analytically and confirm via simulation that merging yields lower prediction error than ensembling when the predictor-outcome relationships are relatively homogeneous across studies. However, as cross-study heterogeneity increases, there exists a transition point beyond which ensembling outperforms merging. We provide analytic expressions for the transition point in various scenarios, study asymptotic properties, and illustrate how transition point theory can be used for deciding when studies should be combined with an application from metagenomics.

Discussion (0). Sign in to comment.

Pith tools