Pith. sign in

REVIEW 2 cited by

Machine Learning with Multi-Site Imaging Data: An Empirical Study on the Impact of Scanner Effects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04597 v1 pith:QVACMBIA submitted 2019-10-10 eess.IV cs.CVcs.LGq-bio.NC

classification eess.IVcs.CVcs.LGq-bio.NC
keywords datalearningmachinemulti-sitebraineffectsempiricalimpact
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This is an empirical study to investigate the impact of scanner effects when using machine learning on multi-site neuroimaging data. We utilize structural T1-weighted brain MRI obtained from two different studies, Cam-CAN and UK Biobank. For the purpose of our investigation, we construct a dataset consisting of brain scans from 592 age- and sex-matched individuals, 296 subjects from each original study. Our results demonstrate that even after careful pre-processing with state-of-the-art neuroimaging pipelines a classifier can easily distinguish between the origin of the data with very high accuracy. Our analysis on the example application of sex classification suggests that current approaches to harmonize data are unable to remove scanner-specific bias leading to overly optimistic performance estimates and poor generalization. We conclude that multi-site data harmonization remains an open challenge and particular care needs to be taken when using such data with advanced machine learning methods for predictive modelling.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking and Explaining Deep Learning Cortical Lesion MRI Segmentation in Multiple Sclerosis

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A standard nnU-Net with blob loss achieves the best cortical lesion segmentation in a multi-center benchmark, with out-of-domain F1 of 0.50.

  2. Learning Semantic Directions for Feature Augmentation in Domain-Generalized Medical Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A feature-augmentation framework with a learned channel selector and covariance-guided intensity sampler that reports the best average domain-generalized segmentation on Prostate and Fundus benchmarks.

Pith tools