Pith. sign in

REVIEW 4 cited by

Machine Learning with Multi-Site Imaging Data: An Empirical Study on the Impact of Scanner Effects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04597 v1 pith:QVACMBIA submitted 2019-10-10 eess.IV cs.CVcs.LGq-bio.NC

classification eess.IVcs.CVcs.LGq-bio.NC
keywords datalearningmachinemulti-sitebraineffectsempiricalimpact
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This is an empirical study to investigate the impact of scanner effects when using machine learning on multi-site neuroimaging data. We utilize structural T1-weighted brain MRI obtained from two different studies, Cam-CAN and UK Biobank. For the purpose of our investigation, we construct a dataset consisting of brain scans from 592 age- and sex-matched individuals, 296 subjects from each original study. Our results demonstrate that even after careful pre-processing with state-of-the-art neuroimaging pipelines a classifier can easily distinguish between the origin of the data with very high accuracy. Our analysis on the example application of sex classification suggests that current approaches to harmonize data are unable to remove scanner-specific bias leading to overly optimistic performance estimates and poor generalization. We conclude that multi-site data harmonization remains an open challenge and particular care needs to be taken when using such data with advanced machine learning methods for predictive modelling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A single whole-volume latent flow matching model, trained jointly on MRI-to-CT, CBCT-to-CT, and MRI-to-MRI tasks, matches task-specific models and gains zero-shot region generalization plus compositional translation.

  2. Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A controlled benchmark of five 3D medical encoders shows that MRI artifacts distort representation geometry differently per model and artifact, with 3DINO most stable and BrainIAC most sensitive.

  3. Benchmarking and Explaining Deep Learning Cortical Lesion MRI Segmentation in Multiple Sclerosis

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A standard nnU-Net with blob loss achieves the best cortical lesion segmentation in a multi-center benchmark, with out-of-domain F1 of 0.50.

  4. Learning Semantic Directions for Feature Augmentation in Domain-Generalized Medical Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A feature-augmentation framework with a learned channel selector and covariance-guided intensity sampler that reports the best average domain-generalized segmentation on Prostate and Fundus benchmarks.

Pith tools