Pith. sign in

REVIEW 6 cited by

How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.11253 v1 pith:5ZTV26I2 submitted 2025-01-20 eess.IV cs.CV

classification eess.IVcs.CV
keywords modelstransferimagenetlearningmodelpre-trainedpre-trainingdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The pre-training and fine-tuning paradigm has become prominent in transfer learning. For example, if the model is pre-trained on ImageNet and then fine-tuned to PASCAL, it can significantly outperform that trained on PASCAL from scratch. While ImageNet pre-training has shown enormous success, it is formed in 2D, and the learned features are for classification tasks; when transferring to more diverse tasks, like 3D image segmentation, its performance is inevitably compromised due to the deviation from the original ImageNet context. A significant challenge lies in the lack of large, annotated 3D datasets rivaling the scale of ImageNet for model pre-training. To overcome this challenge, we make two contributions. Firstly, we construct AbdomenAtlas 1.1 that comprises 9,262 three-dimensional computed tomography (CT) volumes with high-quality, per-voxel annotations of 25 anatomical structures and pseudo annotations of seven tumor types. Secondly, we develop a suite of models that are pre-trained on our AbdomenAtlas 1.1 for transfer learning. Our preliminary analyses indicate that the model trained only with 21 CT volumes, 672 masks, and 40 GPU hours has a transfer learning ability similar to the model trained with 5,050 (unlabeled) CT volumes and 1,152 GPU hours. More importantly, the transfer learning ability of supervised models can further scale up with larger annotated datasets, achieving significantly better performance than preexisting pre-trained models, irrespective of their pre-training methodologies or data sources. We hope this study can facilitate collective efforts in constructing larger 3D medical datasets and more releases of supervised pre-trained models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Segmentation Pre-training for Label-Efficient Lumbar Spine Degeneration Grading

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Segmentation pre-training on automatically generated MRI masks lets a spine-grading model reach near full-supervision performance with only 20% of manual grading labels.

  2. Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Existing 3D CT foundation models generalize poorly to external head and neck cancer cohorts; CT-CLIP was the most robust, and gating fusion of imaging with clinical data gave the best prediction.

  3. Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Source-dataset selection for medical transfer learning is driven by community practice and perceived similarity, and 'more similar is better' does not consistently hold.

  4. CoralBay: A Self-Supervised CT Foundation Model

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    CoralBay extends DINO self-distillation to 3D CT using hierarchical Swin transformers on concatenated multi-scale features, claiming effective transfer to radiological tasks plus a new public leaderboard.

  5. Benchmarking transferability of SSL pretraining to same and different modality segmentation tasks

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    SMIT, which combines masked image modeling with self-distillation, delivers the highest segmentation accuracy, fastest convergence, and best few-shot performance across nine CT and MRI tasks compared to contrastive an...

  6. Adapting Medical Vision Foundation Models for Volumetric Medical Image Segmentation via Active Learning and Selective Semi-supervised Fine-tuning

    eess.IV 2025-09 unverdicted novelty 5.0 of 10

    ASSFT combines active test-time sample selection via diversified knowledge divergence and anatomical segmentation difficulty with selective semi-supervised fine-tuning to adapt medical vision foundation models for vol...

Pith tools