REVIEW 16 cited by
STU-Net: Scalable and Transferable Medical Image Segmentation Models Empowered by Large-Scale Supervised Pre-training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
STU-Net: Scalable and Transferable Medical Image Segmentation Models Empowered by Large-Scale Supervised Pre-training
read the original abstract
Large-scale models pre-trained on large-scale datasets have profoundly advanced the development of deep learning. However, the state-of-the-art models for medical image segmentation are still small-scale, with their parameters only in the tens of millions. Further scaling them up to higher orders of magnitude is rarely explored. An overarching goal of exploring large-scale models is to train them on large-scale medical segmentation datasets for better transfer capacities. In this work, we design a series of Scalable and Transferable U-Net (STU-Net) models, with parameter sizes ranging from 14 million to 1.4 billion. Notably, the 1.4B STU-Net is the largest medical image segmentation model to date. Our STU-Net is based on nnU-Net framework due to its popularity and impressive performance. We first refine the default convolutional blocks in nnU-Net to make them scalable. Then, we empirically evaluate different scaling combinations of network depth and width, discovering that it is optimal to scale model depth and width together. We train our scalable STU-Net models on a large-scale TotalSegmentator dataset and find that increasing model size brings a stronger performance gain. This observation reveals that a large model is promising in medical image segmentation. Furthermore, we evaluate the transferability of our model on 14 downstream datasets for direct inference and 3 datasets for further fine-tuning, covering various modalities and segmentation targets. We observe good performance of our pre-trained model in both direct inference and fine-tuning. The code and pre-trained models are available at https://github.com/Ziyan-Huang/STU-Net.
Forward citations
Cited by 16 Pith papers
-
TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT
TriALS introduces a 150-case four-phase CT dataset and challenge showing top segmentation methods reach 0.754 Dice on venous phase but only 0.57 on non-contrast CT, with external validation gains up to 28%.
-
Camyla: Scaling Autonomous Research in Medical Image Segmentation
Camyla autonomously generates research proposals, experiments, and manuscripts in medical image segmentation, outperforming baselines on 24 of 31 recent datasets while producing 40 human-reviewed papers.
-
Benchmarking Deep Learning for Future Liver Remnant Segmentation in Colorectal Liver Metastasis
The first validated open benchmark for future liver remnant segmentation is created from 197 refined CT volumes, with a cascaded nnU-Net achieving the highest Dice score of 0.767.
-
CardioBench: Do Echocardiography Foundation Models Generalize Beyond the Lab?
CardioBench is a new public benchmark that standardizes eight echocardiography datasets into four regression and five classification tasks to evaluate foundation model generalization.
-
U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation
U-Mamba is a hybrid CNN-SSM architecture that outperforms prior CNN and Transformer networks on biomedical image segmentation tasks by efficiently modeling long-range dependencies.
-
BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens
Boundary-aware mixed-resolution tokens cut volumetric segmentation peak memory by >53% versus MedNeXt-L while averaging within 0.37 Dice across five nnU-Net Revisited datasets.
-
HERMES: A Hybrid Ensemble for Head-and-Neck Tumor Segmentation, TN Staging, and Recurrence-Free Survival on PET/CT
On HECKTOR 2026, mask-derived geometry features for nodal staging gave out-of-fold balanced accuracy 0.720 vs 0.691 for radiomics (paired CI includes zero); on ground-truth masks the gap was 0.897 vs 0.837.
-
BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases
BenchX supplies an 85k-scan benchmark that exposes poor performance of 12 tumor-detection models on underrepresented demographic and protocol subgroups.
-
GLeVE: Graph-Guided Lesion Grounding with Proposal Verification in 3D CT
GLeVE introduces graph-guided lesion grounding with anatomical verification and octree refinement to improve text-to-lesion alignment in 3D CT volumes.
-
Towards Brain MRI Foundation Models for the Clinic: Findings from the FOMO25 Challenge
Self-supervised pretraining on 60K clinical-style brain MRIs improves out-of-domain generalization on classification, segmentation, and regression tasks, with hybrid objectives and small models showing strong results.
-
Primus: Enforcing Attention Usage for 3D Medical Image Segmentation
Primus and PrimusV2 are Transformer-centric models that match or exceed nnU-Net and top CNNs on nine 3D medical segmentation datasets by enforcing attention usage.
-
Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement
Bifurcation Connectedness Score (BCS) measures junction-level vessel connectivity that Dice misses, tracks geometric FFR-CT decision agreement in severe disease, and shows branch recovery and tree connectedness are se...
-
Towards Brain MRI Foundation Models for the Clinic: Findings from the FOMO25 Challenge
Self-supervised pretraining on large unlabeled clinical brain MRI data improves generalization to out-of-domain clinical tasks over supervised in-domain training, with task-specific optimal objectives and limited bene...
-
LETT-NeXt: A Lightweight RECIST-Guided Model for 3D CT Lesion Segmentation
LETT-NeXt uses RECIST line prompts in a cropped MedNeXt-v2 encoder-decoder to predict 3D lesion masks, reaching DSC 73.9 on hidden test data for a CVPR 2026 segmentation competition.
-
Multi-Granularity 3D Kidney Lesion Characterization from CT Volumes
LesionDETR performs per-lesion set prediction on kidney CT volumes, reaching side-level AUC 0.799-0.817 and low per-lesion mAP, with segmentation masks and same-domain pretraining as dominant design choices.
-
Label-Efficient Cross-Modality Generalization for Liver Segmentation in Multi-Phase MRI
A fine-tuned 3D foundation segmentation model combined with cross pseudo supervision achieves robust liver segmentation across labeled and unlabeled multi-phase, multi-vendor MRI without spatial registration.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.