Pith. sign in

REVIEW 16 cited by

nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09556 v2 pith:YYD6PRNZ submitted 2024-04-15 cs.CV

classification cs.CV
keywords nnu-netsegmentationu-netvalidationarchitecturesclaimscnn-basedcurrent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The release of nnU-Net marked a paradigm shift in 3D medical image segmentation, demonstrating that a properly configured U-Net architecture could still achieve state-of-the-art results. Despite this, the pursuit of novel architectures, and the respective claims of superior performance over the U-Net baseline, continued. In this study, we demonstrate that many of these recent claims fail to hold up when scrutinized for common validation shortcomings, such as the use of inadequate baselines, insufficient datasets, and neglected computational resources. By meticulously avoiding these pitfalls, we conduct a thorough and comprehensive benchmarking of current segmentation methods including CNN-based, Transformer-based, and Mamba-based approaches. In contrast to current beliefs, we find that the recipe for state-of-the-art performance is 1) employing CNN-based U-Net models, including ResNet and ConvNeXt variants, 2) using the nnU-Net framework, and 3) scaling models to modern hardware resources. These results indicate an ongoing innovation bias towards novel architectures in the field and underscore the need for more stringent validation standards in the quest for scientific progress.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 19 citations worldwide. Full citation record

  1. Learning Segmentation from Radiology Reports

    eess.IV 2025-07 conditional novelty 7.0 of 10

    R-Super converts tumor count, size, and location information from radiology reports into voxel-wise losses that improve CT tumor segmentation beyond training with masks alone.

  2. CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A new public dataset of 22,022 CT volumes labeled for 167 structures, and a nnU-Net model trained on it, outperform TotalSegmentator on most shared structures and expand coverage.

  3. ArteryX: A Reliable End-to-End Toolbox for Standardized Intracranial Artery Feature Extraction from 3D TOF-MRA

    eess.IV 2025-07 conditional novelty 6.0 of 10

    ArteryX standardizes intracranial artery feature extraction from TOF-MRA with a vessel-fused graph and constrained landmarks, showing closer agreement to reference measurements and lower manual workload than iCafe.

  4. PanTS: The Pancreatic Tumor Segmentation Dataset

    eess.IV 2025-07 conditional novelty 6.0 of 10

    PanTS is a new large CT dataset with expert-drawn pancreatic tumor and anatomy labels, and models trained on it beat prior public benchmarks.

  5. GLOW-FDG: Generalized cancer LesiOn Whole-body segmentation model for $^{18}$F-FDG-PET/CT

    eess.IV 2026-07 accept novelty 5.0 of 10

    An open multi-cancer FDG-PET/CT lesion segmentation model trained on 1,563 scans outperforms public benchmarks on 185 external scans with fewer false positives and robust TTB/TLG.

  6. Towards Interactive Lesion Segmentation in Whole-Body PET/CT with Promptable Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Adding click prompts encoded as Euclidean distance transforms to an nnU-Net improves interactive whole-body PET/CT lesion segmentation over Gaussian encodings and baseline autoPET III models.

  7. ShapeKit

    eess.IV 2025-06 reject novelty 5.0 of 10

    ShapeKit, a rule-based post-processing toolkit, reports Dice score improvements of up to 8.8 percentage points on two CT datasets without retraining the segmentation model.

  8. Explainable Anatomy-Guided AI for Prostate MRI: Foundation Models and In Silico Clinical Trials for Virtual Biopsy-based Risk Assessment

    eess.IV 2025-05 conditional novelty 5.0 of 10

    An anatomy-guided foundation-model pipeline for prostate MRI achieved AUC 0.79 and improved clinician accuracy from 0.72 to 0.77 in an in-silico trial.

  9. Unpaired Modality Translation for Pseudo Labeling of Histology Images

    eess.IV 2024-12 conditional novelty 5.0 of 10

    Unpaired image translation between labeled and unlabeled microscopy domains can produce pseudo labels useful for segmentation, achieving 0.736 mean Dice for axons on SEM via the tutoring path.

  10. autoPET IV challenge: Incorporating organ supervision and human guidance for lesion segmentation in PET/CT

    eess.IV 2025-09 conditional novelty 4.0 of 10

    Combining tracer classification, organ supervision, and stochastic click sampling makes an nnU-Net model segment PET/CT lesions robustly without guidance and progressively better with clicks.

  11. Can Diffusion Models Bridge the Domain Gap in Cardiac MR Imaging?

    cs.CV 2025-08 reject novelty 4.0 of 10

    A source-domain diffusion model with reference-guided sampling is applied to cardiac MRI domain shift, with mixed evidence: surface metrics improve on synthetic test data but the domain-generalisation claim is contrad...

  12. Pre- and Post-Treatment Glioma Segmentation with the Medical Imaging Segmentation Toolkit

    cs.CV 2025-07 conditional novelty 4.0 of 10

    On BraTS 2025 glioma segmentation, postprocessing strategies improve mean Dice/HD95 but worsen the official rank-based score, so the authors submitted the unpostprocessed baseline.

  13. BraTS orchestrator : Democratizing and Disseminating state-of-the-art brain tumor image analysis

    eess.IV 2025-06 conditional novelty 4.0 of 10

    BraTS orchestrator is a new open-source package that provides uniform, tutorial-based access to winning BraTS segmentation and synthesis algorithms for brain tumor MRI.

  14. Robust Renal Mass Segmentation on CT: A Validation Study of an AI-Based Framework

    cs.CV 2025-05 conditional novelty 4.0 of 10

    An nnU-Net model trained only on public CT data matches human-level kidney and abnormality segmentation and beats TotalSegmentator and BAMF on most tested datasets.

  15. Scaling nnU-Net for CBCT Segmentation

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A modified nnU-Net ResEnc L with larger patches, deeper topology, no left/right mirroring, and class-wise postprocessing took first place in the ToothFairy2 CBCT segmentation challenge.

  16. MRI-based Head and Neck Tumor Segmentation Using nnU-Net with 15-fold Cross-Validation Ensemble

    physics.med-ph 2024-12 conditional novelty 3.0 of 10

    Using nnU-Net V2 with a 15-fold ensemble, the authors report aggregated Dice scores of 0.81 (pre-RT) and 0.70 (mid-RT) for head and neck tumor segmentation on MRI.

Pith tools