Pith. sign in

REVIEW 10 cited by

nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09556 v2 pith:YYD6PRNZ submitted 2024-04-15 cs.CV

nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation

classification cs.CV
keywords nnu-netsegmentationu-netvalidationarchitecturesclaimscnn-basedcurrent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The release of nnU-Net marked a paradigm shift in 3D medical image segmentation, demonstrating that a properly configured U-Net architecture could still achieve state-of-the-art results. Despite this, the pursuit of novel architectures, and the respective claims of superior performance over the U-Net baseline, continued. In this study, we demonstrate that many of these recent claims fail to hold up when scrutinized for common validation shortcomings, such as the use of inadequate baselines, insufficient datasets, and neglected computational resources. By meticulously avoiding these pitfalls, we conduct a thorough and comprehensive benchmarking of current segmentation methods including CNN-based, Transformer-based, and Mamba-based approaches. In contrast to current beliefs, we find that the recipe for state-of-the-art performance is 1) employing CNN-based U-Net models, including ResNet and ConvNeXt variants, 2) using the nnU-Net framework, and 3) scaling models to modern hardware resources. These results indicate an ongoing innovation bias towards novel architectures in the field and underscore the need for more stringent validation standards in the quest for scientific progress.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Complete virtual unwrapping and reading of a rolled Herculaneum papyrus

    eess.IV 2026-06 unverdicted novelty 7.0

    First complete digital unwrapping and reading of a Herculaneum papyrus scroll (PHerc. 1667) via synchrotron X-ray CT, virtual unrolling, and machine learning.

  2. SubsurfaceGen: Procedural Generation of Field-Scale Earth Models and Seismic Data

    cs.LG 2026-05 unverdicted novelty 7.0

    The paper introduces SubsurfaceGen, a procedural generator for field-scale 3D velocity models and seismic data, releases a dataset of 4276 2D slices from 42 models across six geological settings, and evaluates neural ...

  3. BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

    cs.CV 2026-06 unverdicted novelty 6.0

    BenchX supplies an 85k-scan benchmark that exposes poor performance of 12 tumor-detection models on underrepresented demographic and protocol subgroups.

  4. MonoUNet: A Robust Tiny Neural Network for Automated Knee Cartilage Segmentation on Point-of-Care Ultrasound Devices

    eess.IV 2026-04 conditional novelty 6.0

    MonoUNet achieves 92.62-94.82% Dice and 0.133-0.254 mm MASD on multi-device knee cartilage ultrasound while using 10-700x fewer parameters than prior lightweight models by adding a trainable monogenic phase block and ...

  5. MonoUNet: A Robust Tiny Neural Network for Automated Knee Cartilage Segmentation on Point-of-Care Ultrasound Devices

    eess.IV 2026-04 unverdicted novelty 6.0

    MonoUNet is a tiny segmentation network that achieves 92-95% Dice scores on multi-device knee cartilage ultrasound while using 10-700x fewer parameters than prior lightweight models by injecting trainable local phase ...

  6. GLOW-FDG: Generalized cancer LesiOn Whole-body segmentation model for $^{18}$F-FDG-PET/CT

    eess.IV 2026-07 accept novelty 5.0

    An open multi-cancer FDG-PET/CT lesion segmentation model trained on 1,563 scans outperforms public benchmarks on 185 external scans with fewer false positives and robust TTB/TLG.

  7. Towards Interactive Lesion Segmentation in Whole-Body PET/CT with Promptable Models

    cs.CV 2025-08 conditional novelty 5.0

    Adding click prompts encoded as Euclidean distance transforms to an nnU-Net improves interactive whole-body PET/CT lesion segmentation over Gaussian encodings and baseline autoPET III models.

  8. Advanced Tumor Segmentation in PET/CT Imaging: A Training Strategy Study with nnU-Net for AutoPET III

    cs.CV 2026-05 unverdicted novelty 4.0

    nnU-Net with ResNet encoder, intensity normalization, batch dice loss, and CraveMix augmentation reaches Dice 0.80 and third place in AutoPET III.

  9. AMO-ENE: Attention-based Multi-Omics Fusion Model for Outcome Prediction in Extra Nodal Extension and HPV-associated Oropharyngeal Cancer

    eess.IV 2026-04 unverdicted novelty 4.0

    An attention-based fusion model combining semi-supervised CT segmentation, radiomics, and clinical features predicts metastatic recurrence, overall survival, and disease-free survival in HPV+ oropharyngeal cancer with...

  10. autoPET IV challenge: Incorporating organ supervision and human guidance for lesion segmentation in PET/CT

    eess.IV 2025-09 conditional novelty 4.0

    Combining tracer classification, organ supervision, and stochastic click sampling makes an nnU-Net model segment PET/CT lesions robustly without guidance and progressively better with clicks.