Pith. sign in

REVIEW 9 cited by

UNETR: Transformers for 3D Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.10504 v3 pith:MA6H67AZ submitted 2021-03-18 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords segmentationencodermedicaldecoderfcnnsimagelearningtransformers
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Fully Convolutional Neural Networks (FCNNs) with contracting and expanding paths have shown prominence for the majority of medical image segmentation applications since the past decade. In FCNNs, the encoder plays an integral role by learning both global and local features and contextual representations which can be utilized for semantic output prediction by the decoder. Despite their success, the locality of convolutional layers in FCNNs, limits the capability of learning long-range spatial dependencies. Inspired by the recent success of transformers for Natural Language Processing (NLP) in long-range sequence learning, we reformulate the task of volumetric (3D) medical image segmentation as a sequence-to-sequence prediction problem. We introduce a novel architecture, dubbed as UNEt TRansformers (UNETR), that utilizes a transformer as the encoder to learn sequence representations of the input volume and effectively capture the global multi-scale information, while also following the successful "U-shaped" network design for the encoder and decoder. The transformer encoder is directly connected to a decoder via skip connections at different resolutions to compute the final semantic segmentation output. We have validated the performance of our method on the Multi Atlas Labeling Beyond The Cranial Vault (BTCV) dataset for multi-organ segmentation and the Medical Segmentation Decathlon (MSD) dataset for brain tumor and spleen segmentation tasks. Our benchmarks demonstrate new state-of-the-art performance on the BTCV leaderboard. Code: https://monai.io/research/unetr

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Physics-Guided Radiotherapy Treatment Planning with Deep Learning

    physics.med-ph 2025-06 conditional novelty 6.0 of 10

    A two-stage deep learning pipeline that first mimics clinical VMAT plans and then uses a frozen learned dose predictor as physics guidance modestly improves dose-metric accuracy on 13 prostate patients, though several...

  2. Generalizable automated ischaemic stroke lesion segmentation with vision transformers

    eess.IV 2025-02 conditional novelty 6.0 of 10

    Swin-UNETR models trained on a large multi-site DWI stroke dataset reach high Dice scores, and a new evaluation framework exposes anatomical, morphological, and noise-dependent performance variation.

  3. Open-Source Manually Annotated Vocal Tract Database for Automatic Segmentation from 3D MRI Using Deep Learning: Benchmarking 2D and 3D Convolutional and Transformer Networks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    An open-source manually annotated 3D MRI vocal tract database for 10 French speakers enables benchmarking showing 3D U-Nets with transfer learning segment vocal tract airspace at Dice around 0.90.

  4. TS-SatFire: A Multi-Task Satellite Image Time-Series Dataset for Wildfire Detection and Prediction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    TS-SatFire is a multi-task satellite time-series dataset for detecting active fires, mapping burned areas, and predicting next-day spread, with benchmarks showing detection works but prediction remains hard.

  5. DM-SegNet: Dual-Mamba Architecture for 3D Medical Image Segmentation with Global Context Modeling

    eess.IV 2025-06 conditional novelty 5.0 of 10

    DM-SegNet combines four-direction Mamba scanning, gated spatial convolutions, and a Mamba decoder to reach reported state-of-the-art Dice scores on Synapse and BraTS2023.

  6. UNICON: UNIfied CONtinual Learning for Medical Foundational Models

    eess.IV 2025-08 unverdicted novelty 4.0 of 10

    UNICON attaches task-specific adapters (LoRA, MLP, decoder, fusion) to a frozen CT foundation model, enabling continual extension to prognosis, segmentation, and PET scans.

  7. Advancing Fetal Ultrasound Image Quality Assessment in Low-Resource Settings

    cs.CV 2025-07 conditional novelty 4.0 of 10

    LoRA fine-tuning of FetalCLIP achieves F1 0.757 for fetal ultrasound frame-quality classification, and a thresholded segmentation variant reaches F1 0.771.

  8. Missing Data Estimation for MR Spectroscopic Imaging via Mask-Free Deep Learning Methods

    eess.IV 2025-05 conditional novelty 4.0 of 10

    A mask-free U-Net framework estimates missing voxels in MRSI metabolic maps, beating interpolation and showing generalization to real patient data.

  9. Efficient Parameter Adaptation for Multi-Modal Medical Image Segmentation and Prognosis

    cs.CV 2025-04 conditional novelty 4.0 of 10

    PEMMA adapts a CT-only transformer segmentation model to CT+PET with LoRA/DoRA, reaching early-fusion-level Dice while training only 0.5-8% of parameters, and extends the same adapters to prognosis.

Pith tools