Pith. sign in

REVIEW 3 cited by

3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.15076 v4 pith:IHA52N7F submitted 2022-09-29 cs.CV cs.LG

classification cs.CVcs.LG
keywords volumetricux-netconvnetlargemedicalmodelsegmentationtransformer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The recent 3D medical ViTs (e.g., SwinUNETR) achieve the state-of-the-art performances on several 3D volumetric data benchmarks, including 3D medical image segmentation. Hierarchical transformers (e.g., Swin Transformers) reintroduced several ConvNet priors and further enhanced the practical viability of adapting volumetric segmentation in 3D medical datasets. The effectiveness of hybrid approaches is largely credited to the large receptive field for non-local self-attention and the large number of model parameters. In this work, we propose a lightweight volumetric ConvNet, termed 3D UX-Net, which adapts the hierarchical transformer using ConvNet modules for robust volumetric segmentation. Specifically, we revisit volumetric depth-wise convolutions with large kernel size (e.g. starting from $7\times7\times7$) to enable the larger global receptive fields, inspired by Swin Transformer. We further substitute the multi-layer perceptron (MLP) in Swin Transformer blocks with pointwise depth convolutions and enhance model performances with fewer normalization and activation layers, thus reducing the number of model parameters. 3D UX-Net competes favorably with current SOTA transformers (e.g. SwinUNETR) using three challenging public datasets on volumetric brain and abdominal imaging: 1) MICCAI Challenge 2021 FLARE, 2) MICCAI Challenge 2021 FeTA, and 3) MICCAI Challenge 2022 AMOS. 3D UX-Net consistently outperforms SwinUNETR with improvement from 0.929 to 0.938 Dice (FLARE2021) and 0.867 to 0.874 Dice (Feta2021). We further evaluate the transfer learning capability of 3D UX-Net with AMOS2022 and demonstrates another improvement of $2.27\%$ Dice (from 0.880 to 0.900). The source code with our proposed model are available at https://github.com/MASILab/3DUX-Net.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Depth-routed LoRA and a depth-shift module lift frozen SAM and SAM2 to 3D and 3D+T segmentation using less than ~3.7% trainable parameters.

  2. Displacement Preserving Relational Distillation for Robust Medical Segmentation

    cs.CV 2026-07 conditional novelty 5.5 of 10

    ROI-masked pairwise displacement alignment lets a tiny nnU-Net student match or exceed MedNeXt Dice and HD95 on AMOS and ISLES with ~5% parameters.

  3. DpDNet: An Dual-Prompt-Driven Network for Universal PET-CT Segmentation

    eess.IV 2025-07 conditional novelty 5.0 of 10

    DpDNet, a dual-prompt network, achieves top average DSC (74.87%) and IoU (62.56%) on a four-cancer-type whole-body PET-CT segmentation benchmark, and its extracted biomarkers stratify breast cancer survival.

Pith tools