Pith. sign in

REVIEW 5 cited by

SegVol: Universal and Interactive Volumetric Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.13385 v5 pith:5SKX7DTZ submitted 2023-11-22 cs.CV

classification cs.CV
keywords segmentationimagemodelfoundationmedicalsegvolvolumetricanatomical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Precise image segmentation provides clinical study with instructive information. Despite the remarkable progress achieved in medical image segmentation, there is still an absence of a 3D foundation segmentation model that can segment a wide range of anatomical categories with easy user interaction. In this paper, we propose a 3D foundation segmentation model, named SegVol, supporting universal and interactive volumetric medical image segmentation. By scaling up training data to 90K unlabeled Computed Tomography (CT) volumes and 6K labeled CT volumes, this foundation model supports the segmentation of over 200 anatomical categories using semantic and spatial prompts. To facilitate efficient and precise inference on volumetric images, we design a zoom-out-zoom-in mechanism. Extensive experiments on 22 anatomical segmentation tasks verify that SegVol outperforms the competitors in 19 tasks, with improvements up to 37.24% compared to the runner-up methods. We demonstrate the effectiveness and importance of specific designs by ablation study. We expect this foundation model can promote the development of volumetric medical image analysis. The model and code are publicly available at: https://github.com/BAAI-DCAI/SegVol.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ORCA compresses 3D CT tokens into organ-guided connected regions with sinusoidal centroid encoding, outperforming grid average and other compressors at matched budgets.

  2. Rethink Domain Generalization in Heterogeneous Sequence MRI Segmentation

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A semi-supervised pretraining method improves cross-sequence pancreas segmentation Dice from 43.55% to 70.39% (NU) and from 35.62% to 66.61% (IH) on the new PancreasDG benchmark.

  3. LETT-NeXt: A Lightweight RECIST-Guided Model for 3D CT Lesion Segmentation

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    LETT-NeXt uses RECIST line prompts in a cropped MedNeXt-v2 encoder-decoder to predict 3D lesion masks, reaching DSC 73.9 on hidden test data for a CVPR 2026 segmentation competition.

  4. RAPS-3D: Efficient interactive segmentation for 3D radiological imaging

    cs.CV 2025-07 conditional novelty 4.0 of 10

    RAPS-3D is a 3D promptable CT segmentation model that reports 86.8 Dice on AMOS-CT with a single 2D bounding-box prompt, using zoom-out/zoom-in inference with no sliding window.

  5. Prompt Mechanisms in Medical Imaging: A Comprehensive Survey

    eess.IV 2025-06 conditional novelty 4.0 of 10

    A broad survey that organizes prompt mechanisms for medical image generation, segmentation, and classification into a two-dimensional taxonomy of core technologies and clinical applications.

Pith tools