Pith. sign in

REVIEW 2 cited by

WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.02403 v5 pith:3A4ZQBMQ submitted 2021-11-03 eess.IV cs.CV

classification eess.IVcs.CV
keywords abdominaldatasetsegmentationorgantextitwholeannotationsclinical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Whole abdominal organ segmentation is important in diagnosing abdomen lesions, radiotherapy, and follow-up. However, oncologists' delineating all abdominal organs from 3D volumes is time-consuming and very expensive. Deep learning-based medical image segmentation has shown the potential to reduce manual delineation efforts, but it still requires a large-scale fine annotated dataset for training, and there is a lack of large-scale datasets covering the whole abdomen region with accurate and detailed annotations for the whole abdominal organ segmentation. In this work, we establish a new large-scale \textit{W}hole abdominal \textit{OR}gan \textit{D}ataset (\textit{WORD}) for algorithm research and clinical application development. This dataset contains 150 abdominal CT volumes (30495 slices). Each volume has 16 organs with fine pixel-level annotations and scribble-based sparse annotations, which may be the largest dataset with whole abdominal organ annotation. Several state-of-the-art segmentation methods are evaluated on this dataset. And we also invited three experienced oncologists to revise the model predictions to measure the gap between the deep learning method and oncologists. Afterwards, we investigate the inference-efficient learning on the WORD, as the high-resolution image requires large GPU memory and a long inference time in the test stage. We further evaluate the scribble-based annotation-efficient learning on this dataset, as the pixel-wise manual annotation is time-consuming and expensive. The work provided a new benchmark for the abdominal multi-organ segmentation task, and these experiments can serve as the baseline for future research and clinical application development.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering

    cs.CV 2025-05 conditional novelty 7.0 of 10

    DeepTumorVQA, a 9,262-volume 3D medical VQA benchmark, shows that current vision-language models handle measurement but remain far from clinical-grade lesion recognition and reasoning.

  2. PanTS: The Pancreatic Tumor Segmentation Dataset

    eess.IV 2025-07 conditional novelty 6.0 of 10

    PanTS is a new large CT dataset with expert-drawn pancreatic tumor and anatomy labels, and models trained on it beat prior public benchmarks.

Pith tools