Pith. sign in

REVIEW 16 cited by

AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.08023 v3 pith:F6SDM2H7 submitted 2022-06-16 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords segmentationabdominalbenchmarkmodelsamosdiverselarge-scalemedical
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting and labeling 3D medical data, most of the deep learning models to date are driven by datasets with a limited number of organs of interest or samples, which still limits the power of modern deep models and makes it difficult to provide a fully comprehensive and fair estimate of various methods. To mitigate the limitations, we present AMOS, a large-scale, diverse, clinical dataset for abdominal organ segmentation. AMOS provides 500 CT and 100 MRI scans collected from multi-center, multi-vendor, multi-modality, multi-phase, multi-disease patients, each with voxel-level annotations of 15 abdominal organs, providing challenging examples and test-bed for studying robust segmentation algorithms under diverse targets and scenarios. We further benchmark several state-of-the-art medical segmentation models to evaluate the status of the existing methods on this new challenging dataset. We have made our datasets, benchmark servers, and baselines publicly available, and hope to inspire future research. Information can be found at https://amos22.grand-challenge.org.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Segmentation from Radiology Reports

    eess.IV 2025-07 conditional novelty 7.0 of 10

    R-Super converts tumor count, size, and location information from radiology reports into voxel-wise losses that improve CT tumor segmentation beyond training with masks alone.

  2. HyperSORT: Self-Organising Robust Training with hyper-networks

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A hyper-network that predicts segmentation UNet weights from per-sample learned latent vectors yields a structured map of annotation styles and a way to flag erroneous labels.

  3. DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new five-level medical imaging benchmark, DrVD-Bench, shows that vision-language models lose accuracy sharply as reasoning complexity grows and often diagnose without grounding in lesion evidence.

  4. MultiverSeg: Scalable Interactive Segmentation of Biomedical Imaging Datasets with In-Context Guidance

    cs.CV 2024-12 conditional novelty 7.0 of 10

    MultiverSeg combines interactive prompting with a growing set of previously segmented image pairs to reduce the number of user interactions needed to segment a new biomedical dataset.

  5. Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Curia-MAE is a multi-modal, multi-anatomy masked autoencoder whose frozen encoder modestly improves 3D segmentation over its MAE baseline, with the largest gains on lesion tasks.

  6. SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    SegMoTE shows that adding token-level mixture-of-experts routing to a frozen SAM decoder can match or beat medical-segmentation models trained on far more data, using 0.15M curated masks and 17M trainable parameters.

  7. Large-scale Multi-sequence Pretraining for Generalizable MRI Analysis in Versatile Clinical Applications

    eess.IV 2025-08 conditional novelty 6.0 of 10

    A four-objective self-supervised pretraining recipe on a 336k-volume multi-sequence MRI corpus yields first-rank transfer on 39 of 44 downstream MRI tasks.

  8. Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Training on synthetic compositional task sequences with sequence-level masking lets a transformer-based in-context learner follow multi-step medical imaging instructions on held-out images, but well below codebook upp...

  9. RadSAM: Segmenting 3D radiological images with a 2D promptable model

    cs.CV 2025-04 conditional novelty 6.0 of 10

    RadSAM segments 3D organs in CT from a single point or box prompt by iteratively propagating the predicted mask as a prompt to neighboring slices, outperforming MedSAM and matching or beating nnU-Net on AMOS.

  10. Leveraging Textual Anatomical Knowledge for Class-Imbalanced Semi-Supervised Multi-Organ Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Injecting GPT-4o-generated textual anatomical priors as segmentation-head parameters improves class-imbalanced semi-supervised multi-organ segmentation.

  11. VOILA: Complexity-Aware Universal Segmentation of CT images by Voxel Interacting with Language

    cs.CV 2025-01 conditional novelty 6.0 of 10

    VOILA performs universal CT segmentation by contrastively aligning voxels with text prompts and training on complexity-graded samples, achieving competitive Dice scores with far fewer trainable parameters.

  12. GIRAFE: Glottal Imaging Dataset for Advanced Segmentation, Analysis, and Facilitative Playbacks Evaluation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GIRAFE releases 65 color high-speed laryngeal videos, 760 manual glottal gap masks, automatic segmentation baselines, and facilitative playbacks.

  13. Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Label quality matters little when pre-training medical segmentation models, but still matters for in-domain deployment; only large quality gaps affect transfer results.

  14. The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology

    cs.AI 2026-07 conditional novelty 4.0 of 10

    A formal orchestration framework for oncology AI pipelines is proposed, demonstrating that routing logic and output schema remain invariant under model substitution via a proof-of-concept with synthetic stubs.

  15. Diffusion-empowered AutoPrompt MedSAM

    eess.IV 2025-02 conditional novelty 4.0 of 10

    A class-index-driven diffusion-style prompt encoder turns MedSAM into a fully automatic segmenter that outputs semantically labeled masks, with reported gains on CT, MRI, endoscopy, and X-ray benchmarks.

  16. A Unified Framework for Foreground and Anonymization Area Segmentation in CT and MRI Data

    eess.IV 2025-01 conditional novelty 4.0 of 10

    A nnU-Net-based toolkit segments body foreground and anonymized regions in 3D CT/MRI with high Dice scores.

Pith tools