Pith. sign in

REVIEW 6 cited by

Computer-Vision Benchmark Segment-Anything Model (SAM) in Medical Images: Accuracy in 12 Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09324 v3 pith:WQFYIL6T submitted 2023-04-18 eess.IV cs.CV

Computer-Vision Benchmark Segment-Anything Model (SAM) in Medical Images: Accuracy in 12 Datasets

classification eess.IV cs.CV
keywords imagesegmentationmedicalaccuracydiceimageswerecontrast
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Background: The segment-anything model (SAM), introduced in April 2023, shows promise as a benchmark model and a universal solution to segment various natural images. It comes without previously-required re-training or fine-tuning specific to each new dataset. Purpose: To test SAM's accuracy in various medical image segmentation tasks and investigate potential factors that may affect its accuracy in medical images. Methods: SAM was tested on 12 public medical image segmentation datasets involving 7,451 subjects. The accuracy was measured by the Dice overlap between the algorithm-segmented and ground-truth masks. SAM was compared with five state-of-the-art algorithms specifically designed for medical image segmentation tasks. Associations of SAM's accuracy with six factors were computed, independently and jointly, including segmentation difficulties as measured by segmentation ability score and by Dice overlap in U-Net, image dimension, size of the target region, image modality, and contrast. Results: The Dice overlaps from SAM were significantly lower than the five medical-image-based algorithms in all 12 medical image segmentation datasets, by a margin of 0.1-0.5 and even 0.6-0.7 Dice. SAM-Semantic was significantly associated with medical image segmentation difficulty and the image modality, and SAM-Point and SAM-Box were significantly associated with image segmentation difficulty, image dimension, target region size, and target-vs-background contrast. All these 3 variations of SAM were more accurate in 2D medical images, larger target region sizes, easier cases with a higher Segmentation Ability score and higher U-Net Dice, and higher foreground-background contrast.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models

    cs.CV 2026-04 unverdicted novelty 7.0

    SAM-family models split into occluder-aware types that avoid predicting into occluded regions and occluder-agnostic types that confidently segment hidden areas, shown via a new benchmark on polyp datasets.

  2. MorVess: Morphology-Aware Pulmonary Vessel Segmentation Network

    cs.CV 2026-06 unverdicted novelty 5.0

    MorVess improves pulmonary vessel segmentation by jointly predicting vessel masks, distance maps, and thickness maps using a 2.5D SAM adapter and global-local fusion for better small-vessel recovery and connectivity.

  3. CardioSAM: Topology-Aware Decoder Design for High-Precision Cardiac MRI Segmentation

    cs.CV 2026-03 accept novelty 5.0

    CardioSAM introduces a topology-aware decoder and boundary refinement module to elevate SAM's performance on cardiac MRI, achieving a 93.39% Dice score.

  4. Clinical utility of foundation models in musculoskeletal MRI for biomarker fidelity and predictive outcomes

    eess.IV 2025-01 unverdicted novelty 4.0

    Fine-tuned foundation models produce reliable MSK MRI biomarkers that support workload-reducing triage and calibrated 48-month prediction of knee replacement and incident OA.

  5. CellNet -- Localizing Cells using Sparse and Noisy Point Annotations

    cs.CV 2026-06 unverdicted novelty 3.0

    CellNet applies regression-based deep learning to count cells from sparse point annotations in microscopy images and claims better performance than zero-shot methods in low-data regimes.

  6. Data-Centric Foundation Models in Computational Healthcare: A Survey

    cs.LG 2024-01 unverdicted novelty 3.0

    The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.