REVIEW 18 cited by
MedSAM2: Segment Anything in 3D Medical Images and Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Medical image and video segmentation is a critical task for precision medicine, which has witnessed considerable progress in developing task or modality-specific and generalist models for 2D images. However, there have been limited studies on building general-purpose models for 3D images and videos with comprehensive user studies. Here, we present MedSAM2, a promptable segmentation foundation model for 3D image and video segmentation. The model is developed by fine-tuning the Segment Anything Model 2 on a large medical dataset with over 455,000 3D image-mask pairs and 76,000 frames, outperforming previous models across a wide range of organs, lesions, and imaging modalities. Furthermore, we implement a human-in-the-loop pipeline to facilitate the creation of large-scale datasets resulting in, to the best of our knowledge, the most extensive user study to date, involving the annotation of 5,000 CT lesions, 3,984 liver MRI lesions, and 251,550 echocardiogram video frames, demonstrating that MedSAM2 can reduce manual costs by more than 85%. MedSAM2 is also integrated into widely used platforms with user-friendly interfaces for local and cloud deployment, making it a practical tool for supporting efficient, scalable, and high-quality segmentation in both research and healthcare environments.
Forward citations
Cited by 18 Pith papers
-
DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
A new five-level medical imaging benchmark, DrVD-Bench, shows that vision-language models lose accuracy sharply as reasoning complexity grows and often diagnose without grounding in lesion evidence.
-
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
An agentic, clinician-editable evidence workflow for thyroid ultrasound matches or beats specialist baselines on segmentation and classification benchmarks and improves report consistency and speed in reader studies.
-
MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation
A unified medical pixel-language model trained on 440K synthesized mask-language samples achieves strong performance on reasoning and explanatory segmentation, with zero-shot transfer to external grounding benchmarks.
-
NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
A prompt-based multi-agent loop combined with an instruction-guided 3D U-Net editor improves neuron segmentation topology and beats state-of-the-art on BigNeuron, CWMBS, and ZBFWB.
-
ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation
An LLM-agent artwork annotation system that combines proactive label suggestions with interaction-driven skill learning reported roughly 50% faster annotation and higher label agreement in a 12-participant study.
-
SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting
Depth-routed LoRA and a depth-shift module lift frozen SAM and SAM2 to 3D and 3D+T segmentation using less than ~3.7% trainable parameters.
-
Do Medical Foundation Models Generalize on the African Brain?
Medical foundation models show no consistent generalization gap on African brain MRI; performance differences track dataset size, not data origin.
-
Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
An atlas-registration pipeline that prompts and fuses a frozen segmentation foundation model achieves one-shot customization, improving Dice on small and uncommon structures across six medical datasets.
-
Live(r) Die: Predicting Survival in Colorectal Liver Metastasis
A fully automated pre/post-contrast MRI framework, combining prompt-based segmentation with autoencoder multiple-instance survival analysis, improves CRLM post-surgery survival prediction over clinical and genomic bio...
-
OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations
OBBSeg segments irregular medical lesions from oriented bounding-box labels via a Mask-to-OBB loss and prompt modules, claiming near fully-supervised accuracy across 13 datasets and 5 modalities.
-
UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation
Adapting SAM3 to ultrasound with 171k image–mask–concept pairs gives a text-promptable multi-organ segmenter that beats general medical concept-segmentation baselines on average, though not on every organ.
-
Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2
Lean-SAM2 combines target-anchored memory pruning, condensed insurance memory, and risk-aware window routing to accelerate SAM2.1 inference ~1.4× with better accuracy than Efficient-SAM2.
-
MIS-HCC: Hierarchical Channel Clustering for Efficient Medical Image Segmentation
MIS-HCC prunes medical segmentation networks by Wasserstein-based hierarchical clustering of channels followed by parameter averaging, reporting near-baseline accuracy at 87.5% pruning.
-
Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer
A segmentation-guided transformer with slice fusion improves four-class prostate MRI classification to 0.633 accuracy and 0.768 joint recall on a 344-patient pathology-grounded cohort.
-
Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation
LoRA on SAM3’s prompt encoder, detector, and tracker (0.98% of parameters) raises surgical concept-segmentation mIoU over zero-shot SAM3 and Medical SAM3 while fitting in ~9 GB GPU memory.
-
Robust Activation Map Rectification for Weakly Supervised Volumetric Segmentation: Temporal Coherence as a Free Lunch
CSSeg rectifies noisy CAMs in volumetric scans by cross-slice averaging and one-sided replacement of suspicious frames, then prompts MedSAM, reporting large but unverified Dice/mIoU gains.
-
ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
ENSAM, a 5.5M-parameter promptable 3D segmentation model trained on a 4,471-volume coreset with one GPU for 6 hours, scored 2.404 DSC AUC on the CVPR 2025 hidden test, beating VISTA3D and SAM-Med3D and nearly matching SegVol.
-
Towards Affordable Tumor Segmentation and Visualization for 3D Breast MRI Using SAM2
SAM2 can segment breast tumors in 3D MRI with a single bounding-box prompt, and center-outward propagation yields the best volumetric Dice.
Discussion (0). Continue with ORCID to comment.