Pith. sign in

REVIEW 7 cited by

A Comprehensive Survey on Segment Anything Model for Vision and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.08196 v2 pith:S2PNUUW6 submitted 2023-05-14 cs.CV cs.AI

classification cs.CVcs.AI
keywords foundationmodelsanythingapplicationsmodeltasksvisionbeyond
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Artificial intelligence (AI) is evolving towards artificial general intelligence, which refers to the ability of an AI system to perform a wide range of tasks and exhibit a level of intelligence similar to that of a human being. This is in contrast to narrow or specialized AI, which is designed to perform specific tasks with a high degree of efficiency. Therefore, it is urgent to design a general class of models, which we term foundation models, trained on broad data that can be adapted to various downstream tasks. The recently proposed segment anything model (SAM) has made significant progress in breaking the boundaries of segmentation, greatly promoting the development of foundation models for computer vision. To fully comprehend SAM, we conduct a survey study. As the first to comprehensively review the progress of segmenting anything task for vision and beyond based on the foundation model of SAM, this work focuses on its applications to various tasks and data types by discussing its historical development, recent progress, and profound impact on broad applications. We first introduce the background and terminology for foundation models including SAM, as well as state-of-the-art methods contemporaneous with SAM that are significant for segmenting anything task. Then, we analyze and summarize the advantages and limitations of SAM across various image processing applications, including software scenes, real-world scenes, and complex scenes. Importantly, many insights are drawn to guide future research to develop more versatile foundation models and improve the architecture of SAM. We also summarize massive other amazing applications of SAM in vision and beyond. Finally, we maintain a continuously updated paper list and an open-source project summary for foundation model SAM at \href{https://github.com/liliu-avril/Awesome-Segment-Anything}{\color{magenta}{here}}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A SAM-based model with shape/intensity adapters, neighboring feature aggregation, and wavelet detail enhancement achieves 73.6 mAP on a new 47-species zooplankton microscopy dataset.

  2. SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    SegMoTE shows that adding token-level mixture-of-experts routing to a frozen SAM decoder can match or beat medical-segmentation models trained on far more data, using 0.15M curated masks and 17M trainable parameters.

  3. Grouped Speculative Decoding for Autoregressive Image Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.

  4. DOMR: Establishing Cross-View Segmentation via Dense Object Matching

    cs.CV 2025-08 conditional novelty 5.0 of 10

    DOMR jointly matches and refines multiple object masks across ego and exo views, reaching 49.7% and 55.2% mean IoU on Ego-Exo4D.

  5. MergeSAM: Unsupervised change detection of remote sensing images based on the Segment Anything Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A SAM-based unsupervised change detection method that matches and splits segmentation masks across two dates, improving F1 over AnyChange on GZ_CD_data.

  6. Fully Automated SAM for Single-source Domain Generalization in Medical Image Segmentation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    FA-SAM automates SAM-based medical segmentation across domains by generating prompt boxes with an uncertainty-enhanced network and fusing image and prompt embeddings.

  7. Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A structured survey of prompt engineering methods for the Segment Anything Model, covering geometric, textual, and multimodal prompts and their applications.

Pith tools