Pith. sign in

REVIEW 3 major objections 4 minor 95 references

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims to be the first comprehensive survey dedicated to prompt engineering in the Segment Anything Model and its variants, organizing 88 methods into a three-family taxonomy.

desk verdict Useful literature map on SAM prompt engineering, but the 'first comprehensive survey' claim does not survive contact with its own reference list. read the letter →

arxiv 2507.09562 v1 pith:JRPCRIMI submitted 2025-07-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords promptengineeringsegmentanythingmodelSAMimagesegmentationfoundationmodelsmultimodalpromptspointbox
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a survey that aims to be the first comprehensive map dedicated to prompt engineering in the Segment Anything Model (SAM) and its variants. It argues that the model's versatility comes from the prompts it receives, and that prompt engineering has evolved from manual clicks to automated, learned, and multimodal prompting. The survey organizes the literature into a three-family taxonomy: geometric prompts (points, boxes, masks), textual semantic prompts, and multimodal fusion prompts. It adds a second axis of automated generation strategies, such as detector-based, reinforcement-learning, and prototype-learning approaches, then applies the framework to medical imaging, remote sensing, and industrial anomaly detection. For a reader, the payoff would be a single structured reference that makes the field's methods and open problems easy to locate.

What carries the argument

The central object is the prompt itself: the point, box, mask, or piece of text that tells SAM what to segment. The paper's organizing device is a three-level taxonomy, which classifies methods into geometric prompts (points, boxes, masks), textual semantic prompts (class descriptions and part-level semantics), and multimodal fusion prompts (vision-language alignment and cross-modal attention). The taxonomy carries the survey's argument by giving every one of the 88 cited methods a home, while a second axis of automated generation strategies (detector-based, reinforcement-learning, and prototype-learning) tracks how prompts are produced.

What would settle it

The taxonomy's completeness would be falsified by finding a substantial set of SAM prompt-engineering papers that cannot be assigned to any of the three prompt families or four generation strategies. The paper's priority claim can be tested directly: because reference [75] is titled 'A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering,' comparing its coverage with this survey's would settle whether the 'first comprehensive survey focused on prompt engineering' claim holds.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that prompt engineering is the load-bearing mechanism behind SAM's flexible segmentation, and that the growing body of work on this mechanism can be captured in a compact taxonomy. According to the paper, every existing prompting approach for SAM falls into one of three families: geometric prompts, textual semantic prompts, and multimodal fusion prompts. A separate strand of work automates prompt generation through detectors, reinforcement-learning agents, or prototype learning, moving the field from manual annotation toward data-driven adaptation. The paper claims to be the first survey centered on this prompt dimension, thereby filling a gap it identifies in broader SAM surveys.

Load-bearing premise

The survey's claim to be a complete map rests on the assumption that the 88 papers it cites are a representative sample of the field and that the proposed taxonomy organizes them without force; because the paper describes no systematic search protocol for choosing those papers, this assumption cannot be checked from the text alone.

Editorial extensions

If this is right

  • A researcher can use the taxonomy as a checklist to classify any new SAM prompting method and compare it with existing work.
  • New automated prompting systems can be positioned against the four generation strategies the survey distinguishes: heuristic, detector-based, reinforcement-learning, and prototype-learning.
  • The survey's challenge list points to concrete research targets: reducing prompt sensitivity, resolving multi-prompt conflicts, handling occlusion and clutter, and cutting computational cost.
  • Its future-directions section suggests specific unexplored techniques, including causal prompting, multi-agent prompt collaboration, diffusion-based progressive refinement, and unsupervised prompt adaptation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the paper does not draw: its two categories of textual semantic prompts and multimodal fusion prompts are not sharply disjoint, since a method like SP-SAM still relies on CLIP text embeddings; the taxonomy is probably best read as idealized types rather than exclusive buckets.
  • One testable extension is to use the taxonomy to predict domain choices: the application tables suggest that medical imaging leans on geometric and prototype-driven prompts, while text-involved prompting clusters in zero-shot anomaly detection; a quantitative meta-analysis of the cited papers could verify this.
  • The survey's 'first' claim is checkable by history: the earlier reference [75] already has prompt engineering in its title, so the claimed novelty should be read as 'first to focus exclusively and comprehensively on prompt engineering,' a narrower statement than it first appears.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript is a survey of prompt engineering for the Segment Anything Model (SAM) and its variants. It proposes a hierarchical taxonomy of prompt types (geometric, textual, multimodal), surveys roughly 88 papers, organizes automated prompt-generation strategies (detector-based, reinforcement-learning, prototype-learning), reviews applications in medical imaging, remote sensing, and industrial anomaly detection, and closes with challenges and future directions. The paper's central claim is that it is the first comprehensive survey specifically focused on prompt engineering in SAM, and that this fills an important gap in the literature.

Significance. If the taxonomy and coverage are made auditable and the novelty claim is properly substantiated, the survey would be a useful reference for researchers working on SAM-based segmentation. The organization by geometric, textual, and multimodal prompts is reasonable, and the survey brings together a large and currently scattered body of work, including several very recent methods. The paper does not derive new empirical results, so its value rests on accuracy of representation and completeness of coverage; those aspects need to be checked and made transparent. The paper also honestly identifies open problems such as prompt sensitivity and multi-prompt conflicts, which are useful pointers for future work.

major comments (3)
  1. [Section 1, Introduction] The 'first comprehensive survey' claim is contradicted by the paper's own reference list: [75] is titled 'A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering.' The text dismisses [75] together with [21] and [76] by asserting that none offer 'a focused or systematic analysis of prompt engineering,' but it provides no comparison of [75]'s scope, organization, or coverage. Because the novelty claim is a stated contribution of the paper, this unsupported dismissal is load-bearing. Please either demonstrate precisely how the present survey differs from [75] or weaken the claim to an accurate statement of contribution.
  2. [Sections 3 and 4] No literature-search protocol is described: the paper does not state which databases were searched, the date range of coverage, the inclusion/exclusion criteria, or how the 88 papers were selected. Consequently, the qualifier 'comprehensive' cannot be audited, and the taxonomy in Sections 3 and 4 cannot be checked for omitted work. Add a methodology paragraph describing the search and selection process, or revise the claims to describe the survey's actual coverage rather than asserting completeness.
  3. [Section 3.2.1 and Section 3.3.1] Citation [15] is used for two different works: GPRN in Section 3.2.1 and GenSAM in Section 3.3.1, but the reference list entry [15] is 'Relax Image-Specific Prompt Requirement in SAM' (the GenSAM paper). GPRN appears to have no corresponding reference entry. This breaks traceability for two methods that are central to the survey's taxonomy and must be corrected.
minor comments (4)
  1. [Section 6.1] The claim that minor prompt variations lead to significant segmentation differences is introduced with 'Studies show' but no citation is supplied; please add supporting references or identify the specific studies.
  2. [Section 3.2.2 and other headings] There are typographical artifacts in headings, such as 'T extual Semantic Prompts' and 'F usion'; please proofread headings for missing spaces.
  3. [Table 1] The method name 'ESP-MEDSAM' in Table 1 differs from 'ESP-MedSAM' used in the text, and the header 'T asks' appears to be a typo for 'Tasks'; please standardize names and fix the header.
  4. [Section 3.1] The category 'Direct Image-based Embedding Generation' is presented as a prompt type, but it is really a strategy for generating prompts from image embeddings; clarify how this fits within the proposed taxonomy of prompt types.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this is a descriptive literature survey with no derivation chain, fitted parameters, or load-bearing self-citations.

full rationale

The paper makes no empirical prediction and contains no equations that could reduce to their own inputs. Its abstract and Section 1 assert that it is 'the first comprehensive survey focusing specifically on prompt engineering techniques for SAM and its variants,' but that is a novelty claim about the literature, not a derived result; even if reference [75] undermines the 'first' claim, the dispute is about scope and completeness, not circularity. Sections 3 and 4 organize 88 cited papers into a taxonomy of geometric, textual, and multimodal prompts, and the organizing categories are descriptive labels rather than fitted or self-defined quantities. No load-bearing step depends on a self-citation: the cited prior surveys are used only to delineate the claimed gap, and the underlying methods' reported results are external evidence. The absence of a stated literature-search protocol affects auditability of the 'comprehensive' claim but does not make the survey's content circular. The survey is therefore a self-contained exposition whose value can be assessed by its coverage and accuracy, and no circular step is exhibited.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

As a review, the paper introduces no free parameters or new entities. It relies on the accuracy of its literature summarization and on the suitability of its proposed taxonomy.

assumptions (2)
  • domain assumption The cited papers are accurately summarized and correctly attributed to the described methods.
    The survey's taxonomy and characterizations rely on the authors' reading of 88 references; no independent verification is provided.
  • ad hoc to paper The proposed taxonomy (geometric, textual, multimodal) is a valid and non-overlapping way to organize prompt engineering methods.
    The categories are introduced by the authors and are not derived from first principles; another reviewer might structure the field differently.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges." pith.science (2026). https://pith.science/paper/JRPCRIMI

@misc{pith2026250709562,
  author       = {Pith},
  title        = {Pith review of: Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JRPCRIMI}},
  note         = {Machine review of arXiv:2507.09562}
}
read the original abstract

The Segment Anything Model (SAM) has revolutionized image segmentation through its innovative prompt-based approach, yet the critical role of prompt engineering in its success remains underexplored. This paper presents the first comprehensive survey focusing specifically on prompt engineering techniques for SAM and its variants. We systematically organize and analyze the rapidly growing body of work in this emerging field, covering fundamental methodologies, practical applications, and key challenges. Our review reveals how prompt engineering has evolved from simple geometric inputs to sophisticated multimodal approaches, enabling SAM's adaptation across diverse domains including medical imaging and remote sensing. We identify unique challenges in prompt optimization and discuss promising research directions. This survey fills an important gap in the literature by providing a structured framework for understanding and advancing prompt engineering in foundation models for segmentation.

Figures

Figures reproduced from arXiv: 2507.09562 by the authors.

Figure 1
Figure 1. Section 2 details SAM’s architecture and prompt [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. Overview of the Survey By addressing this critical gap in the literature, our work provides a reference for researchers and practitioners and lays the groundwork for advancing prompt engineering in segmentation foundation models. 2 Preliminaries 2.1 Analysis of the SAM Architecture The Segment Anything Model (SAM) [23], a promptable segmentation model, is designed to transfer zero-shot to new image distributions and… view at source ↗
Figure 2
Figure 2. SAM framework [23] [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figures from the paper (7 more)
Figure 3
Figure 3. Figure 3: prompts in SAM [23] transforms these masks into rich feature representations that effectively delineate object boundaries and regions of inter￾est [23]. The flexibility of these prompt types enables SAM to adapt to various segmentation challenges, from simple ob￾ject i…
Figure 4
Figure 4. Figure 4: Direct Image-based Embedding Generation 3.2 Single-modality Prompt Strategies 3.2.1 Geometric Prompts Point Prompts In point prompts, inclusive points are used to guide SAM to focus on specific regions, while ex￾clusive points are used to steer the model away from back…
Figure 5
Figure 5. Figure 5: Textual Semantic Prompts embeddings. These part-level embeddings are then selec￾tively fused into whole-instrument representations through category- and image-specific attention weights. SP-SAM’s near-oracle performance on surgical datasets demonstrates that semantic p…
Figure 6
Figure 6. Figure 6: Multimodal Fusion Prompts [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Dynamic Interaction Approach key points) are automatically generated using pre-trained object detectors (e.g., YOLOv8 [22], Grounding DINO [34]), and these structured detection results are then used as in￾put prompts for SAM. This paradigm innovatively combines the eff…
Figure 9
Figure 9. Figure 9: Reinforcement Learning-driven Frameworks [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Prototype Learning Techniques NOv2, computing the distance matrix between blocks, and using bidirectional matching to improve accuracy. For more fine-grained prototype representation at the structural level, APL-SAM [52] introduces superpixel seg￾mentation and K-means…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

95 extracted references · 45 canonical work pages

  1. [75]

    SPPNet: A Single-Point Prompt Network for Nuclei Image Segmentation

    Qing Xu et al. SPPNet: A Single-Point Prompt Net- work for Nuclei Image Segmentation . 2023. arXiv: 2308.12231 [eess.IV]. url: https://arxiv.org/ abs/2308.12231

  2. [21]

    SAM2 for Image and Video Segmentation: A Comprehensive Survey

    Zhang Jiaxing and Tang Hao. SAM2 for Image and Video Segmentation: A Comprehensive Survey . 2025. arXiv: 2503.12781 [cs.CV]. url: https://arxiv. org/abs/2503.12781

  3. [76]

    ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

    Yufei Xu et al. ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation . 2022. arXiv: 2204 . 12484 [cs.CV]. url: https : / / arxiv . org / abs/2204.12484

  4. [15]

    Relax Image-Specific Prompt Require- ment in SAM: A Single Generic Prompt for Segment- ing Camouflaged Objects

    Jian Hu et al. Relax Image-Specific Prompt Require- ment in SAM: A Single Generic Prompt for Segment- ing Camouflaged Objects . 2023. arXiv: 2312 . 07374 [cs.CV]. url: https : / / arxiv . org / abs / 2312 . 07374

  5. [1]

    BioSAM: Generating SAM Prompts From Superpixel Graph for Biological In- stance Segmentation

    Miaomiao Cai et al. “BioSAM: Generating SAM Prompts From Superpixel Graph for Biological In- stance Segmentation”. In: IEEE Journal of Biomed- ical and Health Informatics 29.1 (2025), pp. 273–284. doi: 10.1109/JBHI.2024.3474706

  6. [2]

    Personalizing Vision-Language Models With Hybrid Prompts for Zero-Shot Anomaly Detection

    Yunkang Cao et al. “Personalizing Vision-Language Models With Hybrid Prompts for Zero-Shot Anomaly Detection”. In: IEEE Transactions on Cybernetics 55.4 (Apr. 2025), pp. 1917–1929. issn: 2168-2275. doi: 10.1109/tcyb.2025.3536165. url: http://dx.doi. org/10.1109/TCYB.2025.3536165

  7. [3]

    RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model

    Keyan Chen et al. RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model . 2023. arXiv: 2306 . 16269 [cs.CV]. url: https : / / arxiv . org / abs / 2306 . 16269

  8. [4]

    SAM-OCTA: Prompting Segment-Anything for OCTA Image Segmentation

    Xinrun Chen et al. SAM-OCTA: Prompting Segment- Anything for OCTA Image Segmentation. 2023. arXiv: 2310.07183 [cs.LG]

Show all 95 references
  1. [5]

    Segmentation by registration-enabled SAM prompt engineering using five reference images

    Yaxi Chen et al. Segmentation by registration-enabled SAM prompt engineering using five reference images

  2. [6]

    All-in-SAM: from Weak Annotation to Pixel-wise Nuclei Segmentation with Prompt-based Finetuning

    Can Cui et al. All-in-SAM: from Weak Annotation to Pixel-wise Nuclei Segmentation with Prompt-based Finetuning. 2023. arXiv: 2307.00290 [cs.CV]. url: https://arxiv.org/abs/2307.00290

  3. [7]

    SAMAug: Point Prompt Augmenta- tion for Segment Anything Model

    Haixing Dai et al. SAMAug: Point Prompt Augmenta- tion for Segment Anything Model . 2024. arXiv: 2307. 01187 [cs.CV]. url: https : / / arxiv . org / abs / 2307.01187

  4. [8]

    Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation

    Qiyuan Dai and Sibei Yang. Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation. 2024. arXiv: 2404 . 11998 [cs.CV]. url: https://arxiv.org/abs/2404.11998

  5. [9]

    SAM-U: Multi-box prompts trig- gered uncertainty estimation for reliable SAM in med- ical image

    Guoyao Deng et al. SAM-U: Multi-box prompts trig- gered uncertainty estimation for reliable SAM in med- ical image . 2023. arXiv: 2307 . 04973 [cs.CV]. url: https://arxiv.org/abs/2307.04973

  6. [10]

    K-SAM: A Prompting Method Using Pretrained U-Net to Im- prove Zero Shot Performance of SAM on Lung Seg- mentation in CXR Images

    Mohamed Deriche and Mohammad Marufur. K-SAM: A Prompting Method Using Pretrained U-Net to Im- prove Zero Shot Performance of SAM on Lung Seg- mentation in CXR Images . 2024. arXiv: 2410.06825 [eess.IV]. url: https : / / arxiv . org / abs / 2410 . 06825

  7. [11]

    Automating MedSAM by Learning Prompts with Weak Few-Shot Supervision

    M´ elanie Gaillochet, Christian Desrosiers, and Herv´ e Lombaert. Automating MedSAM by Learning Prompts with Weak Few-Shot Supervision . 2024. arXiv: 2409. 20293 [cs.CV]. url: https : / / arxiv . org / abs / 2409.20293

  8. [12]

    Swin-LiteMedSAM: A Lightweight Box-Based Seg- ment Anything Model for Large-Scale Medical Image Datasets

    Ruochen Gao, Donghang Lyu, and Marius Staring. Swin-LiteMedSAM: A Lightweight Box-Based Seg- ment Anything Model for Large-Scale Medical Image Datasets. 2024. arXiv: 2409 . 07172 [cs.CV]. url: https://arxiv.org/abs/2409.07172. 16

  9. [13]

    Lite Class-Prompt Tiny-VIT for Multi-modality Medical Image Segmentation

    Haotian Guan, Bingze Dai, and Jiajing Zhang. “Lite Class-Prompt Tiny-VIT for Multi-modality Medical Image Segmentation”. In: Medical Image Segmenta- tion Foundation Models. CVPR 2024 Challenge: Seg- ment Anything in Medical Images on Laptop - Med- SAM on Laptop 2024, Held in C...

  10. [14]

    APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation

    Weizhao He et al. APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation

  11. [16]

    08372 [cs.CV]

    arXiv: 2406 . 08372 [cs.CV]. url: https : / / arxiv.org/abs/2406.08372

  12. [17]

    Diffusion-empowered AutoPrompt MedSAM

    Peng Huang et al. Diffusion-empowered AutoPrompt MedSAM. 2025. arXiv: 2502.06817 [eess.IV]. url: https://arxiv.org/abs/2502.06817

  13. [18]

    Learning to Prompt Segment Anything Models

    Jiaxing Huang et al. Learning to Prompt Segment Anything Models. 2024. arXiv: 2401.04651 [cs.CV]. url: https://arxiv.org/abs/2401.04651

  14. [19]

    Robust Box Prompt based SAM for Medical Image Segmentation

    Yuhao Huang et al. Robust Box Prompt based SAM for Medical Image Segmentation . 2024. arXiv: 2407. 21284 [cs.CV]. url: https : / / arxiv . org / abs / 2407.21284

  15. [20]

    Optimizing Efficiency and Effec- tiveness in Sequential Prompt Strategy for SAM us- ing Reinforcement Learning

    Yifei Huang et al. “ Optimizing Efficiency and Effec- tiveness in Sequential Prompt Strategy for SAM us- ing Reinforcement Learning ”. In: proceedings of Med- ical Image Computing and Computer Assisted Inter- vention – MICCAI 2024 . Vol. LNCS 15008. Springer Nature Switzerland...

  16. [22]

    TV-SAM: Increasing Zero-Shot Seg- mentation Performance on Multimodal Medical Im- ages Using GPT-4 Generated Descriptive Prompts Without Human Annotation

    Zekun Jiang et al. TV-SAM: Increasing Zero-Shot Seg- mentation Performance on Multimodal Medical Im- ages Using GPT-4 Generated Descriptive Prompts Without Human Annotation. 2024. arXiv: 2402.15759 [cs.CV]. url: https : / / arxiv . org / abs / 2402 . 15759

  17. [23]

    Segment Anything

    Alexander Kirillov et al. Segment Anything . 2023. arXiv: 2304.02643 [cs.CV]. url: https://arxiv. org/abs/2304.02643

  18. [24]

    Ul- tralytics YOLOv8

    Glenn Jocher, Ayush Chaurasia, and Jing Qiu. Ul- tralytics YOLOv8 . Version 8.0.0. 2023. url: https : //github.com/ultralytics/ultralytics

  19. [25]

    Grounded Language-Image Pre-training

    Liunian Harold Li et al. Grounded Language-Image Pre-training. 2022. arXiv: 2112.03857 [cs.CV]. url: https://arxiv.org/abs/2112.03857

  20. [26]

    AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation

    Chengyin Li et al. “AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation”. In: 2025 IEEE/CVF Winter Conference on Applica- tions of Computer Vision (WACV) . 2025, pp. 3570–

  21. [27]

    A Closer Look at the Explainability of Con- trastive Language-Image Pre-training

    Yi Li et al. A Closer Look at the Explainability of Con- trastive Language-Image Pre-training . 2024. arXiv: 2304 . 05653 [cs.CV]. url: https : / / arxiv . org / abs/2304.05653

  22. [28]

    AM-SAM: Automated Prompting and Mask Calibration for Segment Anything Model

    Yuchen Li et al. AM-SAM: Automated Prompting and Mask Calibration for Segment Anything Model . 2024. arXiv: 2410.09714 [cs.CV]. url: https://arxiv. org/abs/2410.09714

  23. [29]

    ClipSAM: CLIP and SAM collabo- ration for zero-shot anomaly segmentation

    Shengze Li et al. “ClipSAM: CLIP and SAM collabo- ration for zero-shot anomaly segmentation”. In: Neu- rocomputing 618 (2025), p. 129122. issn: 0925-2312. doi: https://doi.org/10.1016/j.neucom.2024. 129122. url: https : / / www . sciencedirect . com / science/article/pii/S0925...

  24. [30]

    SAMRefiner: Taming Segment Any- thing Model for Universal Mask Refinement

    Yuqi Lin et al. SAMRefiner: Taming Segment Any- thing Model for Universal Mask Refinement . 2025. arXiv: 2502.06756 [cs.CV]. url: https://arxiv. org/abs/2502.06756

  25. [31]

    Training- Free Open-Ended Object Detection and Segmentation via Attention as Prompts

    Zhiwei Lin, Yongtao Wang, and Zhi Tang. Training- Free Open-Ended Object Detection and Segmentation via Attention as Prompts . 2024. arXiv: 2410 . 05963 [cs.CV]. url: https : / / arxiv . org / abs / 2410 . 05963

  26. [32]

    SAMCT: Segment Any CT Allow- ing Labor-Free Task-Indicator Prompts

    Xian Lin et al. SAMCT: Segment Any CT Allow- ing Labor-Free Task-Indicator Prompts . 2024. arXiv: 2403 . 13258 [cs.CV]. url: https : / / arxiv . org / abs/2403.13258

  27. [33]

    Rethinking Interactive Image Segmen- tation with Low Latency, High Quality, and Diverse Prompts

    Qin Liu et al. Rethinking Interactive Image Segmen- tation with Low Latency, High Quality, and Diverse Prompts. 2024. arXiv: 2404 . 00741 [cs.CV]. url: https://arxiv.org/abs/2404.00741

  28. [34]

    Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object De- tection

    Shilong Liu et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object De- tection. 2024. arXiv: 2303 . 05499 [cs.CV]. url: https://arxiv.org/abs/2303.05499

  29. [35]

    Unsupervised Continual Anomaly Detection with Contrastively-learned Prompt

    Jiaqi Liu et al. Unsupervised Continual Anomaly Detection with Contrastively-learned Prompt . 2024. arXiv: 2401.01010 [cs.CV]. url: https://arxiv. org/abs/2401.01010

  30. [36]

    Feature-prompting GBMSeg: One- Shot Reference Guided Training-Free Prompt Engi- neering for Glomerular Basement Membrane Seg- mentation

    Xueyu Liu et al. Feature-prompting GBMSeg: One- Shot Reference Guided Training-Free Prompt Engi- neering for Glomerular Basement Membrane Seg- mentation. 2024. arXiv: 2406 . 16271 [cs.CV]. url: https://arxiv.org/abs/2406.16271

  31. [37]

    SAM-RSIS: Progressively Adapt- ing SAM With Box Prompting to Remote Sensing Im- age Instance Segmentation

    Muying Luo et al. “SAM-RSIS: Progressively Adapt- ing SAM With Box Prompting to Remote Sensing Im- age Instance Segmentation”. In: IEEE Transactions on Geoscience and Remote Sensing 62 (2024), pp. 1–

  32. [38]

    Point-supervised Brain Tumor Seg- mentation with Box-prompted MedSAM

    Xiaofeng Liu et al. Point-supervised Brain Tumor Seg- mentation with Box-prompted MedSAM . 2024. arXiv: 2408 . 00706 [cs.CV]. url: https : / / arxiv . org / abs/2408.00706. 17

  33. [39]

    CLISC: Bridging clip and sam by enhanced cam for unsupervised brain tumor segmenta- tion

    Xiaochuan Ma et al. CLISC: Bridging clip and sam by enhanced cam for unsupervised brain tumor segmenta- tion. 2025. arXiv: 2501.16246 [cs.CV]. url: https: //arxiv.org/abs/2501.16246

  34. [40]

    Self-Prompting Polyp Seg- mentation in Colonoscopy using Hybrid Yolo-SAM 2 Model

    Mobina Mansoori et al. Self-Prompting Polyp Seg- mentation in Colonoscopy using Hybrid Yolo-SAM 2 Model. 2024. arXiv: 2409 . 09484 [eess.IV]. url: https://arxiv.org/abs/2409.09484

  35. [41]

    doi: 10.1109/TGRS.2024.3460085

  36. [42]

    GroupPrompter: A Prompting Method for Semantic Segmentation Based on SAM

    Yichuang Luo et al. “GroupPrompter: A Prompting Method for Semantic Segmentation Based on SAM”. In: IEEE Access 11 (2023), pp. 106054–106062. doi: 10.1109/ACCESS.2023.3319740

  37. [43]

    Hyper- correlation Squeeze for Few-Shot Segmentation

    Juhong Min, Dahyun Kang, and Minsu Cho. Hyper- correlation Squeeze for Few-Shot Segmentation . 2021. arXiv: 2104.01538 [cs.CV]. url: https://arxiv. org/abs/2104.01538

  38. [44]

    CycleSAM: One-Shot Surgical Scene Segmentation using Cycle-Consistent Feature Matching to Prompt SAM

    Aditya Murali et al. CycleSAM: One-Shot Surgical Scene Segmentation using Cycle-Consistent Feature Matching to Prompt SAM . 2024. arXiv: 2407.06795 [cs.CV]. url: https : / / arxiv . org / abs / 2407 . 06795

  39. [45]

    Label Anything: Multi- Class Few-Shot Semantic Segmentation with Visual Prompts

    Pasquale De Marinis et al. Label Anything: Multi- Class Few-Shot Semantic Segmentation with Visual Prompts. 2024. arXiv: 2407 . 02075 [cs.CV]. url: https://arxiv.org/abs/2407.02075

  40. [46]

    Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medi- cal Image Segmentation

    Juzheng Miao et al. Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medi- cal Image Segmentation . 2024. arXiv: 2407 . 05416 [cs.CV]. url: https : / / arxiv . org / abs / 2407 . 05416

  41. [47]

    Benchmarking Human and Au- tomated Prompting in the Segment Anything Model

    Jorge Quesada et al. Benchmarking Human and Au- tomated Prompting in the Segment Anything Model

  42. [48]

    Learning Transferable Visual Mod- els From Natural Language Supervision

    Alec Radford et al. Learning Transferable Visual Mod- els From Natural Language Supervision . 2021. arXiv: 2103 . 00020 [cs.CV]. url: https : / / arxiv . org / abs/2103.00020

  43. [49]

    Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation

    Saiyang Na et al. Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation. 2024. arXiv: 2401 . 13220 [eess.IV]. url: https://arxiv.org/abs/2401.13220

  44. [50]

    SAMIC: Segment Anything with In-Context Spatial Prompt Engineering

    Savinay Nagendra et al. SAMIC: Segment Anything with In-Context Spatial Prompt Engineering . 2024. arXiv: 2412.11998 [cs.CV]. url: https://arxiv. org/abs/2412.11998

  45. [51]

    Temporally-Extended Prompts Optimization for SAM in Interactive Medical Im- age Segmentation

    Chuyun Shen et al. “Temporally-Extended Prompts Optimization for SAM in Interactive Medical Im- age Segmentation”. In: 2023 IEEE International Con- ference on Bioinformatics and Biomedicine (BIBM) . 2023, pp. 3550–3557. doi: 10.1109/BIBM58861.2023. 10385291

  46. [52]

    22048 [cs.CV]

    arXiv: 2410 . 22048 [cs.CV]. url: https : / / arxiv.org/abs/2410.22048

  47. [53]

    EP-SAM: Weakly Supervised Histopathology Segmentation via Enhanced Prompt with Segment Anything

    Joonhyeon Song et al. EP-SAM: Weakly Supervised Histopathology Segmentation via Enhanced Prompt with Segment Anything . 2024. arXiv: 2410 . 13621 [cs.CV]. url: https : / / arxiv . org / abs / 2410 . 13621

  48. [54]

    PP-SAM: Perturbed Prompts for Robust Adaptation of Segment Anything Model for Polyp Segmentation

    Md Mostafijur Rahman et al. PP-SAM: Perturbed Prompts for Robust Adaptation of Segment Anything Model for Polyp Segmentation . 2024. arXiv: 2405 . 16740 [cs.CV]. url: https : / / arxiv . org / abs / 2405.16740

  49. [55]

    Vision and Language Reference Prompt into SAM for Few-shot Segmentation

    Kosuke Sakurai, Ryotaro Shimizu, and Masayuki Goto. Vision and Language Reference Prompt into SAM for Few-shot Segmentation . 2025. arXiv: 2502. 00719 [cs.CV]. url: https : / / arxiv . org / abs / 2502.00719

  50. [56]

    Sam2Rad: A Segmentation Model for Medical Images with Learnable Prompts

    Assefa Seyoum Wahd et al. Sam2Rad: A Segmentation Model for Medical Images with Learnable Prompts

  51. [57]

    Adaptive Prompt Learning with SAM for Few-shot Scanning Probe Microscope Image Seg- mentation

    Yao Shen et al. Adaptive Prompt Learning with SAM for Few-shot Scanning Probe Microscope Image Seg- mentation. 2024. arXiv: 2410 . 12562 [cs.CV]. url: https://arxiv.org/abs/2410.12562

  52. [58]

    CogVLM: Visual Expert for Pre- trained Language Models

    Weihan Wang et al. CogVLM: Visual Expert for Pre- trained Language Models . 2024. arXiv: 2311 . 03079 [cs.CV]. url: https : / / arxiv . org / abs / 2311 . 03079

  53. [59]

    Deep High-Resolution Representation Learning for Human Pose Estimation

    Ke Sun et al. Deep High-Resolution Representation Learning for Human Pose Estimation . 2019. arXiv: 1902 . 09212 [cs.CV]. url: https : / / arxiv . org / abs/1902.09212

  54. [60]

    On Efficient Variants of Segment Anything Model: A Survey

    Xiaorui Sun et al. On Efficient Variants of Segment Anything Model: A Survey . 2024. arXiv: 2410.04960 [cs.CV]. url: https : / / arxiv . org / abs / 2410 . 04960

  55. [61]

    TinyViT: Fast Pretraining Distillation for Small Vision Transformers

    Kan Wu et al. TinyViT: Fast Pretraining Distillation for Small Vision Transformers . 2022. arXiv: 2207 . 10666 [cs.CV]. url: https : / / arxiv . org / abs / 2207.10666

  56. [62]

    06821 [cs.CV]

    arXiv: 2409 . 06821 [cs.CV]. url: https : / / arxiv.org/abs/2409.06821

  57. [63]

    Auto-Prompting SAM for Weakly Supervised Landslide Extraction

    Jian Wang et al. Auto-Prompting SAM for Weakly Supervised Landslide Extraction . 2025. arXiv: 2501 . 13426 [cs.CV]. url: https : / / arxiv . org / abs / 2501.13426

  58. [64]

    Self-Prompt SAM: Medical Image Seg- mentation via Automatic Prompt SAM Adaptation

    Bin Xie et al. Self-Prompt SAM: Medical Image Seg- mentation via Automatic Prompt SAM Adaptation

  59. [65]

    CrackESS: A Self-Prompting Crack Segmentation System for Edge Devices

    Yingchu Wang, Ji He, and Shijie Yu. CrackESS: A Self-Prompting Crack Segmentation System for Edge Devices. 2025. arXiv: 2412 . 07205 [cs.CV]. url: https://arxiv.org/abs/2412.07205. 18

  60. [66]

    Optimizing Prompt Strategies for SAM: Advancing lesion Segmentation Across Diverse Medical Imaging Modalities

    Yuli Wang et al. Optimizing Prompt Strategies for SAM: Advancing lesion Segmentation Across Diverse Medical Imaging Modalities. 2024. arXiv: 2412.17943 [eess.IV]

  61. [67]

    ESP-MedSAM: Efficient Self- Prompting SAM for Universal Domain-Generalized Medical Image Segmentation

    Qing Xu et al. ESP-MedSAM: Efficient Self- Prompting SAM for Universal Domain-Generalized Medical Image Segmentation . 2024. arXiv: 2407 . 14153 [eess.IV]. url: https : / / arxiv . org / abs / 2407.14153

  62. [68]

    Self- Prompting Large Vision Models for Few-Shot Med- ical Image Segmentation

    Qi Wu, Yuyao Zhang, and Marawan Elbatel. Self- Prompting Large Vision Models for Few-Shot Med- ical Image Segmentation . 2023. arXiv: 2308 . 07624 [cs.CV]. url: https : / / arxiv . org / abs / 2308 . 07624

  63. [69]

    Integrating multi-scale informa- tion and diverse prompts in large model SAM-Med2D for accurate left ventricular ejection fraction estima- tion

    Yagang Wu et al. “Integrating multi-scale informa- tion and diverse prompts in large model SAM-Med2D for accurate left ventricular ejection fraction estima- tion”. In: Medical & Biological Engineering & Com- puting (Feb. 14, 2025).issn: 1741-0444. doi: 10.1007/ s11517- 025- 03...

  64. [70]

    PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Im- age Segmentation

    Zhonghao Yan et al. PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Im- age Segmentation. 2025. arXiv: 2501.06692 [cs.CV]. url: https://arxiv.org/abs/2501.06692

  65. [71]

    TA VP: Task-Adaptive Visual Prompt for Cross-domain Few-shot Segmentation

    Jiaqi Yang et al. TA VP: Task-Adaptive Visual Prompt for Cross-domain Few-shot Segmentation. 2024. arXiv: 2409 . 05393 [cs.CV]. url: https : / / arxiv . org / abs/2409.05393

  66. [72]

    Char-SAM: Turning Segment Any- thing Model into Scene Text Segmentation Annota- tor with Character-level Visual Prompts

    Enze Xie et al. Char-SAM: Turning Segment Any- thing Model into Scene Text Segmentation Annota- tor with Character-level Visual Prompts . 2024. arXiv: 2412 . 19917 [cs.CV]. url: https : / / arxiv . org / abs/2412.19917

  67. [73]

    SAM-MPA: Applying SAM to Few- shot Medical Image Segmentation using Mask Propa- gation and Auto-prompting

    Jie Xu et al. SAM-MPA: Applying SAM to Few- shot Medical Image Segmentation using Mask Propa- gation and Auto-prompting . 2024. arXiv: 2411.17363 [cs.CV]. url: https : / / arxiv . org / abs / 2411 . 17363

  68. [74]

    SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation

    Wenxi Yue et al. SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation . 2023. arXiv: 2308.08746 [cs.CV]. url: https://arxiv. org/abs/2308.08746

  69. [77]

    COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection

    Xiaoqin Zhang et al. COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection. 2024. arXiv: 2411.18858 [cs.CV]. url: https : / / arxiv . org / abs / 2411 . 18858

  70. [78]

    UV-SAM: Adapting Segment Any- thing Model for Urban Village Identification

    Xin Zhang et al. UV-SAM: Adapting Segment Any- thing Model for Urban Village Identification . 2024. arXiv: 2401.08083 [cs.CV]. url: https://arxiv. org/abs/2401.08083

  71. [79]

    Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization

    Xi Yang et al. “Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization”. In: Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LXIX. Milan, Italy: Springer- Verlag...

  72. [80]

    SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Seg- mentation

    Wenxi Yue et al. SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Seg- mentation. 2024. arXiv: 2312 . 14481 [cs.CV]. url: https://arxiv.org/abs/2312.14481

  73. [81]

    Automatic Seg- mentation Annotation of Space Target Using Segment Anything Model and Object Detection Prompts

    Zhihao Zhang and Zhaohui Dang. “Automatic Seg- mentation Annotation of Space Target Using Segment Anything Model and Object Detection Prompts”. In: IEEE Transactions on Aerospace and Electronic Sys- tems (2024), pp. 1–15. doi: 10 . 1109 / TAES . 2024 . 3512533. 19

  74. [82]

    A Survey on Segment Any- thing Model (SAM): Vision Foundation Model Meets Prompt Engineering

    Chaoning Zhang et al. A Survey on Segment Any- thing Model (SAM): Vision Foundation Model Meets Prompt Engineering . 2024. arXiv: 2306 . 06211 [cs.CV]. url: https : / / arxiv . org / abs / 2306 . 06211

  75. [83]

    A Comprehensive Survey on Seg- ment Anything Model for Vision and Beyond

    Chunhui Zhang et al. A Comprehensive Survey on Seg- ment Anything Model for Vision and Beyond . 2023. arXiv: 2305.08196 [cs.CV]. url: https://arxiv. org/abs/2305.08196

  76. [84]

    Curriculum Prompting Foundation Models for Medical Image Segmentation

    Xiuqi Zheng et al. Curriculum Prompting Foundation Models for Medical Image Segmentation . 2024. arXiv: 2409 . 00695 [cs.CV]. url: https : / / arxiv . org / abs/2409.00695

  77. [85]

    EdgeSAM: Prompt-In-the-Loop Distillation for On-Device Deployment of SAM

    Chong Zhou et al. EdgeSAM: Prompt-In-the-Loop Distillation for On-Device Deployment of SAM . 2024. arXiv: 2312.06660 [cs.CV]. url: https://arxiv. org/abs/2312.06660

  78. [86]

    Towards Segment Any- thing Model (SAM) for Medical Image Segmentation: A Survey

    Yichi Zhang and Rushi Jiao. Towards Segment Any- thing Model (SAM) for Medical Image Segmentation: A Survey . 2023. arXiv: 2305.03678 [eess.IV]. url: https://arxiv.org/abs/2305.03678

  79. [87]

    Enhancing the Reliability of Seg- ment Anything Model for Auto-Prompting Medical Im- age Segmentation with Uncertainty Rectification

    Yichi Zhang et al. Enhancing the Reliability of Seg- ment Anything Model for Auto-Prompting Medical Im- age Segmentation with Uncertainty Rectification. 2024. arXiv: 2311.10529 [cs.CV]. url: https://arxiv. org/abs/2311.10529

  80. [88]

    Segment Everything Everywhere All at Once

    Xueyan Zou et al. Segment Everything Everywhere All at Once . 2023. arXiv: 2304.06718 [cs.CV]. url: https://arxiv.org/abs/2304.06718. 20

  81. [89]

    Semantic-Enhanced Point-Box Joint Prompting for Video Object Segmentation

    Quan Zhao et al. “Semantic-Enhanced Point-Box Joint Prompting for Video Object Segmentation”. In: 2024 IEEE International Conference on Image Pro- cessing (ICIP) . 2024, pp. 2334–2340. doi: 10.1109/ ICIP51287.2024.10648107

  82. [90]

    Fast Segment Anything

    Xu Zhao et al. Fast Segment Anything . 2023. arXiv: 2306 . 12156 [cs.CV]. url: https : / / arxiv . org / abs/2306.12156

  83. [93]

    MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAM

    Nan Zhou et al. MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAM. 2024. arXiv: 2409.00924 [cs.CV]. url: https://arxiv. org/abs/2409.00924

  84. [94]

    ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual De- scriptions

    Deyao Zhu et al. ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual De- scriptions. 2023. arXiv: 2303 . 06594 [cs.CV]. url: https://arxiv.org/abs/2303.06594

  85. [2024]

    17933 [cs.CV]

    arXiv: 2407 . 17933 [cs.CV]. url: https : / / arxiv.org/abs/2407.17933

  86. [2025]

    00630 [cs.CV]

    arXiv: 2502 . 00630 [cs.CV]. url: https : / / arxiv.org/abs/2502.00630

  87. [3580]

    doi: 10.1109/WACV61041.2025.00352

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.