REVIEW 3 major objections 4 minor 95 references
Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims to be the first comprehensive survey dedicated to prompt engineering in the Segment Anything Model and its variants, organizing 88 methods into a three-family taxonomy.
desk verdict Useful literature map on SAM prompt engineering, but the 'first comprehensive survey' claim does not survive contact with its own reference list. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the prompt itself: the point, box, mask, or piece of text that tells SAM what to segment. The paper's organizing device is a three-level taxonomy, which classifies methods into geometric prompts (points, boxes, masks), textual semantic prompts (class descriptions and part-level semantics), and multimodal fusion prompts (vision-language alignment and cross-modal attention). The taxonomy carries the survey's argument by giving every one of the 88 cited methods a home, while a second axis of automated generation strategies (detector-based, reinforcement-learning, and prototype-learning) tracks how prompts are produced.
What would settle it
The taxonomy's completeness would be falsified by finding a substantial set of SAM prompt-engineering papers that cannot be assigned to any of the three prompt families or four generation strategies. The paper's priority claim can be tested directly: because reference [75] is titled 'A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering,' comparing its coverage with this survey's would settle whether the 'first comprehensive survey focused on prompt engineering' claim holds.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that prompt engineering is the load-bearing mechanism behind SAM's flexible segmentation, and that the growing body of work on this mechanism can be captured in a compact taxonomy. According to the paper, every existing prompting approach for SAM falls into one of three families: geometric prompts, textual semantic prompts, and multimodal fusion prompts. A separate strand of work automates prompt generation through detectors, reinforcement-learning agents, or prototype learning, moving the field from manual annotation toward data-driven adaptation. The paper claims to be the first survey centered on this prompt dimension, thereby filling a gap it identifies in broader SAM surveys.
Load-bearing premise
The survey's claim to be a complete map rests on the assumption that the 88 papers it cites are a representative sample of the field and that the proposed taxonomy organizes them without force; because the paper describes no systematic search protocol for choosing those papers, this assumption cannot be checked from the text alone.
Editorial extensions
If this is right
- A researcher can use the taxonomy as a checklist to classify any new SAM prompting method and compare it with existing work.
- New automated prompting systems can be positioned against the four generation strategies the survey distinguishes: heuristic, detector-based, reinforcement-learning, and prototype-learning.
- The survey's challenge list points to concrete research targets: reducing prompt sensitivity, resolving multi-prompt conflicts, handling occlusion and clutter, and cutting computational cost.
- Its future-directions section suggests specific unexplored techniques, including causal prompting, multi-agent prompt collaboration, diffusion-based progressive refinement, and unsupervised prompt adaptation.
Reading between the lines
- A consequence the paper does not draw: its two categories of textual semantic prompts and multimodal fusion prompts are not sharply disjoint, since a method like SP-SAM still relies on CLIP text embeddings; the taxonomy is probably best read as idealized types rather than exclusive buckets.
- One testable extension is to use the taxonomy to predict domain choices: the application tables suggest that medical imaging leans on geometric and prototype-driven prompts, while text-involved prompting clusters in zero-shot anomaly detection; a quantitative meta-analysis of the cited papers could verify this.
- The survey's 'first' claim is checkable by history: the earlier reference [75] already has prompt engineering in its title, so the claimed novelty should be read as 'first to focus exclusively and comprehensively on prompt engineering,' a narrower statement than it first appears.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of prompt engineering for the Segment Anything Model (SAM) and its variants. It proposes a hierarchical taxonomy of prompt types (geometric, textual, multimodal), surveys roughly 88 papers, organizes automated prompt-generation strategies (detector-based, reinforcement-learning, prototype-learning), reviews applications in medical imaging, remote sensing, and industrial anomaly detection, and closes with challenges and future directions. The paper's central claim is that it is the first comprehensive survey specifically focused on prompt engineering in SAM, and that this fills an important gap in the literature.
Significance. If the taxonomy and coverage are made auditable and the novelty claim is properly substantiated, the survey would be a useful reference for researchers working on SAM-based segmentation. The organization by geometric, textual, and multimodal prompts is reasonable, and the survey brings together a large and currently scattered body of work, including several very recent methods. The paper does not derive new empirical results, so its value rests on accuracy of representation and completeness of coverage; those aspects need to be checked and made transparent. The paper also honestly identifies open problems such as prompt sensitivity and multi-prompt conflicts, which are useful pointers for future work.
major comments (3)
- [Section 1, Introduction] The 'first comprehensive survey' claim is contradicted by the paper's own reference list: [75] is titled 'A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering.' The text dismisses [75] together with [21] and [76] by asserting that none offer 'a focused or systematic analysis of prompt engineering,' but it provides no comparison of [75]'s scope, organization, or coverage. Because the novelty claim is a stated contribution of the paper, this unsupported dismissal is load-bearing. Please either demonstrate precisely how the present survey differs from [75] or weaken the claim to an accurate statement of contribution.
- [Sections 3 and 4] No literature-search protocol is described: the paper does not state which databases were searched, the date range of coverage, the inclusion/exclusion criteria, or how the 88 papers were selected. Consequently, the qualifier 'comprehensive' cannot be audited, and the taxonomy in Sections 3 and 4 cannot be checked for omitted work. Add a methodology paragraph describing the search and selection process, or revise the claims to describe the survey's actual coverage rather than asserting completeness.
- [Section 3.2.1 and Section 3.3.1] Citation [15] is used for two different works: GPRN in Section 3.2.1 and GenSAM in Section 3.3.1, but the reference list entry [15] is 'Relax Image-Specific Prompt Requirement in SAM' (the GenSAM paper). GPRN appears to have no corresponding reference entry. This breaks traceability for two methods that are central to the survey's taxonomy and must be corrected.
minor comments (4)
- [Section 6.1] The claim that minor prompt variations lead to significant segmentation differences is introduced with 'Studies show' but no citation is supplied; please add supporting references or identify the specific studies.
- [Section 3.2.2 and other headings] There are typographical artifacts in headings, such as 'T extual Semantic Prompts' and 'F usion'; please proofread headings for missing spaces.
- [Table 1] The method name 'ESP-MEDSAM' in Table 1 differs from 'ESP-MedSAM' used in the text, and the header 'T asks' appears to be a typo for 'Tasks'; please standardize names and fix the header.
- [Section 3.1] The category 'Direct Image-based Embedding Generation' is presented as a prompt type, but it is really a strategy for generating prompts from image embeddings; clarify how this fits within the proposed taxonomy of prompt types.
Circularity Check
No circularity: this is a descriptive literature survey with no derivation chain, fitted parameters, or load-bearing self-citations.
full rationale
The paper makes no empirical prediction and contains no equations that could reduce to their own inputs. Its abstract and Section 1 assert that it is 'the first comprehensive survey focusing specifically on prompt engineering techniques for SAM and its variants,' but that is a novelty claim about the literature, not a derived result; even if reference [75] undermines the 'first' claim, the dispute is about scope and completeness, not circularity. Sections 3 and 4 organize 88 cited papers into a taxonomy of geometric, textual, and multimodal prompts, and the organizing categories are descriptive labels rather than fitted or self-defined quantities. No load-bearing step depends on a self-citation: the cited prior surveys are used only to delineate the claimed gap, and the underlying methods' reported results are external evidence. The absence of a stated literature-search protocol affects auditability of the 'comprehensive' claim but does not make the survey's content circular. The survey is therefore a self-contained exposition whose value can be assessed by its coverage and accuracy, and no circular step is exhibited.
Assumptions & free parameters
assumptions (2)
- domain assumption The cited papers are accurately summarized and correctly attributed to the described methods.
- ad hoc to paper The proposed taxonomy (geometric, textual, multimodal) is a valid and non-overlapping way to organize prompt engineering methods.
Cite this review
Pith. "Pith review of Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges." pith.science (2026). https://pith.science/paper/JRPCRIMI
@misc{pith2026250709562,
author = {Pith},
title = {Pith review of: Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/JRPCRIMI}},
note = {Machine review of arXiv:2507.09562}
}
read the original abstract
The Segment Anything Model (SAM) has revolutionized image segmentation through its innovative prompt-based approach, yet the critical role of prompt engineering in its success remains underexplored. This paper presents the first comprehensive survey focusing specifically on prompt engineering techniques for SAM and its variants. We systematically organize and analyze the rapidly growing body of work in this emerging field, covering fundamental methodologies, practical applications, and key challenges. Our review reveals how prompt engineering has evolved from simple geometric inputs to sophisticated multimodal approaches, enabling SAM's adaptation across diverse domains including medical imaging and remote sensing. We identify unique challenges in prompt optimization and discuss promising research directions. This survey fills an important gap in the literature by providing a structured framework for understanding and advancing prompt engineering in foundation models for segmentation.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[75]
SPPNet: A Single-Point Prompt Network for Nuclei Image Segmentation
Qing Xu et al. SPPNet: A Single-Point Prompt Net- work for Nuclei Image Segmentation . 2023. arXiv: 2308.12231 [eess.IV]. url: https://arxiv.org/ abs/2308.12231
work page Pith review arXiv 2023
-
[21]
SAM2 for Image and Video Segmentation: A Comprehensive Survey
Zhang Jiaxing and Tang Hao. SAM2 for Image and Video Segmentation: A Comprehensive Survey . 2025. arXiv: 2503.12781 [cs.CV]. url: https://arxiv. org/abs/2503.12781
arXiv 2025
-
[76]
ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation
Yufei Xu et al. ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation . 2022. arXiv: 2204 . 12484 [cs.CV]. url: https : / / arxiv . org / abs/2204.12484
arXiv 2022
-
[15]
Relax Image-Specific Prompt Require- ment in SAM: A Single Generic Prompt for Segment- ing Camouflaged Objects
Jian Hu et al. Relax Image-Specific Prompt Require- ment in SAM: A Single Generic Prompt for Segment- ing Camouflaged Objects . 2023. arXiv: 2312 . 07374 [cs.CV]. url: https : / / arxiv . org / abs / 2312 . 07374
2023
-
[1]
BioSAM: Generating SAM Prompts From Superpixel Graph for Biological In- stance Segmentation
Miaomiao Cai et al. “BioSAM: Generating SAM Prompts From Superpixel Graph for Biological In- stance Segmentation”. In: IEEE Journal of Biomed- ical and Health Informatics 29.1 (2025), pp. 273–284. doi: 10.1109/JBHI.2024.3474706
arXiv 2025
-
[2]
Personalizing Vision-Language Models With Hybrid Prompts for Zero-Shot Anomaly Detection
Yunkang Cao et al. “Personalizing Vision-Language Models With Hybrid Prompts for Zero-Shot Anomaly Detection”. In: IEEE Transactions on Cybernetics 55.4 (Apr. 2025), pp. 1917–1929. issn: 2168-2275. doi: 10.1109/tcyb.2025.3536165. url: http://dx.doi. org/10.1109/TCYB.2025.3536165
-
[3]
RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model
Keyan Chen et al. RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model . 2023. arXiv: 2306 . 16269 [cs.CV]. url: https : / / arxiv . org / abs / 2306 . 16269
2023
-
[4]
SAM-OCTA: Prompting Segment-Anything for OCTA Image Segmentation
Xinrun Chen et al. SAM-OCTA: Prompting Segment- Anything for OCTA Image Segmentation. 2023. arXiv: 2310.07183 [cs.LG]
work page Pith review arXiv 2023
Show all 95 references
-
[5]
Segmentation by registration-enabled SAM prompt engineering using five reference images
Yaxi Chen et al. Segmentation by registration-enabled SAM prompt engineering using five reference images
-
[6]
All-in-SAM: from Weak Annotation to Pixel-wise Nuclei Segmentation with Prompt-based Finetuning
Can Cui et al. All-in-SAM: from Weak Annotation to Pixel-wise Nuclei Segmentation with Prompt-based Finetuning. 2023. arXiv: 2307.00290 [cs.CV]. url: https://arxiv.org/abs/2307.00290
2023 arXiv
-
[7]
SAMAug: Point Prompt Augmenta- tion for Segment Anything Model
Haixing Dai et al. SAMAug: Point Prompt Augmenta- tion for Segment Anything Model . 2024. arXiv: 2307. 01187 [cs.CV]. url: https : / / arxiv . org / abs / 2307.01187
2024 arXiv
-
[8]
Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation
Qiyuan Dai and Sibei Yang. Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation. 2024. arXiv: 2404 . 11998 [cs.CV]. url: https://arxiv.org/abs/2404.11998
2024 arXiv
-
[9]
SAM-U: Multi-box prompts trig- gered uncertainty estimation for reliable SAM in med- ical image
Guoyao Deng et al. SAM-U: Multi-box prompts trig- gered uncertainty estimation for reliable SAM in med- ical image . 2023. arXiv: 2307 . 04973 [cs.CV]. url: https://arxiv.org/abs/2307.04973
2023 arXiv
-
[10]
K-SAM: A Prompting Method Using Pretrained U-Net to Im- prove Zero Shot Performance of SAM on Lung Seg- mentation in CXR Images
Mohamed Deriche and Mohammad Marufur. K-SAM: A Prompting Method Using Pretrained U-Net to Im- prove Zero Shot Performance of SAM on Lung Seg- mentation in CXR Images . 2024. arXiv: 2410.06825 [eess.IV]. url: https : / / arxiv . org / abs / 2410 . 06825
2024 arXiv
-
[11]
Automating MedSAM by Learning Prompts with Weak Few-Shot Supervision
M´ elanie Gaillochet, Christian Desrosiers, and Herv´ e Lombaert. Automating MedSAM by Learning Prompts with Weak Few-Shot Supervision . 2024. arXiv: 2409. 20293 [cs.CV]. url: https : / / arxiv . org / abs / 2409.20293
2024 arXiv
-
[12]
Swin-LiteMedSAM: A Lightweight Box-Based Seg- ment Anything Model for Large-Scale Medical Image Datasets
Ruochen Gao, Donghang Lyu, and Marius Staring. Swin-LiteMedSAM: A Lightweight Box-Based Seg- ment Anything Model for Large-Scale Medical Image Datasets. 2024. arXiv: 2409 . 07172 [cs.CV]. url: https://arxiv.org/abs/2409.07172. 16
2024 arXiv
-
[13]
Lite Class-Prompt Tiny-VIT for Multi-modality Medical Image Segmentation
Haotian Guan, Bingze Dai, and Jiajing Zhang. “Lite Class-Prompt Tiny-VIT for Multi-modality Medical Image Segmentation”. In: Medical Image Segmenta- tion Foundation Models. CVPR 2024 Challenge: Seg- ment Anything in Medical Images on Laptop - Med- SAM on Laptop 2024, Held in C...
2024 doi
-
[14]
APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation
Weizhao He et al. APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic Segmentation
- [16]
-
[17]
Diffusion-empowered AutoPrompt MedSAM
Peng Huang et al. Diffusion-empowered AutoPrompt MedSAM. 2025. arXiv: 2502.06817 [eess.IV]. url: https://arxiv.org/abs/2502.06817
2025 arXiv
-
[18]
Learning to Prompt Segment Anything Models
Jiaxing Huang et al. Learning to Prompt Segment Anything Models. 2024. arXiv: 2401.04651 [cs.CV]. url: https://arxiv.org/abs/2401.04651
2024 arXiv
-
[19]
Robust Box Prompt based SAM for Medical Image Segmentation
Yuhao Huang et al. Robust Box Prompt based SAM for Medical Image Segmentation . 2024. arXiv: 2407. 21284 [cs.CV]. url: https : / / arxiv . org / abs / 2407.21284
2024 arXiv
-
[20]
Optimizing Efficiency and Effec- tiveness in Sequential Prompt Strategy for SAM us- ing Reinforcement Learning
Yifei Huang et al. “ Optimizing Efficiency and Effec- tiveness in Sequential Prompt Strategy for SAM us- ing Reinforcement Learning ”. In: proceedings of Med- ical Image Computing and Computer Assisted Inter- vention – MICCAI 2024 . Vol. LNCS 15008. Springer Nature Switzerland...
2024
-
[22]
TV-SAM: Increasing Zero-Shot Seg- mentation Performance on Multimodal Medical Im- ages Using GPT-4 Generated Descriptive Prompts Without Human Annotation
Zekun Jiang et al. TV-SAM: Increasing Zero-Shot Seg- mentation Performance on Multimodal Medical Im- ages Using GPT-4 Generated Descriptive Prompts Without Human Annotation. 2024. arXiv: 2402.15759 [cs.CV]. url: https : / / arxiv . org / abs / 2402 . 15759
2024 arXiv
-
[23]
Segment Anything
Alexander Kirillov et al. Segment Anything . 2023. arXiv: 2304.02643 [cs.CV]. url: https://arxiv. org/abs/2304.02643
2023 arXiv
-
[24]
Ul- tralytics YOLOv8
Glenn Jocher, Ayush Chaurasia, and Jing Qiu. Ul- tralytics YOLOv8 . Version 8.0.0. 2023. url: https : //github.com/ultralytics/ultralytics
2023
-
[25]
Grounded Language-Image Pre-training
Liunian Harold Li et al. Grounded Language-Image Pre-training. 2022. arXiv: 2112.03857 [cs.CV]. url: https://arxiv.org/abs/2112.03857
2022 arXiv
-
[26]
AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation
Chengyin Li et al. “AutoProSAM: Automated Prompting SAM for 3D Multi-Organ Segmentation”. In: 2025 IEEE/CVF Winter Conference on Applica- tions of Computer Vision (WACV) . 2025, pp. 3570–
2025
-
[27]
A Closer Look at the Explainability of Con- trastive Language-Image Pre-training
Yi Li et al. A Closer Look at the Explainability of Con- trastive Language-Image Pre-training . 2024. arXiv: 2304 . 05653 [cs.CV]. url: https : / / arxiv . org / abs/2304.05653
2024 arXiv
-
[28]
AM-SAM: Automated Prompting and Mask Calibration for Segment Anything Model
Yuchen Li et al. AM-SAM: Automated Prompting and Mask Calibration for Segment Anything Model . 2024. arXiv: 2410.09714 [cs.CV]. url: https://arxiv. org/abs/2410.09714
2024
-
[29]
ClipSAM: CLIP and SAM collabo- ration for zero-shot anomaly segmentation
Shengze Li et al. “ClipSAM: CLIP and SAM collabo- ration for zero-shot anomaly segmentation”. In: Neu- rocomputing 618 (2025), p. 129122. issn: 0925-2312. doi: https://doi.org/10.1016/j.neucom.2024. 129122. url: https : / / www . sciencedirect . com / science/article/pii/S0925...
2025 doi
-
[30]
SAMRefiner: Taming Segment Any- thing Model for Universal Mask Refinement
Yuqi Lin et al. SAMRefiner: Taming Segment Any- thing Model for Universal Mask Refinement . 2025. arXiv: 2502.06756 [cs.CV]. url: https://arxiv. org/abs/2502.06756
2025 arXiv
-
[31]
Training- Free Open-Ended Object Detection and Segmentation via Attention as Prompts
Zhiwei Lin, Yongtao Wang, and Zhi Tang. Training- Free Open-Ended Object Detection and Segmentation via Attention as Prompts . 2024. arXiv: 2410 . 05963 [cs.CV]. url: https : / / arxiv . org / abs / 2410 . 05963
2024
-
[32]
SAMCT: Segment Any CT Allow- ing Labor-Free Task-Indicator Prompts
Xian Lin et al. SAMCT: Segment Any CT Allow- ing Labor-Free Task-Indicator Prompts . 2024. arXiv: 2403 . 13258 [cs.CV]. url: https : / / arxiv . org / abs/2403.13258
2024 arXiv
-
[33]
Rethinking Interactive Image Segmen- tation with Low Latency, High Quality, and Diverse Prompts
Qin Liu et al. Rethinking Interactive Image Segmen- tation with Low Latency, High Quality, and Diverse Prompts. 2024. arXiv: 2404 . 00741 [cs.CV]. url: https://arxiv.org/abs/2404.00741
2024 arXiv
-
[34]
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object De- tection
Shilong Liu et al. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object De- tection. 2024. arXiv: 2303 . 05499 [cs.CV]. url: https://arxiv.org/abs/2303.05499
2024 arXiv
-
[35]
Unsupervised Continual Anomaly Detection with Contrastively-learned Prompt
Jiaqi Liu et al. Unsupervised Continual Anomaly Detection with Contrastively-learned Prompt . 2024. arXiv: 2401.01010 [cs.CV]. url: https://arxiv. org/abs/2401.01010
2024 arXiv
-
[36]
Feature-prompting GBMSeg: One- Shot Reference Guided Training-Free Prompt Engi- neering for Glomerular Basement Membrane Seg- mentation
Xueyu Liu et al. Feature-prompting GBMSeg: One- Shot Reference Guided Training-Free Prompt Engi- neering for Glomerular Basement Membrane Seg- mentation. 2024. arXiv: 2406 . 16271 [cs.CV]. url: https://arxiv.org/abs/2406.16271
2024 arXiv
-
[37]
SAM-RSIS: Progressively Adapt- ing SAM With Box Prompting to Remote Sensing Im- age Instance Segmentation
Muying Luo et al. “SAM-RSIS: Progressively Adapt- ing SAM With Box Prompting to Remote Sensing Im- age Instance Segmentation”. In: IEEE Transactions on Geoscience and Remote Sensing 62 (2024), pp. 1–
2024
-
[38]
Point-supervised Brain Tumor Seg- mentation with Box-prompted MedSAM
Xiaofeng Liu et al. Point-supervised Brain Tumor Seg- mentation with Box-prompted MedSAM . 2024. arXiv: 2408 . 00706 [cs.CV]. url: https : / / arxiv . org / abs/2408.00706. 17
2024 arXiv
-
[39]
CLISC: Bridging clip and sam by enhanced cam for unsupervised brain tumor segmenta- tion
Xiaochuan Ma et al. CLISC: Bridging clip and sam by enhanced cam for unsupervised brain tumor segmenta- tion. 2025. arXiv: 2501.16246 [cs.CV]. url: https: //arxiv.org/abs/2501.16246
2025 arXiv
-
[40]
Self-Prompting Polyp Seg- mentation in Colonoscopy using Hybrid Yolo-SAM 2 Model
Mobina Mansoori et al. Self-Prompting Polyp Seg- mentation in Colonoscopy using Hybrid Yolo-SAM 2 Model. 2024. arXiv: 2409 . 09484 [eess.IV]. url: https://arxiv.org/abs/2409.09484
2024 arXiv
-
[41]
doi: 10.1109/TGRS.2024.3460085
2024
-
[42]
GroupPrompter: A Prompting Method for Semantic Segmentation Based on SAM
Yichuang Luo et al. “GroupPrompter: A Prompting Method for Semantic Segmentation Based on SAM”. In: IEEE Access 11 (2023), pp. 106054–106062. doi: 10.1109/ACCESS.2023.3319740
2023
-
[43]
Hyper- correlation Squeeze for Few-Shot Segmentation
Juhong Min, Dahyun Kang, and Minsu Cho. Hyper- correlation Squeeze for Few-Shot Segmentation . 2021. arXiv: 2104.01538 [cs.CV]. url: https://arxiv. org/abs/2104.01538
2021 arXiv
-
[44]
CycleSAM: One-Shot Surgical Scene Segmentation using Cycle-Consistent Feature Matching to Prompt SAM
Aditya Murali et al. CycleSAM: One-Shot Surgical Scene Segmentation using Cycle-Consistent Feature Matching to Prompt SAM . 2024. arXiv: 2407.06795 [cs.CV]. url: https : / / arxiv . org / abs / 2407 . 06795
2024 arXiv
-
[45]
Label Anything: Multi- Class Few-Shot Semantic Segmentation with Visual Prompts
Pasquale De Marinis et al. Label Anything: Multi- Class Few-Shot Semantic Segmentation with Visual Prompts. 2024. arXiv: 2407 . 02075 [cs.CV]. url: https://arxiv.org/abs/2407.02075
2024 arXiv
-
[46]
Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medi- cal Image Segmentation
Juzheng Miao et al. Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medi- cal Image Segmentation . 2024. arXiv: 2407 . 05416 [cs.CV]. url: https : / / arxiv . org / abs / 2407 . 05416
2024
-
[47]
Benchmarking Human and Au- tomated Prompting in the Segment Anything Model
Jorge Quesada et al. Benchmarking Human and Au- tomated Prompting in the Segment Anything Model
-
[48]
Learning Transferable Visual Mod- els From Natural Language Supervision
Alec Radford et al. Learning Transferable Visual Mod- els From Natural Language Supervision . 2021. arXiv: 2103 . 00020 [cs.CV]. url: https : / / arxiv . org / abs/2103.00020
2021 arXiv
-
[49]
Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation
Saiyang Na et al. Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation. 2024. arXiv: 2401 . 13220 [eess.IV]. url: https://arxiv.org/abs/2401.13220
2024 arXiv
-
[50]
SAMIC: Segment Anything with In-Context Spatial Prompt Engineering
Savinay Nagendra et al. SAMIC: Segment Anything with In-Context Spatial Prompt Engineering . 2024. arXiv: 2412.11998 [cs.CV]. url: https://arxiv. org/abs/2412.11998
2024 arXiv
-
[51]
Temporally-Extended Prompts Optimization for SAM in Interactive Medical Im- age Segmentation
Chuyun Shen et al. “Temporally-Extended Prompts Optimization for SAM in Interactive Medical Im- age Segmentation”. In: 2023 IEEE International Con- ference on Bioinformatics and Biomedicine (BIBM) . 2023, pp. 3550–3557. doi: 10.1109/BIBM58861.2023. 10385291
2023
- [52]
-
[53]
EP-SAM: Weakly Supervised Histopathology Segmentation via Enhanced Prompt with Segment Anything
Joonhyeon Song et al. EP-SAM: Weakly Supervised Histopathology Segmentation via Enhanced Prompt with Segment Anything . 2024. arXiv: 2410 . 13621 [cs.CV]. url: https : / / arxiv . org / abs / 2410 . 13621
2024
-
[54]
PP-SAM: Perturbed Prompts for Robust Adaptation of Segment Anything Model for Polyp Segmentation
Md Mostafijur Rahman et al. PP-SAM: Perturbed Prompts for Robust Adaptation of Segment Anything Model for Polyp Segmentation . 2024. arXiv: 2405 . 16740 [cs.CV]. url: https : / / arxiv . org / abs / 2405.16740
2024 arXiv
-
[55]
Vision and Language Reference Prompt into SAM for Few-shot Segmentation
Kosuke Sakurai, Ryotaro Shimizu, and Masayuki Goto. Vision and Language Reference Prompt into SAM for Few-shot Segmentation . 2025. arXiv: 2502. 00719 [cs.CV]. url: https : / / arxiv . org / abs / 2502.00719
2025 arXiv
-
[56]
Sam2Rad: A Segmentation Model for Medical Images with Learnable Prompts
Assefa Seyoum Wahd et al. Sam2Rad: A Segmentation Model for Medical Images with Learnable Prompts
-
[57]
Adaptive Prompt Learning with SAM for Few-shot Scanning Probe Microscope Image Seg- mentation
Yao Shen et al. Adaptive Prompt Learning with SAM for Few-shot Scanning Probe Microscope Image Seg- mentation. 2024. arXiv: 2410 . 12562 [cs.CV]. url: https://arxiv.org/abs/2410.12562
2024 arXiv
-
[58]
CogVLM: Visual Expert for Pre- trained Language Models
Weihan Wang et al. CogVLM: Visual Expert for Pre- trained Language Models . 2024. arXiv: 2311 . 03079 [cs.CV]. url: https : / / arxiv . org / abs / 2311 . 03079
2024
-
[59]
Deep High-Resolution Representation Learning for Human Pose Estimation
Ke Sun et al. Deep High-Resolution Representation Learning for Human Pose Estimation . 2019. arXiv: 1902 . 09212 [cs.CV]. url: https : / / arxiv . org / abs/1902.09212
2019 arXiv
-
[60]
On Efficient Variants of Segment Anything Model: A Survey
Xiaorui Sun et al. On Efficient Variants of Segment Anything Model: A Survey . 2024. arXiv: 2410.04960 [cs.CV]. url: https : / / arxiv . org / abs / 2410 . 04960
2024 arXiv
-
[61]
TinyViT: Fast Pretraining Distillation for Small Vision Transformers
Kan Wu et al. TinyViT: Fast Pretraining Distillation for Small Vision Transformers . 2022. arXiv: 2207 . 10666 [cs.CV]. url: https : / / arxiv . org / abs / 2207.10666
2022 arXiv
- [62]
-
[63]
Auto-Prompting SAM for Weakly Supervised Landslide Extraction
Jian Wang et al. Auto-Prompting SAM for Weakly Supervised Landslide Extraction . 2025. arXiv: 2501 . 13426 [cs.CV]. url: https : / / arxiv . org / abs / 2501.13426
2025 arXiv
-
[64]
Self-Prompt SAM: Medical Image Seg- mentation via Automatic Prompt SAM Adaptation
Bin Xie et al. Self-Prompt SAM: Medical Image Seg- mentation via Automatic Prompt SAM Adaptation
-
[65]
CrackESS: A Self-Prompting Crack Segmentation System for Edge Devices
Yingchu Wang, Ji He, and Shijie Yu. CrackESS: A Self-Prompting Crack Segmentation System for Edge Devices. 2025. arXiv: 2412 . 07205 [cs.CV]. url: https://arxiv.org/abs/2412.07205. 18
2025 arXiv
-
[66]
Optimizing Prompt Strategies for SAM: Advancing lesion Segmentation Across Diverse Medical Imaging Modalities
Yuli Wang et al. Optimizing Prompt Strategies for SAM: Advancing lesion Segmentation Across Diverse Medical Imaging Modalities. 2024. arXiv: 2412.17943 [eess.IV]
2024 arXiv
-
[67]
ESP-MedSAM: Efficient Self- Prompting SAM for Universal Domain-Generalized Medical Image Segmentation
Qing Xu et al. ESP-MedSAM: Efficient Self- Prompting SAM for Universal Domain-Generalized Medical Image Segmentation . 2024. arXiv: 2407 . 14153 [eess.IV]. url: https : / / arxiv . org / abs / 2407.14153
2024 arXiv
-
[68]
Self- Prompting Large Vision Models for Few-Shot Med- ical Image Segmentation
Qi Wu, Yuyao Zhang, and Marawan Elbatel. Self- Prompting Large Vision Models for Few-Shot Med- ical Image Segmentation . 2023. arXiv: 2308 . 07624 [cs.CV]. url: https : / / arxiv . org / abs / 2308 . 07624
2023
-
[69]
Integrating multi-scale informa- tion and diverse prompts in large model SAM-Med2D for accurate left ventricular ejection fraction estima- tion
Yagang Wu et al. “Integrating multi-scale informa- tion and diverse prompts in large model SAM-Med2D for accurate left ventricular ejection fraction estima- tion”. In: Medical & Biological Engineering & Com- puting (Feb. 14, 2025).issn: 1741-0444. doi: 10.1007/ s11517- 025- 03...
2025
-
[70]
PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Im- age Segmentation
Zhonghao Yan et al. PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Im- age Segmentation. 2025. arXiv: 2501.06692 [cs.CV]. url: https://arxiv.org/abs/2501.06692
2025 arXiv
-
[71]
TA VP: Task-Adaptive Visual Prompt for Cross-domain Few-shot Segmentation
Jiaqi Yang et al. TA VP: Task-Adaptive Visual Prompt for Cross-domain Few-shot Segmentation. 2024. arXiv: 2409 . 05393 [cs.CV]. url: https : / / arxiv . org / abs/2409.05393
2024 arXiv
-
[72]
Char-SAM: Turning Segment Any- thing Model into Scene Text Segmentation Annota- tor with Character-level Visual Prompts
Enze Xie et al. Char-SAM: Turning Segment Any- thing Model into Scene Text Segmentation Annota- tor with Character-level Visual Prompts . 2024. arXiv: 2412 . 19917 [cs.CV]. url: https : / / arxiv . org / abs/2412.19917
2024 arXiv
-
[73]
SAM-MPA: Applying SAM to Few- shot Medical Image Segmentation using Mask Propa- gation and Auto-prompting
Jie Xu et al. SAM-MPA: Applying SAM to Few- shot Medical Image Segmentation using Mask Propa- gation and Auto-prompting . 2024. arXiv: 2411.17363 [cs.CV]. url: https : / / arxiv . org / abs / 2411 . 17363
2024 arXiv
-
[74]
SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation
Wenxi Yue et al. SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation . 2023. arXiv: 2308.08746 [cs.CV]. url: https://arxiv. org/abs/2308.08746
2023 arXiv
-
[77]
COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection
Xiaoqin Zhang et al. COMPrompter: reconceptualized segment anything model with multiprompt network for camouflaged object detection. 2024. arXiv: 2411.18858 [cs.CV]. url: https : / / arxiv . org / abs / 2411 . 18858
2024 arXiv
-
[78]
UV-SAM: Adapting Segment Any- thing Model for Urban Village Identification
Xin Zhang et al. UV-SAM: Adapting Segment Any- thing Model for Urban Village Identification . 2024. arXiv: 2401.08083 [cs.CV]. url: https://arxiv. org/abs/2401.08083
2024 arXiv
-
[79]
Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization
Xi Yang et al. “Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization”. In: Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LXIX. Milan, Italy: Springer- Verlag...
2024 doi
-
[80]
SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Seg- mentation
Wenxi Yue et al. SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Seg- mentation. 2024. arXiv: 2312 . 14481 [cs.CV]. url: https://arxiv.org/abs/2312.14481
2024 arXiv
-
[81]
Automatic Seg- mentation Annotation of Space Target Using Segment Anything Model and Object Detection Prompts
Zhihao Zhang and Zhaohui Dang. “Automatic Seg- mentation Annotation of Space Target Using Segment Anything Model and Object Detection Prompts”. In: IEEE Transactions on Aerospace and Electronic Sys- tems (2024), pp. 1–15. doi: 10 . 1109 / TAES . 2024 . 3512533. 19
2024
-
[82]
A Survey on Segment Any- thing Model (SAM): Vision Foundation Model Meets Prompt Engineering
Chaoning Zhang et al. A Survey on Segment Any- thing Model (SAM): Vision Foundation Model Meets Prompt Engineering . 2024. arXiv: 2306 . 06211 [cs.CV]. url: https : / / arxiv . org / abs / 2306 . 06211
2024
-
[83]
A Comprehensive Survey on Seg- ment Anything Model for Vision and Beyond
Chunhui Zhang et al. A Comprehensive Survey on Seg- ment Anything Model for Vision and Beyond . 2023. arXiv: 2305.08196 [cs.CV]. url: https://arxiv. org/abs/2305.08196
2023 arXiv
-
[84]
Curriculum Prompting Foundation Models for Medical Image Segmentation
Xiuqi Zheng et al. Curriculum Prompting Foundation Models for Medical Image Segmentation . 2024. arXiv: 2409 . 00695 [cs.CV]. url: https : / / arxiv . org / abs/2409.00695
2024 arXiv
-
[85]
EdgeSAM: Prompt-In-the-Loop Distillation for On-Device Deployment of SAM
Chong Zhou et al. EdgeSAM: Prompt-In-the-Loop Distillation for On-Device Deployment of SAM . 2024. arXiv: 2312.06660 [cs.CV]. url: https://arxiv. org/abs/2312.06660
2024 arXiv
-
[86]
Towards Segment Any- thing Model (SAM) for Medical Image Segmentation: A Survey
Yichi Zhang and Rushi Jiao. Towards Segment Any- thing Model (SAM) for Medical Image Segmentation: A Survey . 2023. arXiv: 2305.03678 [eess.IV]. url: https://arxiv.org/abs/2305.03678
2023 arXiv
-
[87]
Enhancing the Reliability of Seg- ment Anything Model for Auto-Prompting Medical Im- age Segmentation with Uncertainty Rectification
Yichi Zhang et al. Enhancing the Reliability of Seg- ment Anything Model for Auto-Prompting Medical Im- age Segmentation with Uncertainty Rectification. 2024. arXiv: 2311.10529 [cs.CV]. url: https://arxiv. org/abs/2311.10529
2024 arXiv
-
[88]
Segment Everything Everywhere All at Once
Xueyan Zou et al. Segment Everything Everywhere All at Once . 2023. arXiv: 2304.06718 [cs.CV]. url: https://arxiv.org/abs/2304.06718. 20
2023 arXiv
-
[89]
Semantic-Enhanced Point-Box Joint Prompting for Video Object Segmentation
Quan Zhao et al. “Semantic-Enhanced Point-Box Joint Prompting for Video Object Segmentation”. In: 2024 IEEE International Conference on Image Pro- cessing (ICIP) . 2024, pp. 2334–2340. doi: 10.1109/ ICIP51287.2024.10648107
2024
-
[90]
Fast Segment Anything
Xu Zhao et al. Fast Segment Anything . 2023. arXiv: 2306 . 12156 [cs.CV]. url: https : / / arxiv . org / abs/2306.12156
2023 arXiv
-
[93]
MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAM
Nan Zhou et al. MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAM. 2024. arXiv: 2409.00924 [cs.CV]. url: https://arxiv. org/abs/2409.00924
2024 arXiv
-
[94]
ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual De- scriptions
Deyao Zhu et al. ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual De- scriptions. 2023. arXiv: 2303 . 06594 [cs.CV]. url: https://arxiv.org/abs/2303.06594
2023 arXiv
- [2024]
- [2025]
-
[3580]
doi: 10.1109/WACV61041.2025.00352
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.