REVIEW 36 cited by
Personalize Segment Anything Model with One Shot
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Driven by large-data pre-training, Segment Anything Model (SAM) has been demonstrated as a powerful and promptable framework, revolutionizing the segmentation models. Despite the generality, customizing SAM for specific visual concepts without man-powered prompting is under explored, e.g., automatically segmenting your pet dog in different images. In this paper, we propose a training-free Personalization approach for SAM, termed as PerSAM. Given only a single image with a reference mask, PerSAM first localizes the target concept by a location prior, and segments it within other images or videos via three techniques: target-guided attention, target-semantic prompting, and cascaded post-refinement. In this way, we effectively adapt SAM for private use without any training. To further alleviate the mask ambiguity, we present an efficient one-shot fine-tuning variant, PerSAM-F. Freezing the entire SAM, we introduce two learnable weights for multi-scale masks, only training 2 parameters within 10 seconds for improved performance. To demonstrate our efficacy, we construct a new segmentation dataset, PerSeg, for personalized evaluation, and test our methods on video object segmentation with competitive performance. Besides, our approach can also enhance DreamBooth to personalize Stable Diffusion for text-to-image generation, which discards the background disturbance for better target appearance learning. Code is released at https://github.com/ZrrSkywalker/Personalize-SAM
Forward citations
Cited by 36 Pith papers
-
SAMIC: Segment Anything with In-Context Spatial Prompt Engineering
A 2.6-million-parameter learned prompt generator for SAM achieves state-of-the-art or competitive one-shot segmentation on multiple benchmarks using only 20% of the training data.
-
OGG-FR: Orthogonal Gradient Gaming and Frequency Rectification for Unmanned Aerial Vehicle Infrared Image Super-Resolution
OGG-FR is a plug-and-play training update that separates redundant and innovative parts of the FFT loss gradient and gates the innovative part by a confidence score, improving UAV infrared super-resolution in most tes...
-
Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation
MSSA improves test-time medical image segmentation by storing reliable vision-language predictions in a memory bank and using stored images as prototypes to segment new images, without updating model weights.
-
GFR-SAM: Training-Free Referring Camouflaged Object Segmentation via Cross-Image Prompting
A Generate-Filter-Refine pipeline unlocks SAM3 for training-free Ref-COD, beating prior zero-shot methods by 8.7 points F_β^w on R2C7K and rivaling supervised specialists.
-
Repurposing CLIP to Localize at Pixel Level
CLIPix repurposes CLIP by tracing classification activations, applying noise-resistant correction, and localization embedding to reach SOTA zero-shot binary open-set segmentation on PASCAL-5i and COCO-20i.
-
Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking
A two-stage pure RL method with an information-gap global view and hierarchical grounding loss makes MLLMs truly rely on precise crops and sets SOTA on high-res VQA under tight token budgets.
-
CAD-Prompted SAM3: Geometry-Conditioned Instance Segmentation for Industrial Objects
Multi-view CAD renderings as prompts let SAM3 segment unseen industrial objects by geometry alone, beating appearance-exemplar baselines on 3D-printing and industrial benchmarks.
-
One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt Evolution
OP-SAM turns one labeled polyp image into iterative SAM prompts, reaching 76.93% IoU on Kvasir with no retraining.
-
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
A few-shot segmentation framework that injects reference-image prototypes into SAM's decoder and image encoder, eliminating per-image manual prompts and improving remote sensing segmentation accuracy.
-
ProSAM: Enhancing the Robustness of SAM-based Visual Reference Segmentation with Probabilistic Prompts
ProSAM improves SAM-based visual reference segmentation by learning a Student-t prompt distribution through noise injection, beating VRP-SAM by about one mIoU point on Pascal-5^i and COCO-20^i.
-
Leader360V: The Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment
Leader360V provides a 10,000+ video, 198-class, densely annotated 360-degree video dataset with an LLM-assisted automatic annotation pipeline, and shows fine-tuning on it improves 360 video segmentation and tracking models.
-
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
A new multi-image benchmark and graph-based scoring method show that MLLMs produce weak, poorly-grounded chains of thought on visual physics tasks, and that RL post-training can degrade spatial reasoning.
-
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
The Show or Tell benchmark compares text and visual prompts for semantic segmentation on 14 datasets, finding visual prompts yield higher average mIoU but with high variance and much higher computational cost.
-
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
GeoSense introduces GPI and GPA metrics and a 148-principle hierarchy to jointly measure identification and application of geometric principles in 1,789 bilingual geometry problems.
-
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
A prompting scheme that mines points, elastic boxes, and Gaussian-style masks from coarse masks lets SAM refine those masks more accurately than prior refinement tools.
-
Exploring Few-Shot Defect Segmentation in General Industrial Scenarios with Metric Learning and Vision Foundation Models
A new 12-product industrial benchmark shows foundation models, particularly SAM2 in video track mode, outperform meta-learning for few-shot defect segmentation.
-
SAM-guided Pseudo Label Enhancement for Multi-modal 3D Semantic Segmentation
SAM mask grouping plus majority-vote labeling and geometry-aware propagation densifies pseudo-labels and improves cross-domain 3D semantic segmentation accuracy.
-
SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot Segmentation
GPRN converts SAM-generated masks into graph-reasoned visual prompts and adds test-time SAM point refinement, setting new state-of-the-art results on four cross-domain few-shot segmentation benchmarks.
-
Med-PerSAM: One-Shot Visual Prompt Tuning for Personalized Segment Anything Model in Medical Domain
A one-shot SAM prompting framework that uses test-time image warping to generate mask, point, and box prompts, achieving the highest reported DICE scores among the compared baselines across five medical datasets.
-
Segment Any Class (SAC): Multi-Class Few-Shot Semantic Segmentation via Class Region Proposals
A training-free prompting method that adapts SAM to multi-class few-shot segmentation and reports higher mIoU than trained baselines on COCO-20i as the number of classes grows.
-
Teaching VLMs to Localize Specific Objects from In-context Examples
Fine-tuning VLMs on video-tracking conversations with made-up object names teaches them to localize a specific object in a new image from only a few in-context examples.
-
Slender Object Scene Segmentation in Remote Sensing Image Based on Learnable Morphological Skeleton with Segment Anything Model
A variational morphological skeleton prior, smoothed via log-sum-exp and unrolled into SAM's mask decoder, improves slender object segmentation in remote sensing images by roughly 1-2 F1 points.
-
Proxy Prompt: Endowing SAM and SAM 2 with Auto-Interactive-Prompt for Medical Segmentation
Proxy Prompt lets frozen SAM and SAM 2 segment new medical images and videos using a high-dimensional prompt auto-generated from a non-target image-mask pair.
-
Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
SITN improves cross-domain few-shot detection and segmentation by using weakened-noise diffusion and background inpainting to synthesize helpful training images, outperforming prior methods on all reported benchmarks.
-
Region-aware Depth Scale Adaptation with Sparse Measurements
A non-learning method segments an image and gives each region its own scale and shift, fitted to a few sparse depth points, to turn relative monocular depth predictions into metric depth more accurately than a single ...
-
Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
MA-SAM2 adds context-aware and occlusion-resilient memory to SAM2 and reports Challenge IoU of 62.49 percent on EndoVis2017 and 64.40 percent on EndoVis2018, beating SAM2 by 6.10 and 4.36 points.
-
CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning
A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.
-
SynPo: Boosting Training-Free Few-Shot Medical Segmentation via High-Quality Negative Prompts
A training-free SAM/DINOv2 pipeline with Gaussian-sampled negative prompts achieves near state-of-the-art few-shot medical segmentation on CT and MRI abdominal data.
-
AoP-SAM: Automation of Prompts for Efficient Segmentation
AoP-SAM trains a lightweight prompt predictor on SAM's own image embeddings to emit a point-prompt confidence map, then filters redundant prompts at test time, improving automatic segmentation accuracy and efficiency.
-
Causal Prompt Calibration Guided Segment Anything Model for Open-Vocabulary Multi-Entity Segmentation
CPC-SAM reweights random prompts to enforce segmentation consistency across prompt variants, claiming this yields causal prompts that improve open-vocabulary multi-entity segmentation with SAM.
-
Vision and Language Reference Prompt into SAM for Few-shot Segmentation
Using a frozen vision-language model, VLP-SAM injects text-label semantics into SAM's prompt encoder and raises one-shot segmentation mIoU by 6.3 points on PASCAL-5i and 9.5 on COCO-20i.
-
Multi-Granularity Video Object Segmentation
A new large-scale video dataset with dense masks for objects, parts, and backgrounds, plus a memory-based SAM video segmentation model that wins on that benchmark.
-
SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.
-
PlantSAM: An Object Detection-Driven Segmentation Pipeline for Herbarium Specimens
PlantSAM, a pipeline that uses YOLOv10 bounding boxes to prompt SAM2, segments plants from herbarium backgrounds with IoU 0.94 and improves trait classification accuracy by up to 4.36%.
-
SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation
SRMF is a segmentation framework for UHR satellite images that adds scale-anchored cropping, SAM-HQ based tail-class resampling, and GeoRSCLIP text feature injection, reporting mIoU gains of 3.33, 0.66, and 0.98 on UR...
-
SAM-MPA: Applying SAM to Few-shot Medical Image Segmentation using Mask Propagation and Auto-prompting
A method that combines clustering, B-spline registration, and automatic prompting to adapt SAM for few-shot medical image segmentation with as few as 1 labeled image.
Discussion (0). Continue with ORCID to comment.