REVIEW 80 cited by
Fast Segment Anything
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The recently proposed segment anything model (SAM) has made a significant influence in many computer vision tasks. It is becoming a foundation step for many high-level tasks, like image segmentation, image caption, and image editing. However, its huge computation costs prevent it from wider applications in industry scenarios. The computation mainly comes from the Transformer architecture at high-resolution inputs. In this paper, we propose a speed-up alternative method for this fundamental task with comparable performance. By reformulating the task as segments-generation and prompting, we find that a regular CNN detector with an instance segmentation branch can also accomplish this task well. Specifically, we convert this task to the well-studied instance segmentation task and directly train the existing instance segmentation method using only 1/50 of the SA-1B dataset published by SAM authors. With our method, we achieve a comparable performance with the SAM method at 50 times higher run-time speed. We give sufficient experimental results to demonstrate its effectiveness. The codes and demos will be released at https://github.com/CASIA-IVA-Lab/FastSAM.
Forward citations
Showing 60 of 80 Pith papers that cite this
-
Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation
A point-prompted segmentation model gains efficiency by foveated tokenization, cutting tokens from 4096 to 172 while staying competitive on mIoU benchmarks.
-
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
Introduces PL-VEL, a pixel-mask-based visual entity linking task, and MaskOVEN-Wiki, a 5.2M-annotation dataset built via reverse annotation, plus a semantic tokenization method that yields a 5-point accuracy gain.
-
TestMate: Test-Time Domain Adaptation Aided by Lightweight Vision Foundation Model
TestMate fuses FastSAM mask proposals with a segmentation network via size-ordered soft refinement to achieve backpropagation-free, first-frame TTDA gains on semantic segmentation benchmarks.
-
Zero-OVCD: Bridging Training-Free Foundation Models and Pseudo-Label Learning for Open-Vocabulary Change Detection
Zero-OVCD generates change pseudo-labels from SAM3, DINOv3, and SegEarth-OV3, then trains a change detector on them, lifting F1 to 88.65%, 88.85%, and 57.96% on LEVIR-CD, WHU-CD, and S2Looking without target-domain pi...
-
GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
A 4D scene graph augmented with atomic human-object interactions and goal-driven events improves retrospective question answering about human activities in dynamic scenes.
-
From Transparent Labware Segmentation to Collision Avoidance: A Real-Time Edge-Aware Perception Pipeline
An edge-aware YOLOv5-Seg variant, together with a new real-world dataset, segments transparent labware in real time and enables conservative 3D collision avoidance for robots.
-
AtlasLC: Fast Codec-Ready Compression of Object-Centric 3D Gaussian Splatting
A training-free pipeline prunes object-centric 3D Gaussian splats by local competition and packs them into deterministic codec-ready atlases, cutting preparation time and payload with modest quality loss.
-
MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes
MV-GEL uses a learned view selector and a fine-tuned vision-language segmentation model to localize text-described faces and edges on 3D meshes.
-
StateScribe: Towards Accessible Change Awareness Across Real-World Revisits
StateScribe uses a dual-layer memory architecture for episodic scenes and object-centric changes to deliver live and historical descriptions, achieving 83.1% F1 accuracy across revisits in evaluations and user studies...
-
Active Semantic Perception
An LLM-based scene-graph completion routine can guide a robot to infer and find unseen rooms faster than frontier-based exploration, at least in the three simulated apartments tested.
-
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
AIM-CoT improves multimodal chain-of-thought by selecting image regions that reduce predictive uncertainty and inserting them when attention shifts toward the visual input.
-
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
DreamNav achieves new zero-shot SOTA on VLN-CE with an egocentric-only pipeline that generates candidate trajectories, imagines their futures, and selects the best by language alignment.
-
Probabilistic Human Intent Prediction for Mobile Manipulation: An Evaluation with Human-Inspired Constraints
GUIDER couples navigation-level and manipulation-level probabilistic intent beliefs using map context, visual saliency, grasp-feasibility checks, and end-effector kinematics, and reports higher prediction stability th...
-
Learning human-to-robot handovers through 3D scene reconstruction
A handover policy trained only on images rendered from a sparse-view Gaussian Splatting scene can deploy on a real robot without real-robot training data.
-
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
MTU3D unifies visual grounding and frontier-based exploration in a single transformer, achieving state-of-the-art success rates on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA after large-scale vision-language-exploration p...
-
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
A few-shot segmentation framework that injects reference-image prototypes into SAM's decoder and image encoder, eliminating per-image manual prompts and improving remote sensing segmentation accuracy.
-
LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment
LoD-Loc v2 localizes aerial cameras by aligning predicted building silhouettes with rendered low-detail city-model silhouettes, achieving accurate 4-DoF pose without textured maps.
-
SAM4D: Segment Anything in Camera and LiDAR Streams
SAM4D is a promptable model that segments and tracks objects across camera and LiDAR streams with cross-modal prompts, trained on pseudo-labels generated by an automated data engine.
-
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
A zero-training framework that adaptively selects spatial representation extractors per object and per task stage improves real-world robot manipulation success and efficiency over fixed-representation baselines.
-
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
A SAM2-based framework that uses a fused text-audio-visual token to prompt video segmentation achieves 58.5 J&F on Ref-AVS, outperforming the previous state of the art by 8.5 points.
-
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.
-
Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
A pseudo-supervised pipeline with an affordance-to-part mapping, label refinement, cross-view alignment, and a reasoning module achieves state-of-the-art weakly supervised affordance grounding on AGD20K.
-
RoboCulture: A Robotics Platform for Automated Biological Experimentation
RoboCulture couples a general-purpose robot arm with vision-based pipetting, force-guided tip exchange, and behavior-tree decisions to run a 15-hour yeast culture experiment with automated splitting of saturated wells.
-
Sketch Interface for Teleoperation of Mobile Manipulator to Enable Intuitive and Intended Operation: A Proof of Concept
A sketch-based interface for teleoperating a mobile manipulator lowered perceived workload and increased intuitiveness compared with axis-button control in a proof-of-concept study.
-
L2D2: Robot Learning from 2D Drawings
L2D2 lets humans teach robot tasks by drawing on synthetic images of varied scenes and adding a few physical corrections, achieving teleop-level policy performance with less user effort.
-
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
Dynam3D represents scenes as patch, instance, and zone tokens that update dynamically, and feeds them to a 3.8B vision-language model to improve action prediction in vision-and-language navigation.
-
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation
A zero-shot referring expression segmentation method that uses attribute prompting to reason about object parts and implicit descriptions, outperforming prior zero-shot and several fine-tuned baselines.
-
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
Cues3D achieves view-consistent, globally unique 3D instance IDs by training a NeRF in three phases and correcting IDs from NeRF-rendered 3D masks, without pre-association or contrastive loss.
-
AffordanceSAM: Segment Anything Once More in Affordance Grounding
Adapting EVF-SAM with learnable affordance queries and a coarse-to-fine dataset yields strong affordance grounding on AGD20K, with caveats about test-set tuning.
-
SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
SketchFlex combines sketch-aware prompt recommendation with decompose-and-recompose shape refinement to help novices generate multi-object images from rough region sketches.
-
AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
A camera-based pipeline that automatically produces open-ended 3D semantic occupancy labels via vision-language attention maps and Gaussian splatting, outperforming existing auto-labeling methods.
-
Exploring Few-Shot Defect Segmentation in General Industrial Scenarios with Metric Learning and Vision Foundation Models
A new 12-product industrial benchmark shows foundation models, particularly SAM2 in video track mode, outperform meta-learning for few-shot defect segmentation.
-
FoundationStereo: Zero-Shot Stereo Matching
A new stereo depth model, trained on 1M synthetic pairs with adapted monocular features, reports strong zero-shot accuracy on multiple real-world benchmarks without target-domain fine-tuning.
-
AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label Correction
AdaCo refurbishes noisy VFM-generated 3D pseudo labels using adaptive early-learning correction and robust losses, reaching 25.7% and 31.2% mIoU on SemanticKITTI and nuScenes without 3D annotations.
-
POEX: Towards Policy Executable Jailbreak Attacks Against the LLM-based Robots
POEX generates short adversarial suffixes that make LLM-based robots turn harmful instructions into executable robot policies, with about 60% average execution success across tested models.
-
Stereo Hand-Object Reconstruction for Human-to-Robot Handover
StereoHO combines two RGB views of a hand holding an object into a joint 3D reconstruction via learned shape codebooks, enabling robots to grasp diverse household objects, including transparent ones.
-
SparseGrasp: Robotic Grasping via 3D Semantic Gaussian Splatting from Sparse Multi-View RGB Images
A language-guided robotic grasping system that reconstructs a 3D semantic scene from three RGB views and updates moved objects in about 200 ms, reporting higher grasp success than F3RM and LERF-TOGO.
-
UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation
UrbanCAD retrieves a matching CAD model from a single car image, optimizes its materials, and inserts it into reconstructed urban scenes, showing that perception models degrade when the cars are edited into out-of-dis...
-
UNOPose: Unseen Object Pose Estimation with an Unposed RGB-D Reference Image
A single unposed RGB-D reference image is enough to estimate the 6D pose of an unseen object, outperforming prior reference-based methods on BOP datasets.
-
Bringing the Context Back into Object Recognition, Robustly
Localizing the foreground before classification and fusing its classifier output with the full-image prediction improves accuracy and robustness to background shifts in supervised and zero-shot VLM recognition.
-
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
Training-free CorrCLIP reconstructs patch correlations in CLIP with SAM masks and DINO similarity, raising averaged mIoU across eight benchmarks from 48.6 to 53.6.
-
Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2
Lean-SAM2 combines target-anchored memory pruning, condensed insurance memory, and risk-aware window routing to accelerate SAM2.1 inference ~1.4× with better accuracy than Efficient-SAM2.
-
MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors
A joint multi-view self-attention matcher over SAM segments beats pairwise matchers at wide baselines and lifts HM3D navigation success from 50% to 70%, while a LightGlue-style pairwise head wins outdoors and at narro...
-
Analysis of the Dick Effect for AI-based Dynamic Gravimeter
A 0.12 s accelerometer dead time in an atom-interferometer dynamic gravimeter introduces roughly 8 mGal of measurement noise, which the paper attributes to high-frequency aliasing and analyzes with a derived frequency...
-
DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory
A zero-shot VLM navigation policy with dynamic waypoint selection and graph memory reports state-of-the-art results among VLM-based methods on ObjectNav and GOAT-Bench, plus real-world tests on a quadruped.
-
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
A personalized, proactive LLM planner suggests helpful next actions during human-robot lunch packing, reporting 38.7% faster task execution at the cost of an extra 5.6-minute setup phase.
-
Accurate and efficient zero-shot 6D pose estimation with frozen foundation models
A training-free 6D pose estimator using sparse-to-dense matching of frozen foundation model features achieves new state-of-the-art results on BOP with large speedups.
-
Dynamic Sub-region Search in Homogeneous Collections Using CLIP
Dynamic region detection plus IoU-based geometric constraints roughly doubles recall over static grid partitioning for underwater known-item search, given accurate query boxes.
-
Towards Terrain-Aware Task-Driven 3D Scene Graph Generation in Outdoor Environments
An outdoor 3D scene graph pipeline using LiDAR-camera fusion, CLIP embeddings, and per-terrain Voronoi graphs is demonstrated on a campus dataset with qualitative results.
-
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
PAM extends SAM 2 with a frozen LLM and a Semantic Perceiver to jointly segment and describe regions in images, videos, and streaming video, and contributes a 0.6M-sample region-level streaming video caption dataset.
-
Experimental Study on Automatically Assembling Custom Catering Packages With a 3-DOF Delta Robot Using Deep Learning Methods
A YOLOv5-plus-FastSAM pipeline with eigenvector-derived grasp points lets a Delta robot assemble custom catering packages with about 82% physical grasping success.
-
Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
MatchCD performs homography-based registration and building change detection on large unregistered bi-temporal remote sensing images using contrastive pre-training, frozen matching, and prior masks from FastSAM.
-
Image Embedding Sampling Method for Diverse Captioning
A training-free hierarchical embedding sampling method (HBoP) lets a small BLIP model generate captions as diverse as human ones, beating much larger VLMs on diversity metrics.
-
Vision and Language Reference Prompt into SAM for Few-shot Segmentation
Using a frozen vision-language model, VLP-SAM injects text-label semantics into SAM's prompt encoder and raises one-shot segmentation mIoU by 6.3 points on PASCAL-5i and 9.5 on COCO-20i.
-
Lifting by Gaussians: A Simple, Fast and Flexible Method for 3D Instance Segmentation
LBG segments 3D Gaussian Splatting scenes into objects, parts, and subparts by assigning each pixel's maximum-contributing Gaussian a 2D mask ID and merging fragments across frames using geometric and semantic similarity.
-
GFreeDet: Exploiting Gaussian Splatting and Foundation Models for Model-free Unseen Object Detection in the BOP Challenge 2024
Model-free unseen object detection that reconstructs objects as 3D Gaussians from onboarding videos and matches SAM proposals to rendered templates with DINOv2, reaching 31.9% AP on BOP-H3 without CAD models.
-
SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory
A training-free adaptation of SAM 2 that adds Kalman-filter motion scoring and motion-aware memory selection improves zero-shot visual tracking across multiple benchmarks.
-
Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data
Adding monocular depth to EfficientViT-SAM improves point-prompted segmentation at 3 and 5 clicks after fine-tuning on 11.2k images, but universal gains and data-efficiency are not established.
-
SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.
-
Multi-modal video data-pipelines for machine learning with minimal human supervision
An open-source video pipeline automatically extracts 13+ visual modalities from raw video with no human annotation, and a sub-1M-parameter distilled model reaches near-Mask2Former accuracy on an aerial scene benchmark.
Discussion (0). Continue with ORCID to comment.