Pith. sign in

REVIEW 20 cited by

SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.11922 v2 pith:F5GEWC6F submitted 2024-11-18 cs.CV

classification cs.CV
keywords samuraitrackingobjectmemorymodelvisualachievesanything
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The Segment Anything Model 2 (SAM 2) has demonstrated strong performance in object segmentation tasks but faces challenges in visual object tracking, particularly when managing crowded scenes with fast-moving or self-occluding objects. Furthermore, the fixed-window memory approach in the original model does not consider the quality of memories selected to condition the image features for the next frame, leading to error propagation in videos. This paper introduces SAMURAI, an enhanced adaptation of SAM 2 specifically designed for visual object tracking. By incorporating temporal motion cues with the proposed motion-aware memory selection mechanism, SAMURAI effectively predicts object motion and refines mask selection, achieving robust, accurate tracking without the need for retraining or fine-tuning. SAMURAI operates in real-time and demonstrates strong zero-shot performance across diverse benchmark datasets, showcasing its ability to generalize without fine-tuning. In evaluations, SAMURAI achieves significant improvements in success rate and precision over existing trackers, with a 7.1% AUC gain on LaSOT$_{\text{ext}}$ and a 3.5% AO gain on GOT-10k. Moreover, it achieves competitive results compared to fully supervised methods on LaSOT, underscoring its robustness in complex tracking scenarios and its potential for real-world applications in dynamic environments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mitigating Error Accumulation in Continuous Navigation via Memory-Augmented Kalman Filtering

    cs.RO 2026-01 unverdicted novelty 7.0 of 10

    NeuroKalman mitigates state drift in vision-language UAV navigation by using memory-augmented Kalman filtering where attention retrieves historical anchors to correct predictions without gradient updates.

  2. 3AM: 3egment Anything with Geometric Consistency in Videos

    cs.CV 2026-01 unverdicted novelty 7.0 of 10

    3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.

  3. SAM 3: Segment Anything with Concepts

    cs.CV 2025-11 unverdicted novelty 7.0 of 10

    SAM 3 introduces promptable concept segmentation that doubles accuracy of prior systems on images and videos while improving standard SAM segmentation performance.

  4. Efficient Tracking and Understanding Object Transformations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    FluxGraph detects object transformations reactively via SAM2's multi-mask disagreement, cutting TubeletGraph's inference cost by 3.3–10.7x with comparable tracking and state-graph quality.

  5. SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A modified video segmentation architecture decouples processing latency from target count, enabling real-time (>36 FPS) tracking of 10+ objects simultaneously while preserving individual identities.

  6. SUMO: Segment and Track Any Motion with Nonlinear State Space Models

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    SUMO is a training-free unified framework using nonlinear SSM and Selective Unscented Filter for VOT and MOS, reporting SOTA results.

  7. ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ViewSAM achieves state-of-the-art weakly supervised performance on cross-view referring multi-object tracking by refining SAM tracklets via affinity-guided re-prompting and modeling view-induced variations as learnabl...

  8. HOIGS: Human-Object Interaction Gaussian Splatting

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    HOIGS adds a cross-attention HOI module to Gaussian Splatting that combines HexPlane human features with Cubic Hermite Spline object features to model interaction-induced deformations.

  9. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    HVG-3D uses a 3D-aware diffusion architecture with ControlNet to synthesize high-fidelity hand-object interaction videos from 3D control signals, achieving state-of-the-art spatial fidelity and temporal coherence on t...

  10. HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization

    cs.CV 2025-08 conditional novelty 6.0 of 10

    HERO-VQL combines top-down attention guidance with egocentric augmentations and consistency training, improving visual query localization accuracy on VQ2D over prior state-of-the-art methods.

  11. SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SAMITE improves zero-shot visual object tracking by selecting trustworthy past frames via prototype similarity and adding positional mask prompts, outperforming prior SAM2-based trackers on most benchmarks.

  12. Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Lean-SAM2 combines target-anchored memory pruning, condensed insurance memory, and risk-aware window routing to accelerate SAM2.1 inference ~1.4× with better accuracy than Efficient-SAM2.

  13. Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    TEP-SAM adapts SAM for multiframe infrared small target detection by generating temporal-emerged cues from joint global-local motion modeling.

  14. Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    SAMOSA adapts SAM 2 for complex visual object tracking by integrating explicit nonlinear motion prediction, semantic cues for failure recovery, and geometric constraints for stability, outperforming prior SAM 2-based ...

  15. 4D Vessel Reconstruction for Benchtop Thrombectomy Analysis

    eess.IV 2026-04 conditional novelty 5.0 of 10

    A nine-camera multi-view workflow with 4D Gaussian Splatting reconstructs dynamic vessel surfaces in thrombectomy phantoms to enable standardized comparative displacement and stress-proxy tracking.

  16. SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking

    cs.CV 2025-11 conditional novelty 5.0 of 10

    SAM2S, a SAM2 variant trained on the new 61k-frame SA-SV surgical benchmark, improves average J&F to 80.42 at 68 FPS for interactive surgical-video object segmentation.

  17. Zero-Shot Multi-Animal Tracking in the Wild

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A zero-shot multi-animal tracker combining Grounding DINO + SAM 2 with three hand-designed heuristics beats prior methods on four animal-tracking benchmarks with fixed hyperparameters.

  18. FreeVPS: Repurposing Training-Free SAM2 for Generalizable Video Polyp Segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    FreeVPS pairs a per-frame polyp segmenter with frozen SAM2 tracking and two filtering modules to reduce error accumulation, improving in-domain and out-of-domain video polyp segmentation.

  19. A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

    cs.RO 2025-07 unverdicted novelty 5.0 of 10

    The survey frames VLA models as pipelines that generate progressively grounded action tokens and classifies those tokens into eight types to guide future development.

  20. Cosmos World Foundation Model Platform for Physical AI

    cs.CV 2025-01 unverdicted novelty 3.0 of 10

    The Cosmos platform supplies open-source pre-trained world models and supporting tools for building fine-tunable digital world simulations to train Physical AI.

Pith tools