Pith. sign in

REVIEW 22 cited by

Segment and Track Anything

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06558 v1 pith:TZBTAI3X submitted 2023-05-11 cs.CV

classification cs.CV
keywords sam-tracksegmentanythinginteractionmodeltracktrackingframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report presents a framework called Segment And Track Anything (SAMTrack) that allows users to precisely and effectively segment and track any object in a video. Additionally, SAM-Track employs multimodal interaction methods that enable users to select multiple objects in videos for tracking, corresponding to their specific requirements. These interaction methods comprise click, stroke, and text, each possessing unique benefits and capable of being employed in combination. As a result, SAM-Track can be used across an array of fields, ranging from drone technology, autonomous driving, medical imaging, augmented reality, to biological analysis. SAM-Track amalgamates Segment Anything Model (SAM), an interactive key-frame segmentation model, with our proposed AOT-based tracking model (DeAOT), which secured 1st place in four tracks of the VOT 2022 challenge, to facilitate object tracking in video. In addition, SAM-Track incorporates Grounding-DINO, which enables the framework to support text-based interaction. We have demonstrated the remarkable capabilities of SAM-Track on DAVIS-2016 Val (92.0%), DAVIS-2017 Test (79.2%)and its practicability in diverse applications. The project page is available at: https://github.com/z-x-yang/Segment-and-Track-Anything.

Discussion (0). Sign in to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    LangDriveCTRL decomposes driving videos into 3D scene graphs and uses an agentic pipeline with specialized multi-modal agents to perform language-controlled object and behavior edits, achieving nearly 2x higher instru...

  2. Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement

    cs.CV 2024-11 unverdicted novelty 7.0 of 10

    VideoRepair detects text-video misalignments via MLLM-generated questions and performs localized, region-preserving refinement to improve alignment in existing T2V diffusion models.

  3. CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CROSS adds SAM-derived structural distillation and spatial-contrastive negatives to a SigLIP-SAM pipeline, achieving state-of-the-art cIoU on two remote sensing referring segmentation benchmarks.

  4. One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    CineMEC performs multimodal entity coreference by clustering visual entities and aligning them with text role mentions to boost captioning and grounding performance on an extended VidSitu dataset.

  5. AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    AdaTracker enables zero-shot cross-embodiment active visual tracking by encoding embodiment constraints from history to modulate a context-aware policy.

  6. When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    NUMINA improves counting accuracy in text-to-video diffusion models by up to 7.4% via a training-free identify-then-guide framework on the new CountBench dataset.

  7. Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A glove-based system learns human grasp demonstrations from joint angles and contact forces, then controls various robotic hands to grasp diverse objects without vision or retraining.

  8. Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A value-guided MPC policy trained on 2 million synthetic trajectories improves closed-loop 6-DoF grasping in clutter and adapts to object perturbations.

  9. VoCap: Video Object Captioning and Segmentation from Any Prompt

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.

  10. Grouped Speculative Decoding for Autoregressive Image Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.

  11. High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 3D Gaussian inpainting framework with automatic mask refinement and depth-initialized uncertainty weighting balances multi-view consistency and visual detail, reporting the best LPIPS on the SPIn-NeRF dataset.

  12. SAM 2: Segment Anything in Images and Videos

    cs.CV 2024-08 conditional novelty 6.0 of 10

    SAM 2 delivers more accurate video segmentation with 3x fewer user interactions and 6x faster image segmentation than the original SAM by training a streaming-memory transformer on the largest video segmentation datas...

  13. DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

    cs.CV 2023-08 unverdicted novelty 6.0 of 10

    DragNUWA integrates text, image, and trajectory controls into a diffusion video model using a Trajectory Sampler, Multiscale Fusion, and Adaptive Training to enable fine-grained open-domain video generation.

  14. Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    TEP-SAM adapts SAM for multiframe infrared small target detection by generating temporal-emerged cues from joint global-local motion modeling.

  15. ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis

    cs.CV 2026-03 accept novelty 5.0 of 10

    A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.

  16. Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation

    eess.IV 2025-10 unverdicted novelty 5.0 of 10

    RL4Seg3D applies reinforcement learning with novel reward functions and fusion to adapt echocardiography segmentation models across domains, improving accuracy, anatomical validity, and temporal consistency on over 30...

  17. ViSTR-GP: Online Cyberattack Detection via Vision-to-State Tensor Regression and Gaussian Processes in Automated Robotic Operations

    cs.RO 2025-09 conditional novelty 5.0 of 10

    ViSTR-GP uses an overhead camera, a learned vision-to-joint-angle map, and a Gaussian-process residual test to detect replay attacks on industrial robots from small physical deviations.

  18. DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.

  19. ViPE: Video Pose Engine for 3D Geometric Perception

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.

  20. Representative Volume Element: Existence and Extent in Cracked Heterogeneous Medium

    cs.CE 2025-08 unverdicted novelty 5.0 of 10

    Modified periodic boundary conditions that add strain periodicity to displacement periodicity are claimed to reduce mesh and size sensitivity in cracked-composite RVE simulations, tested on 1,200 samples.

  21. On Efficient Variants of Segment Anything Model: A Survey

    cs.CV 2024-10 unverdicted novelty 5.0 of 10

    A survey that reviews efficient variants of the Segment Anything Model, categorizes acceleration strategies, and provides a unified hardware evaluation on benchmarks.

  22. Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data

    cs.CV 2026-02 reject novelty 4.0 of 10

    Adding monocular depth to EfficientViT-SAM improves point-prompted segmentation at 3 and 5 clicks after fine-tuning on 11.2k images, but universal gains and data-efficiency are not established.

Pith tools