REVIEW 22 cited by
Segment and Track Anything
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This report presents a framework called Segment And Track Anything (SAMTrack) that allows users to precisely and effectively segment and track any object in a video. Additionally, SAM-Track employs multimodal interaction methods that enable users to select multiple objects in videos for tracking, corresponding to their specific requirements. These interaction methods comprise click, stroke, and text, each possessing unique benefits and capable of being employed in combination. As a result, SAM-Track can be used across an array of fields, ranging from drone technology, autonomous driving, medical imaging, augmented reality, to biological analysis. SAM-Track amalgamates Segment Anything Model (SAM), an interactive key-frame segmentation model, with our proposed AOT-based tracking model (DeAOT), which secured 1st place in four tracks of the VOT 2022 challenge, to facilitate object tracking in video. In addition, SAM-Track incorporates Grounding-DINO, which enables the framework to support text-based interaction. We have demonstrated the remarkable capabilities of SAM-Track on DAVIS-2016 Val (92.0%), DAVIS-2017 Test (79.2%)and its practicability in diverse applications. The project page is available at: https://github.com/z-x-yang/Segment-and-Track-Anything.
Forward citations
Cited by 22 Pith papers
-
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
LangDriveCTRL decomposes driving videos into 3D scene graphs and uses an agentic pipeline with specialized multi-modal agents to perform language-controlled object and behavior edits, achieving nearly 2x higher instru...
-
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
VideoRepair detects text-video misalignments via MLLM-generated questions and performs localized, region-preserving refinement to improve alignment in existing T2V diffusion models.
-
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
CROSS adds SAM-derived structural distillation and spatial-contrastive negatives to a SigLIP-SAM pipeline, achieving state-of-the-art cIoU on two remote sensing referring segmentation benchmarks.
-
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
CineMEC performs multimodal entity coreference by clustering visual entities and aligning them with text role mentions to boost captioning and grounding performance on an extended VidSitu dataset.
-
AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking
AdaTracker enables zero-shot cross-embodiment active visual tracking by encoding embodiment constraints from history to modulate a context-aware policy.
-
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
NUMINA improves counting accuracy in text-to-video diffusion models by up to 7.4% via a training-free identify-then-guide framework on the new CountBench dataset.
-
Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration
A glove-based system learns human grasp demonstrations from joint angles and contact forces, then controls various robotic hands to grasp diverse objects without vision or retraining.
-
Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control
A value-guided MPC policy trained on 2 million synthetic trajectories improves closed-loop 6-DoF grasping in clutter and adapts to object perturbations.
-
VoCap: Video Object Captioning and Segmentation from Any Prompt
VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.
-
Grouped Speculative Decoding for Autoregressive Image Generation
Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.
-
High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details
A 3D Gaussian inpainting framework with automatic mask refinement and depth-initialized uncertainty weighting balances multi-view consistency and visual detail, reporting the best LPIPS on the SPIn-NeRF dataset.
-
SAM 2: Segment Anything in Images and Videos
SAM 2 delivers more accurate video segmentation with 3x fewer user interactions and 6x faster image segmentation than the original SAM by training a streaming-memory transformer on the largest video segmentation datas...
-
DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory
DragNUWA integrates text, image, and trajectory controls into a diffusion video model using a Trajectory Sampler, Multiscale Fusion, and Adaptive Training to enable fine-grained open-domain video generation.
-
Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection
TEP-SAM adapts SAM for multiframe infrared small target detection by generating temporal-emerged cues from joint global-local motion modeling.
-
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.
-
Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation
RL4Seg3D applies reinforcement learning with novel reward functions and fusion to adapt echocardiography segmentation models across domains, improving accuracy, anatomical validity, and temporal consistency on over 30...
-
ViSTR-GP: Online Cyberattack Detection via Vision-to-State Tensor Regression and Gaussian Processes in Automated Robotic Operations
ViSTR-GP uses an overhead camera, a learned vision-to-joint-angle map, and a Gaussian-process residual test to detect replay attacks on industrial robots from small physical deviations.
-
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.
-
ViPE: Video Pose Engine for 3D Geometric Perception
ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.
-
Representative Volume Element: Existence and Extent in Cracked Heterogeneous Medium
Modified periodic boundary conditions that add strain periodicity to displacement periodicity are claimed to reduce mesh and size sensitivity in cracked-composite RVE simulations, tested on 1,200 samples.
-
On Efficient Variants of Segment Anything Model: A Survey
A survey that reviews efficient variants of the Segment Anything Model, categorizes acceleration strategies, and provides a unified hardware evaluation on benchmarks.
-
Efficient Segment Anything with Depth-Aware Fusion and Limited Training Data
Adding monocular depth to EfficientViT-SAM improves point-prompted segmentation at 3 and 5 clicks after fine-tuning on 11.2k images, but universal gains and data-efficiency are not established.
Discussion (0). Sign in to comment.