REVIEW 35 cited by
Track Anything: Segment Anything Meets Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recently, the Segment Anything Model (SAM) gains lots of attention rapidly due to its impressive segmentation performance on images. Regarding its strong ability on image segmentation and high interactivity with different prompts, we found that it performs poorly on consistent segmentation in videos. Therefore, in this report, we propose Track Anything Model (TAM), which achieves high-performance interactive tracking and segmentation in videos. To be detailed, given a video sequence, only with very little human participation, i.e., several clicks, people can track anything they are interested in, and get satisfactory results in one-pass inference. Without additional training, such an interactive design performs impressively on video object tracking and segmentation. All resources are available on {https://github.com/gaomingqi/Track-Anything}. We hope this work can facilitate related research.
Forward citations
Cited by 35 Pith papers
-
First-frame Supervised Video Polyp Segmentation via Propagative and Semantic Dual-teacher Network
A first-frame-only supervised polyp segmentation method, PSDNet, reaches within roughly 1.6 to 3.1 Dice points of fully supervised performance on SUN-SEG while using one annotated frame per video.
-
SimVS: Simulating World Inconsistencies for Robust View Synthesis
Video diffusion models simulate world inconsistencies, and a harmonization network trained on the simulated data reconciles sparse inconsistent multi-view images into consistent 3D scenes.
-
OGG-FR: Orthogonal Gradient Gaming and Frequency Rectification for Unmanned Aerial Vehicle Infrared Image Super-Resolution
OGG-FR is a plug-and-play training update that separates redundant and innovative parts of the FFT loss gradient and gates the innovative part by a confidence score, improving UAV infrared super-resolution in most tes...
-
EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking
A modular egocentric 3D tracking framework that lifts segmentation masks into 3D, scores motion with point trajectories, merges duplicate tracks, and improves PCL by 11 percent over the strongest baseline on ADT.
-
VoCap: Video Object Captioning and Segmentation from Any Prompt
VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.
-
Grouped Speculative Decoding for Autoregressive Image Generation
Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.
-
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
A new interactive method lets users click on an object in a video and generates audio for just that object, using mask-conditioned contrastive fine-tuning and latent diffusion.
-
ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs
A diffusion-enhanced 4D reconstruction pipeline for monocular video that achieves state-of-the-art results on DyCheck by supervising Gaussian splatting with personalized diffusion-generated pseudo-views.
-
R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision
The survey formalizes degradation-aware rendering for 3D Low-Level Vision and organizes roughly 100 methods on super-resolution, deblurring, weather removal, restoration, and enhancement in NeRF and 3DGS pipelines.
-
FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
FocalClick-XL, a three-subnet extension of FocalClick, achieves state-of-the-art click-based interactive segmentation and supports boxes, scribbles, and coarse masks through a single prompting layer.
-
SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost
SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.
-
CU-Multi: A Dataset for Multi-Robot Data Association
CU-Multi is a new public multi-robot dataset with controlled trajectory overlaps, dense semantic LiDAR labels, and geospatially aligned poses for evaluating data association in collaborative SLAM.
-
SnapNCode: An Integrated Development Environment for Programming Physical Objects Interactions
SnapNCode lets programmers write code with image tokens of real objects and attach executable snippets to objects so they run when a camera recognizes the object.
-
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
A prompting scheme that mines points, elastic boxes, and Gaussian-style masks from coarse masks lets SAM refine those masks more accurately than prior refinement tools.
-
SAM-guided Pseudo Label Enhancement for Multi-modal 3D Semantic Segmentation
SAM mask grouping plus majority-vote labeling and geometry-aware propagation densifies pseudo-labels and improves cross-domain 3D semantic segmentation accuracy.
-
Generative Physical AI in Vision: A Survey
A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.
-
FoundationStereo: Zero-Shot Stereo Matching
A new stereo depth model, trained on 1M synthetic pairs with adapted monocular features, reports strong zero-shot accuracy on multiple real-world benchmarks without target-domain fine-tuning.
-
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
A new collection of 100-plus Unreal Engine worlds and tools for embodied AI shows that environment diversity improves tracking agents, while exposing gaps in navigation, cross-embodiment transfer, and latency.
-
Efficient Track Anything
A lightweight video segmentation model with a vanilla ViT encoder and pooled memory cross-attention matches SAM 2 closely while running twice as fast and using 2.4x fewer parameters.
-
CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
CAT4D uses a multi-view video diffusion model to convert monocular video into multi-view video and reconstruct a dynamic 3D Gaussian scene, with competitive results on 4D reconstruction benchmarks.
-
Segment Anything in Light Fields for Real-Time Applications via Constrained Prompting
Segment Anything Model 2 is adapted to light fields by disparity-based mask propagation, semantic occlusion filtering, and reprompting, achieving view-consistent masks at real-time speed without retraining.
-
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.
-
3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction
A 3D Gaussian Splatting model whose Gaussian centers are represented as a learned combination of shared global motion bases recovers dynamic scenes and motion trajectories from monocular video.
-
Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
MA-SAM2 adds context-aware and occlusion-resilient memory to SAM2 and reports Challenge IoU of 62.49 percent on EndoVis2017 and 64.40 percent on EndoVis2018, beating SAM2 by 6.10 and 4.36 points.
-
CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning
A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.
-
UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References
UA-Pose estimates 6D poses from partial object references by labeling seen and unseen model regions, using that uncertainty to filter poses and trigger online 3D completion.
-
UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting
UAV4D reconstructs 4D scenes from monocular drone video by fitting a single global scale to align human meshes with the background mesh, then renders with separate Gaussian splats.
-
Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization
A weakly supervised localizer that builds a coarse class-aware map with a transformer, then picks the best SAM mask from a grid-prompted gallery by pixel-level overlap.
-
Dynamic Arthroscopic Navigation System for Anterior Cruciate Ligament Reconstruction Based on Multi-level Memory Architecture
A memory-based dynamic tracking system for arthroscopic ACL navigation runs at 25 FPS and reduces tracking error by roughly 20-58% relative to the authors' previous static system.
-
SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition
SplitGaussian reconstructs dynamic 3D scenes from monocular video by decomposing Gaussians into a rigid static branch and a deformable dynamic branch, claiming better motion separation and rendering quality than prior...
-
Continuous Marine Tracking via Autonomous UAV Handoff
A two-drone shark-tracking system using OSTrack and ORB feature matching is reported, but the headline handoff result comes from template matching in a simulated environment, not from a real inter-UAV flight.
-
Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model
A mixed-precision post-training quantization method for SAM that allocates bit-widths via an integer quadratic program guided by KL-divergence importance scores and a cross-layer synergy heuristic.
-
OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models
A multi-branch video-language model with person-specific ViTPose encoders and a frozen LLaMA 2 generates open-vocabulary descriptions of human interactions, and a merged 103-class benchmark is introduced.
-
There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks
SAM image encoders are poor at classification under linear probing, and lightweight in-context fine-tuning overfits to seen classes, so semantic information must be injected externally, for example via DINOv2.
-
A Survey on 3D Reconstruction Techniques in Plant Phenotyping: From Classical Methods to Neural Radiance Fields (NeRF), 3D Gaussian Splatting (3DGS), and Beyond
A review of 3D reconstruction techniques for plant phenotyping, comparing classical methods, NeRF, and 3D Gaussian Splatting on methodology, applications, and future directions.
Discussion (0). Continue with ORCID to comment.