Pith. sign in

REVIEW 35 cited by

Track Anything: Segment Anything Meets Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11968 v2 pith:K2WUEM3B submitted 2023-04-24 cs.CV

classification cs.CV
keywords anythingsegmentationtrackvideosinteractivemodelperformssegment
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, the Segment Anything Model (SAM) gains lots of attention rapidly due to its impressive segmentation performance on images. Regarding its strong ability on image segmentation and high interactivity with different prompts, we found that it performs poorly on consistent segmentation in videos. Therefore, in this report, we propose Track Anything Model (TAM), which achieves high-performance interactive tracking and segmentation in videos. To be detailed, given a video sequence, only with very little human participation, i.e., several clicks, people can track anything they are interested in, and get satisfactory results in one-pass inference. Without additional training, such an interactive design performs impressively on video object tracking and segmentation. All resources are available on {https://github.com/gaomingqi/Track-Anything}. We hope this work can facilitate related research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 95 citations worldwide. Full citation record

  1. First-frame Supervised Video Polyp Segmentation via Propagative and Semantic Dual-teacher Network

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A first-frame-only supervised polyp segmentation method, PSDNet, reaches within roughly 1.6 to 3.1 Dice points of fully supervised performance on SUN-SEG while using one annotated frame per video.

  2. SimVS: Simulating World Inconsistencies for Robust View Synthesis

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Video diffusion models simulate world inconsistencies, and a harmonization network trained on the simulated data reconciles sparse inconsistent multi-view images into consistent 3D scenes.

  3. OGG-FR: Orthogonal Gradient Gaming and Frequency Rectification for Unmanned Aerial Vehicle Infrared Image Super-Resolution

    cs.CV 2026-08 conditional novelty 6.0 of 10

    OGG-FR is a plug-and-play training update that separates redundant and innovative parts of the FFT loss gradient and gates the innovative part by a confidence score, improving UAV infrared super-resolution in most tes...

  4. EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A modular egocentric 3D tracking framework that lifts segmentation masks into 3D, scores motion with point trajectories, merges duplicate tracks, and improves PCL by 11 percent over the strongest baseline on ADT.

  5. VoCap: Video Object Captioning and Segmentation from Any Prompt

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.

  6. Grouped Speculative Decoding for Autoregressive Image Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Accepting clusters of visually valid tokens during speculative decoding yields about 3.7x training-free speedup for autoregressive image generation with quality preserved.

  7. Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new interactive method lets users click on an object in a video and generates audio for just that object, using mask-conditioned contrastive fine-tuning and latent diffusion.

  8. ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion-enhanced 4D reconstruction pipeline for monocular video that achieves state-of-the-art results on DyCheck by supervising Gaussian splatting with personalized diffusion-generated pseudo-views.

  9. R3eVision: A Survey on Robust Rendering, Restoration, and Enhancement for 3D Low-Level Vision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The survey formalizes degradation-aware rendering for 3D Low-Level Vision and organizes roughly 100 methods on super-resolution, deblurring, weather removal, restoration, and enhancement in NeRF and 3DGS pipelines.

  10. FocalClick-XL: Towards Unified and High-quality Interactive Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FocalClick-XL, a three-subnet extension of FocalClick, achieves state-of-the-art click-based interactive segmentation and supports boxes, scribbles, and coarse masks through a single prompting layer.

  11. SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SAM-I2V upgrades SAM to video segmentation with three lightweight modules (temporal integrator, selective memory, memory prompts), reaching about 90% of SAM 2.1's average J&F at 0.2% of its training cost.

  12. CU-Multi: A Dataset for Multi-Robot Data Association

    cs.RO 2025-05 conditional novelty 6.0 of 10

    CU-Multi is a new public multi-robot dataset with controlled trajectory overlaps, dense semantic LiDAR labels, and geospatially aligned poses for evaluating data association in collaborative SLAM.

  13. SnapNCode: An Integrated Development Environment for Programming Physical Objects Interactions

    cs.HC 2025-05 conditional novelty 6.0 of 10

    SnapNCode lets programmers write code with image tokens of real objects and attach executable snippets to objects so they run when a camera recognizes the object.

  14. SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A prompting scheme that mines points, elastic boxes, and Gaussian-style masks from coarse masks lets SAM refine those masks more accurately than prior refinement tools.

  15. SAM-guided Pseudo Label Enhancement for Multi-modal 3D Semantic Segmentation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    SAM mask grouping plus majority-vote labeling and geometry-aware propagation densifies pseudo-labels and improves cross-domain 3D semantic segmentation accuracy.

  16. Generative Physical AI in Vision: A Survey

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.

  17. FoundationStereo: Zero-Shot Stereo Matching

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new stereo depth model, trained on 1M synthetic pairs with adapted monocular features, reports strong zero-shot accuracy on multiple real-world benchmarks without target-domain fine-tuning.

  18. UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A new collection of 100-plus Unreal Engine worlds and tools for embodied AI shows that environment diversity improves tracking agents, while exposing gaps in navigation, cross-embodiment transfer, and latency.

  19. Efficient Track Anything

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A lightweight video segmentation model with a vanilla ViT encoder and pooled memory cross-attention matches SAM 2 closely while running twice as fast and using 2.4x fewer parameters.

  20. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    CAT4D uses a multi-view video diffusion model to convert monocular video into multi-view video and reconstruct a dynamic 3D Gaussian scene, with competitive results on 4D reconstruction benchmarks.

  21. Segment Anything in Light Fields for Real-Time Applications via Constrained Prompting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Segment Anything Model 2 is adapted to light fields by disparity-based mask propagation, semantic occlusion filtering, and reprompting, achieving view-consistent masks at real-time speed without retraining.

  22. ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis

    cs.CV 2026-03 accept novelty 5.0 of 10

    A lightweight point decoder plus task-prompt joint training yields ~3× faster SAM-based hierarchical text detection with competitive HierText results and +11% average F-score on three single-level benchmarks.

  23. 3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A 3D Gaussian Splatting model whose Gaussian centers are represented as a learned combination of shared global motion bases recovers dynamic scenes and motion trajectories from monocular video.

  24. Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MA-SAM2 adds context-aware and occlusion-resilient memory to SAM2 and reports Challenge IoU of 62.49 percent on EndoVis2017 and 64.40 percent on EndoVis2018, beating SAM2 by 6.10 and 4.36 points.

  25. CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning

    eess.IV 2025-07 conditional novelty 5.0 of 10

    A CLIP-based encoder with RL residual refinement and curriculum learning reaches 81% mIoU on EndoVis 2018 and 74.12% on EndoVis 2017 surgical segmentation.

  26. UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References

    cs.CV 2025-06 conditional novelty 5.0 of 10

    UA-Pose estimates 6D poses from partial object references by labeling seen and unseen model regions, using that uncertainty to filter poses and trigger online 3D completion.

  27. UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    UAV4D reconstructs 4D scenes from monocular drone video by fitting a single global scale to align human meshes with the background mesh, then renders with separate Gaussian splats.

  28. Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A weakly supervised localizer that builds a coarse class-aware map with a transformer, then picks the best SAM mask from a grid-prompted gallery by pixel-level overlap.

  29. Dynamic Arthroscopic Navigation System for Anterior Cruciate Ligament Reconstruction Based on Multi-level Memory Architecture

    cs.CV 2025-04 conditional novelty 5.0 of 10

    A memory-based dynamic tracking system for arthroscopic ACL navigation runs at 25 FPS and reduces tracking error by roughly 20-58% relative to the authors' previous static system.

  30. SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    SplitGaussian reconstructs dynamic 3D scenes from monocular video by decomposing Gaussians into a rigid static branch and a deformable dynamic branch, claiming better motion separation and rendering quality than prior...

  31. Continuous Marine Tracking via Autonomous UAV Handoff

    cs.CV 2025-07 reject novelty 4.0 of 10

    A two-drone shark-tracking system using OSTrack and ORB feature matching is reported, but the headline handoff result comes from template matching in a simulated environment, not from a real inter-UAV flight.

  32. Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A mixed-precision post-training quantization method for SAM that allocates bit-widths via an integer quadratic program guided by KL-divergence importance scores and a cross-layer synergy heuristic.

  33. OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models

    cs.CV 2024-12 reject novelty 4.0 of 10

    A multi-branch video-language model with person-specific ViTPose encoders and a frozen LLaMA 2 generates open-vocabulary descriptions of human interactions, and a merged 103-class benchmark is introduced.

  34. There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks

    cs.CV 2024-11 conditional novelty 4.0 of 10

    SAM image encoders are poor at classification under linear probing, and lightweight in-context fine-tuning overfits to seen classes, so semantic information must be injected externally, for example via DINOv2.

  35. A Survey on 3D Reconstruction Techniques in Plant Phenotyping: From Classical Methods to Neural Radiance Fields (NeRF), 3D Gaussian Splatting (3DGS), and Beyond

    eess.IV 2025-04 conditional novelty 1.0 of 10

    A review of 3D reconstruction techniques for plant phenotyping, comparing classical methods, NeRF, and 3D Gaussian Splatting on methodology, applications, and future directions.

Pith tools