Pith. sign in

REVIEW 14 cited by

CoTracker: It is Better to Track Together

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.07635 v3 pith:HJIIE7CP submitted 2023-07-14 cs.CV

CoTracker: It is Better to Track Together

classification cs.CV
keywords cotrackerpointstracktracksintroducejointlylongoccluded
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points independently, CoTracker tracks them jointly, accounting for their dependencies. We show that joint tracking significantly improves tracking accuracy and robustness, and allows CoTracker to track occluded points and points outside of the camera view. We also introduce several innovations for this class of trackers, including using token proxies that significantly improve memory efficiency and allow CoTracker to track 70k points jointly and simultaneously at inference on a single GPU. CoTracker is an online algorithm that operates causally on short windows. However, it is trained utilizing unrolled windows as a recurrent network, maintaining tracks for long periods of time even when points are occluded or leave the field of view. Quantitatively, CoTracker substantially outperforms prior trackers on standard point-tracking benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Functionalization via Structure Completion and Motion Rectification

    cs.CV 2026-05 unverdicted novelty 7.0

    Object functionalization is cast as neural graph completion over a functional graph of parts, contacts, and motions, followed by geometry realization that also rectifies erroneous motions, demonstrated on furniture wi...

  2. ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

    cs.RO 2024-09 conditional novelty 7.0

    ReKep encodes robotic tasks as optimizable Python functions over 3D keypoints that are generated automatically from language and RGB-D input, enabling real-time hierarchical planning on single- and dual-arm platforms ...

  3. Any-point Trajectory Modeling for Policy Learning

    cs.RO 2023-12 conditional novelty 7.0

    ATM pre-trains models to predict trajectories of any points in videos, then uses those predictions to learn strong visuomotor policies from minimal action labels, beating baselines by 80% on 130+ tasks.

  4. Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

    cs.RO 2026-06 unverdicted novelty 6.0

    Video2Sim2Real turns a single human video into a deployable robot manipulation skill by reconstructing a digital twin, anchoring motions to object-centric simulator configurations, and bridging sim-to-real gaps with i...

  5. Zero-shot World Models Are Developmentally Efficient Learners

    cs.AI 2026-04 unverdicted novelty 6.0

    A zero-shot visual world model trained on one child's experience achieves broad competence on physical understanding benchmarks while matching developmental behavioral patterns.

  6. 3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos

    cs.RO 2026-03 conditional novelty 6.0

    Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.

  7. TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

    cs.RO 2024-12 conditional novelty 6.0

    Visual trace prompting improves spatial-temporal awareness in VLA models, delivering 10% gains on SimplerEnv and 3.5x on real-robot tasks.

  8. RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos

    cs.CV 2024-12 unverdicted novelty 6.0

    RoDyGS separates static and dynamic elements in monocular videos using Gaussian splatting with regularization and introduces the Kubric-MRig benchmark for pose-free dynamic novel view synthesis.

  9. Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

    cs.RO 2024-09 unverdicted novelty 6.0

    Gen2Act enables generalizable robot manipulation for unseen objects and novel motions by using zero-shot human video generation from web data to condition a policy trained on an order of magnitude less robot interaction data.

  10. 3D Reconstruction with Spatial Memory

    cs.CV 2024-08 unverdicted novelty 6.0

    Spann3R uses a learned spatial memory to regress per-image pointmaps directly in a shared global coordinate system, removing the need for optimization-based alignment after per-pair predictions.

  11. PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

    cs.CV 2026-06 unverdicted novelty 5.0

    PhysisForcing applies trajectory and relational alignment losses to DiT features in video models, improving physical plausibility on R-Bench, PAI-Bench, and EZS-Bench while raising closed-loop robotic success rates fr...

  12. SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

    cs.CV 2026-05 conditional novelty 5.0

    SWoMo decouples symbolic rule-based motion modeling via scene graphs from visual realism via diffusion models, trained through inverse pairing of real cataract surgery videos reconstructed in the simulator for sim-to-...

  13. SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

    cs.CV 2026-05 unverdicted novelty 5.0

    SWoMo decouples symbolic rule-based motion modeling from diffusion-based visual realism using inverse pairing of reconstructed real videos to enable sim-to-real translation and generalization in cataract surgery simulations.

  14. World Action Models: A Survey

    cs.RO 2026-06 unverdicted novelty 3.0

    A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.