REVIEW 4 major objections 5 minor 10 cited by
SpatialTrackerV2: 3D Point Tracking Made Easy
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A single feed-forward model now sets the 3D tracking record on TAPVid-3D, matching a slow optimizer at 50x its speed.
desk verdict A credible SOTA for 3D point tracking with a genuinely new dual-branch architecture and an impressive joint-training recipe, but the headline egocentric gain rests on an unvalidated MoGe teacher and the paper needs missing ablations and code before I'd fully trust the numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
SyncFormer, an iterative transformer with separate 2D (image-space) and 3D (camera-coordinate) branches whose representations are exchanged through cross-attention between compact proxy tokens. The 2D branch updates tracks in UV space while the 3D branch updates positions in camera coordinates using correlations computed on normalized point maps. A weighted Procrustes alignment, weighted by the predicted dynamic scores, registers the 3D trajectories to the world frame, and a bundle-adjustment step optimizes camera poses from the filtered static points. The front end combines a temporal encoder (built by extending a monocular depth architecture with alternating intra- and inter-frame attention and learnable pose/scale tokens) that outputs depth and initial camera poses. The design exists to keep 2D tracking, 3D geometry, and ego-motion consistent without letting any one task's errors corrupt the others.
What would settle it
Retrain the pipeline with teacher depth supervision removed from the pose-only and unlabeled datasets and re-measure 3D tracking and video-depth accuracy; a large drop would confirm the teacher is load-bearing. Separately, compare the predicted depth and 3D tracks against LiDAR ground truth on egocentric sequences to test whether teacher bias in that domain propagates into the final geometry.
Extended reading notes
Core claim
SpatialTrackerV2 is a feed-forward 3D point tracker that unifies point tracking, monocular depth estimation, and camera pose estimation into a single differentiable network. It decomposes world-space 3D motion into video depth, camera ego-motion, and per-pixel object motion; a front end encodes temporal video context to produce scale-aligned depth and an initial camera trajectory, and a back-end module iteratively refines 2D and 3D trajectories together with visibility and dynamic/static scores. The refinements feed back into a differentiable bundle-adjustment step that re-estimates camera poses, and the entire loop is trained jointly on 17 datasets spanning full supervision (RGB-D with 3D track labels) to pose-only and unlabeled footage. The paper reports state-of-the-art results on a standard 3D tracking benchmark (21.2 average Jaccard, 31.0 3D position error, 90.6 occlusion accuracy), video depth that improves on both a strong feed-forward model and an optimization-based system (0.081 absolute relative error), and camera poses on par with the optimization-based system at roughly 50x lower inference time.
Load-bearing premise
The scaling result rests on a teacher monocular depth model being right on pose-only and unlabeled videos, since the paper never measures that teacher's accuracy on egocentric or Internet footage and its errors would be baked into the learned geometry and 3D trajectories.
Editorial extensions
If this is right
- Per-video optimization pipelines can be replaced by a single forward pass for many dynamic-reconstruction tasks, shrinking inference from minutes to seconds.
- Because the pipeline trains on pose-only and unlabeled video, depth and 3D tracking should keep improving as more in-the-wild data is added, without needing expensive 3D track annotations.
- The decoupled 2D/3D branches mean that adding 3D supervision no longer degrades 2D tracking accuracy: the paper's naive-lifting baseline drops average Jaccard from 64.4 to 51.6, while SyncFormer keeps it at 64.9.
- The explicit camera-motion decomposition yields the largest gains on egocentric and background-point-heavy subsets, so the method should be especially useful for wearable and robotic video.
Reading between the lines
- The dynamic-score weighting in the Procrustes alignment is a self-supervised bootstrap; an ablation comparing against an oracle static/dynamic segmentation would quantify how much error the circularity introduces.
- If the teacher depth model is biased on egocentric or Internet footage, the paper's scaling story predicts that a better teacher or a second, independent geometric prior would directly improve 3D tracking on those domains, which is a testable prediction.
- The same decomposition (depth plus ego-motion plus object motion) could serve as a representation for downstream tasks like video editing, generation control, and robotic grasping, potentially letting those tasks consume trajectory and geometry directly rather than training their own motion estimators.
- The 50x speed advantage suggests that feed-forward geometry-and-motion models could enable real-time structure-from-motion-like applications on mobile devices, which the paper does not discuss.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SpatialTrackerV2, a fully feed-forward model for monocular 3D point tracking that jointly estimates video depth, camera poses, and long-term 2D/3D point trajectories. The architecture combines a temporal encoder front-end for scale-aligned depth and camera initialization with a SyncFormer back-end that alternates between 2D and 3D correlation branches and performs differentiable bundle adjustment, using learned dynamic and visibility scores to weight the optimization. Training mixes posed RGB-D data with tracking annotations, posed RGB-D without tracking, and pose-only or unlabeled video where a monocular teacher provides relative depth. The paper reports state-of-the-art results on TAPVid-3D (21.2 AJ, 31.0 APD3D), improvements over prior 3D trackers such as DELTA, and faster feed-forward inference than optimization-based reconstruction systems like MegaSAM. The main empirical tables support the headline tracking claims, while the depth and pose comparisons are more mixed than the abstract suggests.
Significance. If the benchmark results hold, the paper offers a meaningful advance: a single model that jointly produces metric-scale video depth, camera poses, and 3D point trajectories, with a unified training scheme across heterogeneous datasets. The evaluation protocol is largely fair, with comparisons controlled for the depth and pose provider (e.g., Ours-offl+ versus Type I baselines), and the paper reports results across three benchmark subsets. The biggest reported gain, 21.2 versus 13.1 AJ over DELTA, is large and directionally consistent across Aria, DriveTrack, and PStudio. The paper also provides a public demo and includes useful implementation details such as dataset categories and training stages. The main caveats are that the 'matches MegaSAM' claim is not uniformly supported by Tables 2 and 3, and that several load-bearing components (the MoGe teacher, the dynamic-score bootstrap, and the missing ablation rows) are not yet validated or documented. The submitted version is a solid empirical contribution but needs revision before the claims can be taken at face value.
major comments (4)
- [Abstract; Section 4.2, Tables 2 and 3] The abstract's statement that the method 'matches the accuracy of leading dynamic 3D reconstruction approaches' is not supported by the paper's own tables. In Table 3, on Sintel, Ours (ATE 0.054, RPEt 0.027, RPEr 0.288) is substantially worse than MegaSAM (ATE 0.023, RPEt 0.008, RPEr 0.060), and on Lightspeed, Ours has ATE 0.134 versus MegaSAM's 0.105. In Table 2, on Sintel, MegaSAM is also better on AbsRel (0.185 vs. 0.199) and δ1.25 (0.746 vs. 0.703). The average depth numbers favor the proposed method, but the 'matches' claim should be qualified to the datasets and metrics where it actually holds, or the text should explain why the Sintel/Lightspeed gaps are acceptable.
- [Section 4.3, Tables 2 and 5] The ablation analysis is incomplete as presented. In the 3D Point Tracking ablation, the text states that 'our final version is clearly better than these two' but Table 5 contains only Base@K and Base@V,K,P, with no row for the final model. In the Depth Estimation ablation, the text says 'As shown in Tab. 2' for rows Ours-Synthetic and Ours-Real-Full, but Table 2 does not contain these rows. These missing rows make the scaling and joint-training claims in Section 4.3 unverifiable. The authors should either add the missing rows or remove the references.
- [Section 3.3, training data category (3)] The paper's scalable-training story depends on the MoGe teacher [78] for preserving relative depth on HOI4D, Ego4D, and Stereo4D, which are the main egocentric and Internet-video training sources. The TAPVid-3D Aria subset, where the largest gains are reported (24.6 vs. 23.5 AJ for the next best method), is egocentric, yet the paper provides no evaluation of MoGe's depth accuracy on these domains and no ablation that removes or replaces the teacher. Since a systematic teacher bias would be distilled into the video depth and then into the 3D tracks via Eq. (2), I ask the authors to either include a teacher-quality analysis (e.g., relative depth consistency on a held-out egocentric split) or an ablation with an alternative teacher such as DepthAnythingV2 or with the teacher removed.
- [Section 3.2, camera motion optimization] The dynamic-score bootstrap is a potential source of bias. In the weighted Procrustes alignment, the learned dynamic probability pdyn decides which points contribute to the camera pose optimization, and the updated poses then feed back into the dynamic and visibility scores in the next SyncFormer iteration. On category (3) data there are no ground-truth dynamic masks, so the loop is entirely self-supervised; if dynamic points are misclassified, the biased poses could reinforce the error in the scores. The paper should at least analyze this failure mode (e.g., by comparing the learned pdyn against ground-truth masks on category (1) data and by ablating the dynamic filtering) or explicitly justify why the self-consistency signal is sufficient.
minor comments (5)
- [Abstract] The phrase 'outperforms existing 3D tracking methods by 30%' is ambiguous: relative to DELTA the gain is 61.8% in AJ, while relative to TAPIP3D it is 12.8%; please specify the reference point and metric.
- [Table 1] The 'Full-ours' row lacks Type and Depth/Cam Pose entries, making its protocol (world-space, own depth and pose) implicit; please add them for clarity.
- [Table 2] The citation for MegaSAM is given as [74] but should be [42].
- [Section 4.1, Table 4] The baseline name is written inconsistently as 'Cotracker3' in the text and 'CoTracker3' elsewhere; please standardize the spelling.
- [General] The paper does not mention whether code and trained models will be released; given the 17-dataset training recipe and the dependence on the MoGe teacher, this is important for reproducibility.
Circularity Check
No significant circularity: the benchmark claims are externally evaluated and the training pipeline uses heterogeneous external supervision; no equation reduces to a fitted parameter.
full rationale
SpatialTrackerV2's central claims are empirical numbers on held-out benchmarks: TAPVid-3D for 3D tracking, KITTI/TUM/Bonn/Sintel for video depth, and TUM-dynamic/Lightspeed/Sintel for camera poses. These datasets provide ground truth independent of the model's fitted parameters. The training recipe uses external GT tracks, depth, and poses on categories (1)-(2), and an external MoGe teacher for relative depth on category (3); the paper's ablations, e.g., Base@K vs. Base@V,K,P and joint-training comparisons, isolate the contribution of data scaling and joint optimization. The dynamic-score-weighted Procrustes loop in Section 3.2 is a bootstrapped training objective that uses the model's own predictions as alignment weights; it is not a reported prediction against ground truth, and the evaluated AJ/APD3D/OA numbers come from TAPVid-3D annotations, so no evaluated result is forced by construction. The self-citations to VGGT, CoTracker3, SpatialTracker, and Video Depth Anything are architectural baselines, initializations, or comparative references; they are externally benchmarked and are not invoked as a uniqueness theorem or as the sole justification for the state-of-the-art claim. The unvalidated MoGe teacher on egocentric/Internet domains and the absence of a no-MoGe ablation are legitimate empirical robustness concerns, but they are not definitional circularity.
Assumptions & free parameters
free parameters (4)
- Correlation window radius =
3 (Delta = 3)
- Number of correlation scales =
4 (s = 1,2,3,4)
- SyncFormer iteration count =
unspecified
- Joint training loss weights =
unspecified
assumptions (4)
- domain assumption A monocular video can be decomposed into static scene geometry, camera ego-motion, and pixel-wise object motion.
- domain assumption Monocular depth is recoverable up to scale and shift, and the learned scale-shift parameters a,b sufficiently align it with camera poses.
- domain assumption The teacher depth model MoGe provides sufficiently accurate relative depth for unlabeled and pose-only training data.
- ad hoc to paper Dynamic points can be identified by a learned probability pdyn and filtering them before Procrustes alignment improves pose estimation.
Cite this review
Pith. "Pith review of SpatialTrackerV2: 3D Point Tracking Made Easy." pith.science (2026). https://pith.science/paper/VLBOM5IW
@misc{pith2026250712462,
author = {Pith},
title = {Pith review of: SpatialTrackerV2: 3D Point Tracking Made Easy},
year = {2026},
howpublished = {\url{https://pith.science/paper/VLBOM5IW}},
note = {Machine review of arXiv:2507.12462}
}
abstract
We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50$\times$ faster.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 10 Pith papers
-
EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
World Action Model co-training with DINO or 3D-flow targets scales human-to-robot transfer on bimanual tasks far better than behavior cloning, while pixel prediction transfers weakly.
-
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
A single video-diffusion framework composites both static images and dynamic footage along user-defined trajectories by transporting canonical foreground latents directly into the background latent sequence.
-
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
A training-free two-pass adaptation of VGGT, with attention-based motion masking and inverse-variance depth fusion, improves dynamic-scene point-cloud reconstruction on DyCheck.
-
SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
SymphoMotion jointly controls camera trajectories and depth-aware object dynamics inside one video diffusion model, supported by the new RealCOD-25K real-world paired-motion dataset.
-
4DGS360: 360{\deg} Gaussian Reconstruction of Dynamic Objects from a Single Video
Combining high-confidence 2D tracking anchors with a 3D point tracker improves initialization and monocular 360-degree dynamic object reconstruction, demonstrated on a new far-viewpoint benchmark.
-
IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
A causal streaming transformer that jointly predicts camera motion, 3D geometry, and persistent object-instance features from video, trained on a new 147K-sequence 4D dataset.
-
TrackDeform3D: Markerless and Autonomous 3D Keypoint Tracking and Dataset Collection for Deformable Objects
A markerless RGB-D tracking pipeline for deformable objects plus a released 110-minute, six-object trajectory dataset.
-
Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation
A trajectory-conditioned retrieval system discovers multiple motion descriptions in videos without user queries and grounds them to point tracks, evaluated mainly on MeViS.
-
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.
-
Reconstructing 4D Spatial Intelligence: A Survey
A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.
Reference graph
Works this paper leans on
-
[78]
Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision, 2024. 5
work page 2024
-
[1]
L4p: Low-level 4d vision perception unified
Abhishek Badki, Hang Su, Bowen Wen, and Orazio Gallo. L4p: Low-level 4d vision perception unified. arXiv preprint arXiv:2502.13078, 2025. 2
arXiv 2025
-
[2]
Speeded-Up Robust Features (SURF)
Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool. Speeded-Up Robust Features (SURF). CVIU, 110(3),
-
[3]
Zoedepth: Zero-shot trans- fer by combining relative and metric depth
Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias M ¨uller. Zoedepth: Zero-shot trans- fer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023. 3
arXiv 2023
-
[4]
Weikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yijin Li, Fu- Yun Wang, and Hongsheng Li. Gs-dit: Advancing video generation with pseudo 4d gaussian fields through efficient dense 3d point tracking. CoRR, abs/2501.02690, 2025. 2
arXiv 2025
-
[5]
Midas v3.1 - A model zoo for robust monocular relative depth estimation
Reiner Birkl, Diana Wofk, and Matthias M¨uller. Midas v3.1 - A model zoo for robust monocular relative depth estimation. CoRR, abs/2307.14460, 2023. 3
arXiv 2023
-
[6]
Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang
Michael J. Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang. BEDLAM: A synthetic dataset of bodies exhibit- ing detailed lifelike animated motion. InCVPR, pages 8726–
-
[7]
Vir- tual kitti 2, 2020
Yohann Cabon, Naila Murray, and Martin Humenberger. Vir- tual kitti 2, 2020. 5, 9
2020
Show all 97 references
-
[8]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 3
2021
-
[9]
Video depth anything: Consistent depth estimation for super-long videos
Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zi- long Huang, Jiashi Feng, and Bingyi Kang. Video depth anything: Consistent depth estimation for super-long videos. CoRR, abs/2501.12375, 2025. 3, 5, 6, 7
2025 arXiv
-
[10]
Pref3r: Pose- free feed-forward 3d gaussian splatting from variable-length image sequence
Zequn Chen, Jiezhi Yang, and Heng Yang. Pref3r: Pose- free feed-forward 3d gaussian splatting from variable-length image sequence. CoRR, abs/2411.16877, 2024. 2
2024 arXiv
-
[11]
Local all-pair correspon- dence for point tracking
Seokju Cho, Jiahui Huang, Jisu Nam, Honggyu An, Seun- gryong Kim, and Joon-Young Lee. Local all-pair correspon- dence for point tracking. In European Conference on Com- puter Vision, pages 306–325. Springer, 2024. 2
2024
-
[12]
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017. 5, 8
2017
-
[13]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Informa- tion Processing Systems, 34:8780–8794, 2021. 3
2021
-
[14]
Tap-vid: A benchmark for track- ing any point in a video
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adria Re- casens, Lucas Smaira, Yusuf Aytar, Joao Carreira, Andrew Zisserman, and Yi Yang. Tap-vid: A benchmark for track- ing any point in a video. Advances in Neural Information Processing Systems, 35:13610–13626, 2022. 2
2022
-
[15]
Tapir: Tracking any point with per-frame initialization and 9 temporal refinement
Carl Doersch, Yi Yang, Mel Vecerik, Dilara Gokay, Ankush Gupta, Yusuf Aytar, Joao Carreira, and Andrew Zisserman. Tapir: Tracking any point with per-frame initialization and 9 temporal refinement. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , p...
2023
-
[16]
Bootstap: Bootstrapped training for tracking-any-point
Carl Doersch, Pauline Luc, Yi Yang, Dilara Gokay, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ignacio Rocco, Ross Goroshin, Jo˜ao Carreira, et al. Bootstap: Bootstrapped training for tracking-any-point. In Proceedings of the Asian Conference on Computer Vision, pages 3257–32...
2024
-
[17]
Depth map prediction from a single image using a multi-scale deep net- work
David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep net- work. In NeurIPS, pages 2366–2374, 2014. 3
2014
-
[18]
Deep ordinal regression net- work for monocular depth estimation
Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Bat- manghelich, and Dacheng Tao. Deep ordinal regression net- work for monocular depth estimation. InCVPR, pages 2002–
2002
-
[19]
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. The in- ternational journal of robotics research, 32(11):1231–1237,
-
[20]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vi- sio...
2022
-
[21]
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapra- gasam, Florian Golemo, Charles Herrmann, et al. Kubric: A scalable dataset generator. In Proceedings of the IEEE/CVF conference on computer vision and pattern re...
2022
-
[22]
Diffusion as shader: 3d- aware video diffusion for versatile video generation control
Zekai Gu, Rui Yan, Jiahao Lu, Peng Li, Zhiyang Dou, Chenyang Si, Zhen Dong, Qifeng Liu, Cheng Lin, Ziwei Liu, Wenping Wang, and Yuan Liu. Diffusion as shader: 3d- aware video diffusion for versatile video generation control. CoRR, abs/2501.03847, 2025. 2
2025 arXiv
-
[23]
Harley, Zhaoyuan Fang, and Katerina Fragkiadaki
Adam W. Harley, Zhaoyuan Fang, and Katerina Fragkiadaki. Particle video revisited: Tracking through occlusions using point trajectories. In ECCV, 2022. 2
2022
-
[24]
Multiple View Ge- ometry in Computer Vision
Richard Hartley and Andrew Zisserman. Multiple View Ge- ometry in Computer Vision . Cambridge University Press, ISBN: 0521540518, second edition, 2004. 3
2004
-
[25]
In defense of the eight-point algorithm
Richard I Hartley. In defense of the eight-point algorithm. IEEE Transactions on pattern analysis and machine intelli- gence, 19(6):580–593, 1997. 3
1997
-
[26]
Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface nor- mal estimation
Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface nor- mal estimation. IEEE TPAMI, 46(12):10579–10596, 2024. 3
2024
-
[27]
Depthcrafter: Generating consistent long depth sequences for open-world videos
Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xiaodong Cun, Yong Zhang, Long Quan, and Ying Shan. Depthcrafter: Generating consistent long depth sequences for open-world videos. CoRR, abs/2409.02095, 2024. 2, 3, 6, 7
2024 arXiv
-
[28]
Deepmvs: Learning multi-view stereopsis
Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang. Deepmvs: Learning multi-view stereopsis. In CVPR, pages 2821–2830. Computer Vision Foundation / IEEE Computer Society, 2018. 5
2018
-
[29]
Mitra, and Duygu Ceylan
Hyeonho Jeong, Chun-Hao Paul Huang, Jong Chul Ye, Niloy J. Mitra, and Duygu Ceylan. Track4gen: Teaching video diffusion models to track points improves video gen- eration. CoRR, abs/2412.06016, 2024. 2
2024 arXiv
-
[30]
Stereo4d: Learning how things move in 3d from internet stereo videos.arXiv preprint,
Linyi Jin, Richard Tucker, Zhengqi Li, David Fouhey, Noah Snavely, and Aleksander Holynski. Stereo4d: Learning how things move in 3d from internet stereo videos.arXiv preprint,
-
[31]
Dy- namicstereo: Consistent dynamic depth from stereo videos
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Dy- namicstereo: Consistent dynamic depth from stereo videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13229–13239, 2023. 5
2023
-
[32]
Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos
Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos. arXiv preprint arXiv:2410.11831 ,
-
[33]
Co- tracker: It is better to track together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker: It is better to track together. InEuropean Conference on Computer Vision, pages 18–35. Springer, 2024. 2, 3, 4, 8
2024
-
[34]
Fast encoder- based 3d from casual videos via point track processing.arXiv preprint arXiv:2404.07097, 2024
Yoni Kasten, Wuyue Lu, and Haggai Maron. Fast encoder- based 3d from casual videos via point track processing.arXiv preprint arXiv:2404.07097, 2024. 2
2024 arXiv
-
[35]
Ssd-6d: Making rgb-based 3d detec- tion and 6d pose estimation great again
Wadim Kehl, Fabian Manhardt, Federico Tombari, Slobodan Ilic, and Nassir Navab. Ssd-6d: Making rgb-based 3d detec- tion and 6d pose estimation great again. In Proceedings of the IEEE international conference on computer vision, pages 1521–1529, 2017. 3
2017
-
[36]
Robust consistent video depth estimation
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Robust consistent video depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, vir- tual, June 19-25, 2021, pages 1611–1621. Computer Vision Foundation / IEEE, 2021. 8
2021
-
[37]
Tapvid-3d: A benchmark for tracking any point in 3d
Skanda Koppula, Ignacio Rocco, Yi Yang, Joseph Heyward, Jo˜ao Carreira, Andrew Zisserman, Gabriel Brostow, and Carl Doersch. Tapvid-3d: A benchmark for tracking any point in 3d. In NeurIPS, 2024. 2, 5
2024
-
[38]
Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds
Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas, and Kostas Daniilidis. Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds. arXiv preprint arXiv:2405.17421, 2024. 2
2024 arXiv
-
[39]
Five-point motion esti- mation made easy
Hongdong Li and Richard Hartley. Five-point motion esti- mation made easy. In 18th International Conference on Pat- tern Recognition (ICPR’06), pages 630–633. IEEE, 2006. 3
2006
-
[40]
Taptr: Tracking any point with transformers as detection
Hongyang Li, Hao Zhang, Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, and Lei Zhang. Taptr: Tracking any point with transformers as detection. In European Confer- ence on Computer Vision, pages 57–75. Springer, 2024. 2
2024
-
[41]
Drip: Unleashing diffusion priors for joint foreground and alpha prediction in image matting.Advances in Neural Information Processing Systems, 37:79868–79888, 2024
Xiaodi Li, Zongxin Yang, Ruijie Quan, and Yi Yang. Drip: Unleashing diffusion priors for joint foreground and alpha prediction in image matting.Advances in Neural Information Processing Systems, 37:79868–79888, 2024. 3 10
2024
-
[42]
Megasam: Accurate, fast and robust structure and motion from casual dynamic videos
Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang, Linyi Jin, Vickie Ye, Angjoo Kanazawa, Aleksander Holyn- ski, and Noah Snavely. Megasam: Accurate, fast and robust structure and motion from casual dynamic videos. In CVPR,
-
[43]
Relpose++: Recovering 6d poses from sparse-view observations
Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tul- siani. Relpose++: Recovering 6d poses from sparse-view observations. arXiv preprint arXiv:2305.04926, 2023. 3
2023 arXiv
-
[44]
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , p...
2024
-
[45]
Sift flow: Dense correspondence across scenes and its applications
Ce Liu, Jenny Yuen, and Antonio Torralba. Sift flow: Dense correspondence across scenes and its applications. IEEE transactions on pattern analysis and machine intelligence , 33(5):978–994, 2010. 2
2010
-
[46]
HOI4D: A 4d egocentric dataset for category-level human- object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4d egocentric dataset for category-level human- object interaction. In CVPR, pages 20981–20990. IEEE,
-
[47]
Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, 2015. 2
2015
-
[48]
Robo-gs: A physics consistent spatial-temporal model for robotic arm with hybrid representation
Haozhe Lou, Yurong Liu, Yike Pan, Yiran Geng, Jianteng Chen, Wenlong Ma, Chenglong Li, Lin Wang, Hengzhen Feng, Lu Shi, Liyi Luo, and Yongliang Shi. Robo-gs: A physics consistent spatial-temporal model for robotic arm with hybrid representation. CoRR, abs/2408.14873, 2024. 2
2024 arXiv
-
[49]
David G. Lowe. Object Recognition from Local Scale- Invariant Features. In Proc. ICCV, 1999. 3
1999
-
[50]
David G. Lowe. Distinctive Image Features from Scale- Invariant Keypoints. IJCV, 60(2), 2004. 3
2004
-
[51]
Virtual correspondence: Hu- mans as a cue for extreme-view geometry
Wei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun, and Antonio Torralba. Virtual correspondence: Hu- mans as a cue for extreme-view geometry. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15924–15934, 2022. 3
2022
-
[52]
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and ...
2016
-
[53]
Spring: A high-resolution high- detail dataset and benchmark for scene flow, optical flow and stereo
Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nali- vayko, and Andr ´es Bruhn. Spring: A high-resolution high- detail dataset and benchmark for scene flow, optical flow and stereo. In CVPR, 2023. 5
2023
-
[54]
Delta: Dense efficient long-range 3d track- ing for any video
Tuan Duc Ngo, Peiye Zhuang, Chuang Gan, Evange- los Kalogerakis, Sergey Tulyakov, Hsin-Ying Lee, and Chaoyang Wang. Delta: Dense efficient long-range 3d track- ing for any video. arXiv preprint arXiv:2410.24211, 2024. 2, 3, 6
2024 arXiv
-
[55]
An efficient solution to the five-point relative pose problem
David Nist ´er. An efficient solution to the five-point relative pose problem. IEEE transactions on pattern analysis and machine intelligence, 26(6):756–770, 2004. 3
2004
-
[56]
Pre- training auto-regressive robotic models with 4d representa- tions
Dantong Niu, Yuvan Sharma, Haoru Xue, Giscard Biamby, Junyi Zhang, Ziteng Ji, Trevor Darrell, and Roei Herzig. Pre- training auto-regressive robotic models with 4d representa- tions. arXiv preprint arXiv:2502.13142, 2025. 2
2025 arXiv
-
[57]
Codef: Content deformation fields for temporally consistent video processing
Hao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai, Jun- tao Zhang, Kecheng Zheng, Xiaowei Zhou, Qifeng Chen, and Yujun Shen. Codef: Content deformation fields for temporally consistent video processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2024
-
[58]
A survey of structure from motion*
Onur ¨Ozyes ¸il, Vladislav V oroninski, Ronen Basri, and Amit Singer. A survey of structure from motion*. Acta Numerica, 26:305–364, 2017. 3
2017
-
[59]
Palazzolo, J
E. Palazzolo, J. Behley, P. Lottes, P. Gigu `ere, and C. Stach- niss. ReFusion: 3D Reconstruction in Dynamic Environ- ments for RGB-D Cameras Exploiting Residuals. In IROS,
-
[60]
Unidepth: Universal monocular metric depth estimation
Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Seg`u, Siyuan Li, Luc Van Gool, and Fisher Yu. Unidepth: Universal monocular metric depth estimation. In CVPR, pages 10106–10116. IEEE, 2024. 3, 6
2024
-
[61]
Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotn ´y. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In ICCV. IEEE, 2021. 5
2021
-
[62]
Fouhey, and Chen-Hsuan Lin
Chris Rockwell, Joseph Tung, Tsung-Yi Lin, Ming-Yu Liu, David F. Fouhey, and Chen-Hsuan Lin. Dynamic camera poses and where to find them. In CVPR, 2025. 8
2025
-
[63]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bod- ies together. ACM Transactions on Graphics, (Proc. SIG- GRAPH Asia), 36(6), 2017. 2
2017
-
[64]
Particle video: Long-range mo- tion estimation using point trajectories.International journal of computer vision, 80:72–91, 2008
Peter Sand and Seth Teller. Particle video: Long-range mo- tion estimation using point trajectories.International journal of computer vision, 80:72–91, 2008. 2
2008
-
[65]
Common pets in 3d: Dynamic new-view synthesis of real-life deformable categories
Samarth Sinha, Roman Shapovalov, Jeremy Reizenstein, Ig- nacio Rocco, Natalia Neverova, Andrea Vedaldi, and David Novotn´y. Common pets in 3d: Dynamic new-view synthesis of real-life deformable categories. In CVPR, pages 4881–
-
[66]
A benchmark for the eval- uation of RGB-D SLAM systems
J ¨urgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers. A benchmark for the eval- uation of RGB-D SLAM systems. In IROS, pages 573–580. IEEE, 2012. 6, 7, 8
2012
-
[67]
Kalib: Markerless hand-eye calibration with keypoint track- ing
Tutian Tang, Minghao Liu, Wenqiang Xu, and Cewu Lu. Kalib: Markerless hand-eye calibration with keypoint track- ing. CoRR, abs/2408.10562, 2024. 2
2024
-
[68]
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part II 16, pages 402–419. Springer,
2020
-
[69]
Bundle adjustment—a mod- 11 ern synthesis
Bill Triggs, Philip F McLauchlan, Richard I Hartley, and Andrew W Fitzgibbon. Bundle adjustment—a mod- 11 ern synthesis. In Vision Algorithms: Theory and Prac- tice: International Workshop on Vision Algorithms Corfu, Greece, September 21–22, 1999 Proceedings , pages 298–
1999
-
[70]
Marigold-dc: Zero-shot monocular depth completion with guided diffusion
Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker, Konrad Schindler, and Anton Obukhov. Marigold-dc: Zero-shot monocular depth completion with guided diffusion. CoRR, abs/2412.13389, 2024. 3
2024 arXiv
-
[71]
Scenetracker: Long-term scene flow estimation network
Bo Wang, Jian Li, Yang Yu, Li Liu, Zhenping Sun, and Dewen Hu. Scenetracker: Long-term scene flow estimation network. arXiv preprint arXiv:2403.19924, 2024. 3
2024 arXiv
-
[72]
Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment
Jianyuan Wang, Christian Rupprecht, and David Novotny. Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 9773–9783,
-
[73]
Vggsfm: Visual geometry grounded deep structure from motion
Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. Vggsfm: Visual geometry grounded deep structure from motion. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 21686–21697, 2024. 3
2024
-
[74]
Vggt: Vi- sual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Vi- sual geometry grounded transformer. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 5294–5306, 2025. 2, 3, 6, 7, 8
2025
-
[75]
Tracking everything everywhere all at once
Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, and Noah Snavely. Tracking everything everywhere all at once. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 19795–19806, 2023. 2, 3
2023
-
[76]
Shape of motion: 4d reconstruc- tion from a single video
Qianqian Wang, Vickie Ye, Hang Gao, Jake Austin, Zhengqi Li, and Angjoo Kanazawa. Shape of motion: 4d reconstruc- tion from a single video. arXiv preprint arXiv:2407.13764,
-
[77]
Continuous 3d perception model with persistent state
Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A Efros, and Angjoo Kanazawa. Continuous 3d perception model with persistent state. arXiv preprint arXiv:2501.12387, 2025. 6, 7, 8
2025 arXiv
-
[79]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20697– 20709, 2024. 6, 7, 8
2024
-
[80]
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Se- bastian Scherer. Tartanair: A dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4909–49...
2020
-
[81]
Objctrl-2.5 d: Training-free object con- trol with camera poses
Zhouxia Wang, Yushi Lan, Shangchen Zhou, and Chen Change Loy. Objctrl-2.5 d: Training-free object con- trol with camera poses. arXiv preprint arXiv:2412.07721 ,
-
[82]
Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation
Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Jiawei Ren, Liang Pan, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, Dahua Lin, and Ziwei Liu. Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation. In CVPR, pages 803–814. IEEE, 2023. 5
2023
-
[83]
RGBD objects in the wild: Scaling real-world 3d object learning from RGB-D videos
Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang. RGBD objects in the wild: Scaling real-world 3d object learning from RGB-D videos. In CVPR, pages 22378– 22389. IEEE, 2024. 5
2024
-
[84]
Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv preprint arXiv:1711.00199, 2017. 3
2017 arXiv
-
[85]
Level- S2fM: Structure From Motion on Neural Level Set of Im- plicit Surfaces
Yuxi Xiao, Nan Xue, Tianfu Wu, and Gui-Song Xia. Level- S2fM: Structure From Motion on Neural Level Set of Im- plicit Surfaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17205– 17214, 2023. 3
2023
-
[86]
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20406–20417, 2024. 2, 3, 6
2024
-
[87]
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10371–10381, 2024. 2, 3, 7
2024
-
[88]
Depth any- thing v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2. Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2025. 2, 3
2025
-
[89]
Neural window fully-connected crfs for monocular depth estimation
Weihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu, and Ping Tan. Neural window fully-connected crfs for monocular depth estimation. In CVPR, pages 3906–3915. IEEE, 2022. 3
2022
-
[90]
Tapip3d: Tracking any point in persistent 3d geome- try
Bowei Zhang, Lei Ke, Adam W Harley, and Katerina Fragki- adaki. Tapip3d: Tracking any point in persistent 3d geome- try. arXiv preprint arXiv:2504.14717, 2025. 3, 6
2025
-
[91]
Monst3r: A simple approach for estimat- ing geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jam- pani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming- Hsuan Yang. Monst3r: A simple approach for estimat- ing geometry in the presence of motion. arXiv preprint arXiv:2410.03825, 2024. 6, 7, 8
-
[92]
Rel- pose: Predicting probabilistic relative rotation for single ob- jects in the wild
Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani. Rel- pose: Predicting probabilistic relative rotation for single ob- jects in the wild. In ECCV, pages 592–611. Springer, 2022. 3
2022
-
[93]
Cameras as rays: Pose estimation via ray diffusion
Jason Y Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. arXiv preprint arXiv:2402.14817, 2024. 3
2024 arXiv
-
[94]
Zhoutong Zhang, Forrester Cole, Zhengqi Li, Michael Ru- binstein, Noah Snavely, and William T. Freeman. Struc- ture and motion from casual videos. In ECCV, pages 20–37. Springer, 2022. 8 12
2022
-
[95]
Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild
Wang Zhao, Shaohui Liu, Hengkai Guo, Wenping Wang, and Yong-Jin Liu. Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild. In ECCV. Springer, 2022. 2, 8
2022
-
[96]
Pointodyssey: A large-scale synthetic dataset for long-term point tracking
Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wet- zstein, and Leonidas J Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 19855–19865, 2023. 5, 9 13
2023
-
[2011]
Computer Vision Foundation / IEEE Computer Soci- ety, 2018. 3
2018
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.