Pith. sign in

REVIEW 4 major objections 5 minor 10 cited by

SpatialTrackerV2: 3D Point Tracking Made Easy

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A single feed-forward model now sets the 3D tracking record on TAPVid-3D, matching a slow optimizer at 50x its speed.

desk verdict A credible SOTA for 3D point tracking with a genuinely new dual-branch architecture and an impressive joint-training recipe, but the headline egocentric gain rests on an unvalidated MoGe teacher and the paper needs missing ablations and code before I'd fully trust the numbers. read the letter →

arxiv 2507.12462 v2 pith:VLBOM5IW submitted 2025-07-16 cs.CV

classification cs.CV
keywords 3Dpointtrackingmonocularvideodepthestimationcameraposefeed-forwardarchitectureSyncFormerbundleadjustmentdynamicscenereconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that 3D point tracking in monocular video becomes more accurate and scalable when the problem is decomposed into three jointly learned pieces: scene geometry (video depth), camera ego-motion, and pixel-wise object motion. The authors build a fully differentiable, end-to-end pipeline that estimates all three from a single forward pass, so no per-video optimization is needed. On the TAPVid-3D benchmark it reports 21.2 average Jaccard and 31.0 3D position error, surpassing the previous best tracker by 61.8% and 50.5%, and matching the dynamic-reconstruction accuracy of a leading optimization-based method while running about 50 times faster. If correct, a single model trained on heterogeneous data—synthetic sequences, RGB-D videos, and unlabeled Internet footage—can replace modular pipelines that chain together off-the-shelf 2D trackers, depth estimators, and pose solvers.

What carries the argument

SyncFormer, an iterative transformer with separate 2D (image-space) and 3D (camera-coordinate) branches whose representations are exchanged through cross-attention between compact proxy tokens. The 2D branch updates tracks in UV space while the 3D branch updates positions in camera coordinates using correlations computed on normalized point maps. A weighted Procrustes alignment, weighted by the predicted dynamic scores, registers the 3D trajectories to the world frame, and a bundle-adjustment step optimizes camera poses from the filtered static points. The front end combines a temporal encoder (built by extending a monocular depth architecture with alternating intra- and inter-frame attention and learnable pose/scale tokens) that outputs depth and initial camera poses. The design exists to keep 2D tracking, 3D geometry, and ego-motion consistent without letting any one task's errors corrupt the others.

What would settle it

Retrain the pipeline with teacher depth supervision removed from the pose-only and unlabeled datasets and re-measure 3D tracking and video-depth accuracy; a large drop would confirm the teacher is load-bearing. Separately, compare the predicted depth and 3D tracks against LiDAR ground truth on egocentric sequences to test whether teacher bias in that domain propagates into the final geometry.

Watch

Extended reading notes

Core claim

SpatialTrackerV2 is a feed-forward 3D point tracker that unifies point tracking, monocular depth estimation, and camera pose estimation into a single differentiable network. It decomposes world-space 3D motion into video depth, camera ego-motion, and per-pixel object motion; a front end encodes temporal video context to produce scale-aligned depth and an initial camera trajectory, and a back-end module iteratively refines 2D and 3D trajectories together with visibility and dynamic/static scores. The refinements feed back into a differentiable bundle-adjustment step that re-estimates camera poses, and the entire loop is trained jointly on 17 datasets spanning full supervision (RGB-D with 3D track labels) to pose-only and unlabeled footage. The paper reports state-of-the-art results on a standard 3D tracking benchmark (21.2 average Jaccard, 31.0 3D position error, 90.6 occlusion accuracy), video depth that improves on both a strong feed-forward model and an optimization-based system (0.081 absolute relative error), and camera poses on par with the optimization-based system at roughly 50x lower inference time.

Load-bearing premise

The scaling result rests on a teacher monocular depth model being right on pose-only and unlabeled videos, since the paper never measures that teacher's accuracy on egocentric or Internet footage and its errors would be baked into the learned geometry and 3D trajectories.

Editorial extensions

If this is right

  • Per-video optimization pipelines can be replaced by a single forward pass for many dynamic-reconstruction tasks, shrinking inference from minutes to seconds.
  • Because the pipeline trains on pose-only and unlabeled video, depth and 3D tracking should keep improving as more in-the-wild data is added, without needing expensive 3D track annotations.
  • The decoupled 2D/3D branches mean that adding 3D supervision no longer degrades 2D tracking accuracy: the paper's naive-lifting baseline drops average Jaccard from 64.4 to 51.6, while SyncFormer keeps it at 64.9.
  • The explicit camera-motion decomposition yields the largest gains on egocentric and background-point-heavy subsets, so the method should be especially useful for wearable and robotic video.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The dynamic-score weighting in the Procrustes alignment is a self-supervised bootstrap; an ablation comparing against an oracle static/dynamic segmentation would quantify how much error the circularity introduces.
  • If the teacher depth model is biased on egocentric or Internet footage, the paper's scaling story predicts that a better teacher or a second, independent geometric prior would directly improve 3D tracking on those domains, which is a testable prediction.
  • The same decomposition (depth plus ego-motion plus object motion) could serve as a representation for downstream tasks like video editing, generation control, and robotic grasping, potentially letting those tasks consume trajectory and geometry directly rather than training their own motion estimators.
  • The 50x speed advantage suggests that feed-forward geometry-and-motion models could enable real-time structure-from-motion-like applications on mobile devices, which the paper does not discuss.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents SpatialTrackerV2, a fully feed-forward model for monocular 3D point tracking that jointly estimates video depth, camera poses, and long-term 2D/3D point trajectories. The architecture combines a temporal encoder front-end for scale-aligned depth and camera initialization with a SyncFormer back-end that alternates between 2D and 3D correlation branches and performs differentiable bundle adjustment, using learned dynamic and visibility scores to weight the optimization. Training mixes posed RGB-D data with tracking annotations, posed RGB-D without tracking, and pose-only or unlabeled video where a monocular teacher provides relative depth. The paper reports state-of-the-art results on TAPVid-3D (21.2 AJ, 31.0 APD3D), improvements over prior 3D trackers such as DELTA, and faster feed-forward inference than optimization-based reconstruction systems like MegaSAM. The main empirical tables support the headline tracking claims, while the depth and pose comparisons are more mixed than the abstract suggests.

Significance. If the benchmark results hold, the paper offers a meaningful advance: a single model that jointly produces metric-scale video depth, camera poses, and 3D point trajectories, with a unified training scheme across heterogeneous datasets. The evaluation protocol is largely fair, with comparisons controlled for the depth and pose provider (e.g., Ours-offl+ versus Type I baselines), and the paper reports results across three benchmark subsets. The biggest reported gain, 21.2 versus 13.1 AJ over DELTA, is large and directionally consistent across Aria, DriveTrack, and PStudio. The paper also provides a public demo and includes useful implementation details such as dataset categories and training stages. The main caveats are that the 'matches MegaSAM' claim is not uniformly supported by Tables 2 and 3, and that several load-bearing components (the MoGe teacher, the dynamic-score bootstrap, and the missing ablation rows) are not yet validated or documented. The submitted version is a solid empirical contribution but needs revision before the claims can be taken at face value.

major comments (4)
  1. [Abstract; Section 4.2, Tables 2 and 3] The abstract's statement that the method 'matches the accuracy of leading dynamic 3D reconstruction approaches' is not supported by the paper's own tables. In Table 3, on Sintel, Ours (ATE 0.054, RPEt 0.027, RPEr 0.288) is substantially worse than MegaSAM (ATE 0.023, RPEt 0.008, RPEr 0.060), and on Lightspeed, Ours has ATE 0.134 versus MegaSAM's 0.105. In Table 2, on Sintel, MegaSAM is also better on AbsRel (0.185 vs. 0.199) and δ1.25 (0.746 vs. 0.703). The average depth numbers favor the proposed method, but the 'matches' claim should be qualified to the datasets and metrics where it actually holds, or the text should explain why the Sintel/Lightspeed gaps are acceptable.
  2. [Section 4.3, Tables 2 and 5] The ablation analysis is incomplete as presented. In the 3D Point Tracking ablation, the text states that 'our final version is clearly better than these two' but Table 5 contains only Base@K and Base@V,K,P, with no row for the final model. In the Depth Estimation ablation, the text says 'As shown in Tab. 2' for rows Ours-Synthetic and Ours-Real-Full, but Table 2 does not contain these rows. These missing rows make the scaling and joint-training claims in Section 4.3 unverifiable. The authors should either add the missing rows or remove the references.
  3. [Section 3.3, training data category (3)] The paper's scalable-training story depends on the MoGe teacher [78] for preserving relative depth on HOI4D, Ego4D, and Stereo4D, which are the main egocentric and Internet-video training sources. The TAPVid-3D Aria subset, where the largest gains are reported (24.6 vs. 23.5 AJ for the next best method), is egocentric, yet the paper provides no evaluation of MoGe's depth accuracy on these domains and no ablation that removes or replaces the teacher. Since a systematic teacher bias would be distilled into the video depth and then into the 3D tracks via Eq. (2), I ask the authors to either include a teacher-quality analysis (e.g., relative depth consistency on a held-out egocentric split) or an ablation with an alternative teacher such as DepthAnythingV2 or with the teacher removed.
  4. [Section 3.2, camera motion optimization] The dynamic-score bootstrap is a potential source of bias. In the weighted Procrustes alignment, the learned dynamic probability pdyn decides which points contribute to the camera pose optimization, and the updated poses then feed back into the dynamic and visibility scores in the next SyncFormer iteration. On category (3) data there are no ground-truth dynamic masks, so the loop is entirely self-supervised; if dynamic points are misclassified, the biased poses could reinforce the error in the scores. The paper should at least analyze this failure mode (e.g., by comparing the learned pdyn against ground-truth masks on category (1) data and by ablating the dynamic filtering) or explicitly justify why the self-consistency signal is sufficient.
minor comments (5)
  1. [Abstract] The phrase 'outperforms existing 3D tracking methods by 30%' is ambiguous: relative to DELTA the gain is 61.8% in AJ, while relative to TAPIP3D it is 12.8%; please specify the reference point and metric.
  2. [Table 1] The 'Full-ours' row lacks Type and Depth/Cam Pose entries, making its protocol (world-space, own depth and pose) implicit; please add them for clarity.
  3. [Table 2] The citation for MegaSAM is given as [74] but should be [42].
  4. [Section 4.1, Table 4] The baseline name is written inconsistently as 'Cotracker3' in the text and 'CoTracker3' elsewhere; please standardize the spelling.
  5. [General] The paper does not mention whether code and trained models will be released; given the 17-dataset training recipe and the dependence on the MoGe teacher, this is important for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the benchmark claims are externally evaluated and the training pipeline uses heterogeneous external supervision; no equation reduces to a fitted parameter.

full rationale

SpatialTrackerV2's central claims are empirical numbers on held-out benchmarks: TAPVid-3D for 3D tracking, KITTI/TUM/Bonn/Sintel for video depth, and TUM-dynamic/Lightspeed/Sintel for camera poses. These datasets provide ground truth independent of the model's fitted parameters. The training recipe uses external GT tracks, depth, and poses on categories (1)-(2), and an external MoGe teacher for relative depth on category (3); the paper's ablations, e.g., Base@K vs. Base@V,K,P and joint-training comparisons, isolate the contribution of data scaling and joint optimization. The dynamic-score-weighted Procrustes loop in Section 3.2 is a bootstrapped training objective that uses the model's own predictions as alignment weights; it is not a reported prediction against ground truth, and the evaluated AJ/APD3D/OA numbers come from TAPVid-3D annotations, so no evaluated result is forced by construction. The self-citations to VGGT, CoTracker3, SpatialTracker, and Video Depth Anything are architectural baselines, initializations, or comparative references; they are externally benchmarked and are not invoked as a uniqueness theorem or as the sole justification for the state-of-the-art claim. The unvalidated MoGe teacher on egocentric/Internet domains and the absence of a no-MoGe ablation are legitimate empirical robustness concerns, but they are not definitional circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim is an empirical benchmark result, so the ledger records the modeling assumptions and hand-set hyperparameters that the result depends on. The network weights themselves are learned from supervision and are not hidden free parameters.

free parameters (4)
  • Correlation window radius = 3 (Delta = 3)
    Hand-picked for 3D correlation features; no ablation shown.
  • Number of correlation scales = 4 (s = 1,2,3,4)
    Hand-picked for multi-scale 3D correlations; no ablation shown.
  • SyncFormer iteration count = unspecified
    Number of iterative updates is not stated in the paper; affects quality and compute.
  • Joint training loss weights = unspecified
    Weights for depth, pose, 2D, 3D, dynamic, visibility losses are not reported.
assumptions (4)
  • domain assumption A monocular video can be decomposed into static scene geometry, camera ego-motion, and pixel-wise object motion.
    This is the core modeling assumption of the paper (Section 3). It fails for scenes with fluids, mirrors, or non-rigid camera-relative motion that is not separable.
  • domain assumption Monocular depth is recoverable up to scale and shift, and the learned scale-shift parameters a,b sufficiently align it with camera poses.
    Equations (1) and (2) assume a single global scale and shift per video can align depth to poses; this is a common monocular depth assumption but not always valid.
  • domain assumption The teacher depth model MoGe provides sufficiently accurate relative depth for unlabeled and pose-only training data.
    Section 3.3 category (3) relies on this without independent verification in the target domains.
  • ad hoc to paper Dynamic points can be identified by a learned probability pdyn and filtering them before Procrustes alignment improves pose estimation.
    Section 3.2 uses dynamic scores to weight correspondences; if this assumption fails, pose optimization degrades, a risk the paper does not ablate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SpatialTrackerV2: 3D Point Tracking Made Easy." pith.science (2026). https://pith.science/paper/VLBOM5IW

@misc{pith2026250712462,
  author       = {Pith},
  title        = {Pith review of: SpatialTrackerV2: 3D Point Tracking Made Easy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VLBOM5IW}},
  note         = {Machine review of arXiv:2507.12462}
}
abstract

We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50$\times$ faster.

Figures

Figures reproduced from arXiv: 2507.12462 by the authors.

Figure 1
Figure 1. SpatialTrackerV2 produces consistent 3D scene geometry, camera poses, and 3D point trajectories all at once from monocular videos of arbitrary scenarios, e.g., robotic manipulation, first-person egocentric views, and dynamic sports (drifting and skating) shown in this figure. Try our online demo at https://huggingface.co/spaces/Yuxihenry/SpatialTrackerV2. Abstract We present SpatialTrackerV2, a feed-forward 3D point… view at source ↗
Figure 2
Figure 2. Pipeline Overview. Our method adopts a front-end and back-end architecture. The front-end estimates scale-aligned depth and camera poses from the input video, which are used to construct initial static 3D tracks. The back-end then iteratively refines both tracks and poses via joint motion optimization. 3D Embedding 2D Embedding 3D Proxy tokens Cross A'en*on 2D Proxy tokens Diff BA Self Attention T 2d k+1 T 3d k+1 pd… view at source ↗
Figure 3
Figure 3. SyncFormer. The model takes previous estimates and their corresponding embeddings as input, and updates them iter￾atively. The 2D and 3D embeddings are processed in separate branches that interact via cross-attention. to as SyncFormer. In parallel, the SyncFormer also dynam￾ically estimates the visibility probability p vis and dynamic probability p dyn for each trajectory, enabling an efficient bundle adjustment pro… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Iterative Process of SyncFormer. The left subfig￾ure shows the convergence curve of reprojection error, illustrat￾ing rapid reprojection error reduction. The right subfigures visu￾alizes the progressive alignment between 2D tracking results (in green) and projections o…
Figure 5
Figure 5. Figure 5: Fused Point Clouds, Camera Poses, and 3D Point Trajectories. We visualize the fused point clouds reconstructed from our video depth and camera poses, along with long-term 3D point trajectories in world space. … VGGT SpatialTrackerV2 … [PITH_FULL_IMAGE:figures/full_fig…
Figure 6
Figure 6. Figure 6: Qualitative Comparisons on Internet Videos. To as￾sess generalization, we compare our method with VGGT [74] on challenging Internet videos. flects the accuracy of the underlying 2D tracks and the associated depth predictions. Intriguingly, CoTracker3 + MegaSAM achieves…
Figure 7
Figure 7. Figure 7: The influence of Joint Training [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data

    cs.RO 2026-07 accept novelty 6.5 of 10

    World Action Model co-training with DINO or 3D-flow targets scales human-to-robot transfer on bimanual tasks far better than behavior cloning, while pixel prediction transfers weakly.

  2. FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single video-diffusion framework composites both static images and dynamic footage along user-defined trajectories by transporting canonical foreground latents directly into the background latent sequence.

  3. 4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    A training-free two-pass adaptation of VGGT, with attention-based motion masking and inverse-variance depth fusion, improves dynamic-scene point-cloud reconstruction on DyCheck.

  4. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SymphoMotion jointly controls camera trajectories and depth-aware object dynamics inside one video diffusion model, supported by the new RealCOD-25K real-world paired-motion dataset.

  5. 4DGS360: 360{\deg} Gaussian Reconstruction of Dynamic Objects from a Single Video

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Combining high-confidence 2D tracking anchors with a 3D point tracker improves initialization and monocular 360-degree dynamic object reconstruction, demonstrated on a new far-viewpoint benchmark.

  6. IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A causal streaming transformer that jointly predicts camera motion, 3D geometry, and persistent object-instance features from video, trained on a new 147K-sequence 4D dataset.

  7. TrackDeform3D: Markerless and Autonomous 3D Keypoint Tracking and Dataset Collection for Deformable Objects

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A markerless RGB-D tracking pipeline for deformable objects plus a released 110-minute, six-object trajectory dataset.

  8. Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A trajectory-conditioned retrieval system discovers multiple motion descriptions in videos without user queries and grounds them to point tracks, evaluated mainly on MeViS.

  9. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

  10. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

Reference graph

Works this paper leans on

97 extracted references · 54 canonical work pages · cited by 10 Pith papers

  1. [78]

    Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision, 2024

    Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision, 2024. 5

  2. [1]

    L4p: Low-level 4d vision perception unified

    Abhishek Badki, Hang Su, Bowen Wen, and Orazio Gallo. L4p: Low-level 4d vision perception unified. arXiv preprint arXiv:2502.13078, 2025. 2

  3. [2]

    Speeded-Up Robust Features (SURF)

    Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool. Speeded-Up Robust Features (SURF). CVIU, 110(3),

  4. [3]

    Zoedepth: Zero-shot trans- fer by combining relative and metric depth

    Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias M ¨uller. Zoedepth: Zero-shot trans- fer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023. 3

  5. [4]

    Gs-dit: Advancing video generation with pseudo 4d gaussian fields through efficient dense 3d point tracking

    Weikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yijin Li, Fu- Yun Wang, and Hongsheng Li. Gs-dit: Advancing video generation with pseudo 4d gaussian fields through efficient dense 3d point tracking. CoRR, abs/2501.02690, 2025. 2

  6. [5]

    Midas v3.1 - A model zoo for robust monocular relative depth estimation

    Reiner Birkl, Diana Wofk, and Matthias M¨uller. Midas v3.1 - A model zoo for robust monocular relative depth estimation. CoRR, abs/2307.14460, 2023. 3

  7. [6]

    Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang

    Michael J. Black, Priyanka Patel, Joachim Tesch, and Jin- long Yang. BEDLAM: A synthetic dataset of bodies exhibit- ing detailed lifelike animated motion. InCVPR, pages 8726–

  8. [7]

    Vir- tual kitti 2, 2020

    Yohann Cabon, Naila Murray, and Martin Humenberger. Vir- tual kitti 2, 2020. 5, 9

Show all 97 references
  1. [8]

    Emerg- ing properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 3

  2. [9]

    Video depth anything: Consistent depth estimation for super-long videos

    Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zi- long Huang, Jiashi Feng, and Bingyi Kang. Video depth anything: Consistent depth estimation for super-long videos. CoRR, abs/2501.12375, 2025. 3, 5, 6, 7

  3. [10]

    Pref3r: Pose- free feed-forward 3d gaussian splatting from variable-length image sequence

    Zequn Chen, Jiezhi Yang, and Heng Yang. Pref3r: Pose- free feed-forward 3d gaussian splatting from variable-length image sequence. CoRR, abs/2411.16877, 2024. 2

  4. [11]

    Local all-pair correspon- dence for point tracking

    Seokju Cho, Jiahui Huang, Jisu Nam, Honggyu An, Seun- gryong Kim, and Joon-Young Lee. Local all-pair correspon- dence for point tracking. In European Conference on Com- puter Vision, pages 306–325. Springer, 2024. 2

  5. [12]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017. 5, 8

  6. [13]

    Diffusion models beat gans on image synthesis

    Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Informa- tion Processing Systems, 34:8780–8794, 2021. 3

  7. [14]

    Tap-vid: A benchmark for track- ing any point in a video

    Carl Doersch, Ankush Gupta, Larisa Markeeva, Adria Re- casens, Lucas Smaira, Yusuf Aytar, Joao Carreira, Andrew Zisserman, and Yi Yang. Tap-vid: A benchmark for track- ing any point in a video. Advances in Neural Information Processing Systems, 35:13610–13626, 2022. 2

  8. [15]

    Tapir: Tracking any point with per-frame initialization and 9 temporal refinement

    Carl Doersch, Yi Yang, Mel Vecerik, Dilara Gokay, Ankush Gupta, Yusuf Aytar, Joao Carreira, and Andrew Zisserman. Tapir: Tracking any point with per-frame initialization and 9 temporal refinement. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , p...

  9. [16]

    Bootstap: Bootstrapped training for tracking-any-point

    Carl Doersch, Pauline Luc, Yi Yang, Dilara Gokay, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ignacio Rocco, Ross Goroshin, Jo˜ao Carreira, et al. Bootstap: Bootstrapped training for tracking-any-point. In Proceedings of the Asian Conference on Computer Vision, pages 3257–32...

  10. [17]

    Depth map prediction from a single image using a multi-scale deep net- work

    David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep net- work. In NeurIPS, pages 2366–2374, 2014. 3

  11. [18]

    Deep ordinal regression net- work for monocular depth estimation

    Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Bat- manghelich, and Dacheng Tao. Deep ordinal regression net- work for monocular depth estimation. InCVPR, pages 2002–

  12. [19]

    Vision meets robotics: The kitti dataset

    Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. The in- ternational journal of robotics research, 32(11):1231–1237,

  13. [20]

    Ego4d: Around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vi- sio...

  14. [21]

    Kubric: A scalable dataset generator

    Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapra- gasam, Florian Golemo, Charles Herrmann, et al. Kubric: A scalable dataset generator. In Proceedings of the IEEE/CVF conference on computer vision and pattern re...

  15. [22]

    Diffusion as shader: 3d- aware video diffusion for versatile video generation control

    Zekai Gu, Rui Yan, Jiahao Lu, Peng Li, Zhiyang Dou, Chenyang Si, Zhen Dong, Qifeng Liu, Cheng Lin, Ziwei Liu, Wenping Wang, and Yuan Liu. Diffusion as shader: 3d- aware video diffusion for versatile video generation control. CoRR, abs/2501.03847, 2025. 2

  16. [23]

    Harley, Zhaoyuan Fang, and Katerina Fragkiadaki

    Adam W. Harley, Zhaoyuan Fang, and Katerina Fragkiadaki. Particle video revisited: Tracking through occlusions using point trajectories. In ECCV, 2022. 2

  17. [24]

    Multiple View Ge- ometry in Computer Vision

    Richard Hartley and Andrew Zisserman. Multiple View Ge- ometry in Computer Vision . Cambridge University Press, ISBN: 0521540518, second edition, 2004. 3

  18. [25]

    In defense of the eight-point algorithm

    Richard I Hartley. In defense of the eight-point algorithm. IEEE Transactions on pattern analysis and machine intelli- gence, 19(6):580–593, 1997. 3

  19. [26]

    Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface nor- mal estimation

    Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface nor- mal estimation. IEEE TPAMI, 46(12):10579–10596, 2024. 3

  20. [27]

    Depthcrafter: Generating consistent long depth sequences for open-world videos

    Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xiaodong Cun, Yong Zhang, Long Quan, and Ying Shan. Depthcrafter: Generating consistent long depth sequences for open-world videos. CoRR, abs/2409.02095, 2024. 2, 3, 6, 7

  21. [28]

    Deepmvs: Learning multi-view stereopsis

    Po-Han Huang, Kevin Matzen, Johannes Kopf, Narendra Ahuja, and Jia-Bin Huang. Deepmvs: Learning multi-view stereopsis. In CVPR, pages 2821–2830. Computer Vision Foundation / IEEE Computer Society, 2018. 5

  22. [29]

    Mitra, and Duygu Ceylan

    Hyeonho Jeong, Chun-Hao Paul Huang, Jong Chul Ye, Niloy J. Mitra, and Duygu Ceylan. Track4gen: Teaching video diffusion models to track points improves video gen- eration. CoRR, abs/2412.06016, 2024. 2

  23. [30]

    Stereo4d: Learning how things move in 3d from internet stereo videos.arXiv preprint,

    Linyi Jin, Richard Tucker, Zhengqi Li, David Fouhey, Noah Snavely, and Aleksander Holynski. Stereo4d: Learning how things move in 3d from internet stereo videos.arXiv preprint,

  24. [31]

    Dy- namicstereo: Consistent dynamic depth from stereo videos

    Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Dy- namicstereo: Consistent dynamic depth from stereo videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13229–13239, 2023. 5

  25. [32]

    Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos

    Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos. arXiv preprint arXiv:2410.11831 ,

  26. [33]

    Co- tracker: It is better to track together

    Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker: It is better to track together. InEuropean Conference on Computer Vision, pages 18–35. Springer, 2024. 2, 3, 4, 8

  27. [34]

    Fast encoder- based 3d from casual videos via point track processing.arXiv preprint arXiv:2404.07097, 2024

    Yoni Kasten, Wuyue Lu, and Haggai Maron. Fast encoder- based 3d from casual videos via point track processing.arXiv preprint arXiv:2404.07097, 2024. 2

  28. [35]

    Ssd-6d: Making rgb-based 3d detec- tion and 6d pose estimation great again

    Wadim Kehl, Fabian Manhardt, Federico Tombari, Slobodan Ilic, and Nassir Navab. Ssd-6d: Making rgb-based 3d detec- tion and 6d pose estimation great again. In Proceedings of the IEEE international conference on computer vision, pages 1521–1529, 2017. 3

  29. [36]

    Robust consistent video depth estimation

    Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Robust consistent video depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, vir- tual, June 19-25, 2021, pages 1611–1621. Computer Vision Foundation / IEEE, 2021. 8

  30. [37]

    Tapvid-3d: A benchmark for tracking any point in 3d

    Skanda Koppula, Ignacio Rocco, Yi Yang, Joseph Heyward, Jo˜ao Carreira, Andrew Zisserman, Gabriel Brostow, and Carl Doersch. Tapvid-3d: A benchmark for tracking any point in 3d. In NeurIPS, 2024. 2, 5

  31. [38]

    Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds

    Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas, and Kostas Daniilidis. Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds. arXiv preprint arXiv:2405.17421, 2024. 2

  32. [39]

    Five-point motion esti- mation made easy

    Hongdong Li and Richard Hartley. Five-point motion esti- mation made easy. In 18th International Conference on Pat- tern Recognition (ICPR’06), pages 630–633. IEEE, 2006. 3

  33. [40]

    Taptr: Tracking any point with transformers as detection

    Hongyang Li, Hao Zhang, Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, and Lei Zhang. Taptr: Tracking any point with transformers as detection. In European Confer- ence on Computer Vision, pages 57–75. Springer, 2024. 2

  34. [41]

    Drip: Unleashing diffusion priors for joint foreground and alpha prediction in image matting.Advances in Neural Information Processing Systems, 37:79868–79888, 2024

    Xiaodi Li, Zongxin Yang, Ruijie Quan, and Yi Yang. Drip: Unleashing diffusion priors for joint foreground and alpha prediction in image matting.Advances in Neural Information Processing Systems, 37:79868–79888, 2024. 3 10

  35. [42]

    Megasam: Accurate, fast and robust structure and motion from casual dynamic videos

    Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang, Linyi Jin, Vickie Ye, Angjoo Kanazawa, Aleksander Holyn- ski, and Noah Snavely. Megasam: Accurate, fast and robust structure and motion from casual dynamic videos. In CVPR,

  36. [43]

    Relpose++: Recovering 6d poses from sparse-view observations

    Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tul- siani. Relpose++: Recovering 6d poses from sparse-view observations. arXiv preprint arXiv:2305.04926, 2023. 3

  37. [44]

    Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

    Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , p...

  38. [45]

    Sift flow: Dense correspondence across scenes and its applications

    Ce Liu, Jenny Yuen, and Antonio Torralba. Sift flow: Dense correspondence across scenes and its applications. IEEE transactions on pattern analysis and machine intelligence , 33(5):978–994, 2010. 2

  39. [46]

    HOI4D: A 4d egocentric dataset for category-level human- object interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4d egocentric dataset for category-level human- object interaction. In CVPR, pages 20981–20990. IEEE,

  40. [47]

    Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, 2015. 2

  41. [48]

    Robo-gs: A physics consistent spatial-temporal model for robotic arm with hybrid representation

    Haozhe Lou, Yurong Liu, Yike Pan, Yiran Geng, Jianteng Chen, Wenlong Ma, Chenglong Li, Lin Wang, Hengzhen Feng, Lu Shi, Liyi Luo, and Yongliang Shi. Robo-gs: A physics consistent spatial-temporal model for robotic arm with hybrid representation. CoRR, abs/2408.14873, 2024. 2

  42. [49]

    David G. Lowe. Object Recognition from Local Scale- Invariant Features. In Proc. ICCV, 1999. 3

  43. [50]

    David G. Lowe. Distinctive Image Features from Scale- Invariant Keypoints. IJCV, 60(2), 2004. 3

  44. [51]

    Virtual correspondence: Hu- mans as a cue for extreme-view geometry

    Wei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun, and Antonio Torralba. Virtual correspondence: Hu- mans as a cue for extreme-view geometry. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15924–15934, 2022. 3

  45. [52]

    A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation

    Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and ...

  46. [53]

    Spring: A high-resolution high- detail dataset and benchmark for scene flow, optical flow and stereo

    Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nali- vayko, and Andr ´es Bruhn. Spring: A high-resolution high- detail dataset and benchmark for scene flow, optical flow and stereo. In CVPR, 2023. 5

  47. [54]

    Delta: Dense efficient long-range 3d track- ing for any video

    Tuan Duc Ngo, Peiye Zhuang, Chuang Gan, Evange- los Kalogerakis, Sergey Tulyakov, Hsin-Ying Lee, and Chaoyang Wang. Delta: Dense efficient long-range 3d track- ing for any video. arXiv preprint arXiv:2410.24211, 2024. 2, 3, 6

  48. [55]

    An efficient solution to the five-point relative pose problem

    David Nist ´er. An efficient solution to the five-point relative pose problem. IEEE transactions on pattern analysis and machine intelligence, 26(6):756–770, 2004. 3

  49. [56]

    Pre- training auto-regressive robotic models with 4d representa- tions

    Dantong Niu, Yuvan Sharma, Haoru Xue, Giscard Biamby, Junyi Zhang, Ziteng Ji, Trevor Darrell, and Roei Herzig. Pre- training auto-regressive robotic models with 4d representa- tions. arXiv preprint arXiv:2502.13142, 2025. 2

  50. [57]

    Codef: Content deformation fields for temporally consistent video processing

    Hao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai, Jun- tao Zhang, Kecheng Zheng, Xiaowei Zhou, Qifeng Chen, and Yujun Shen. Codef: Content deformation fields for temporally consistent video processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  51. [58]

    A survey of structure from motion*

    Onur ¨Ozyes ¸il, Vladislav V oroninski, Ronen Basri, and Amit Singer. A survey of structure from motion*. Acta Numerica, 26:305–364, 2017. 3

  52. [59]

    Palazzolo, J

    E. Palazzolo, J. Behley, P. Lottes, P. Gigu `ere, and C. Stach- niss. ReFusion: 3D Reconstruction in Dynamic Environ- ments for RGB-D Cameras Exploiting Residuals. In IROS,

  53. [60]

    Unidepth: Universal monocular metric depth estimation

    Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Seg`u, Siyuan Li, Luc Van Gool, and Fisher Yu. Unidepth: Universal monocular metric depth estimation. In CVPR, pages 10106–10116. IEEE, 2024. 3, 6

  54. [61]

    Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction

    Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotn ´y. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In ICCV. IEEE, 2021. 5

  55. [62]

    Fouhey, and Chen-Hsuan Lin

    Chris Rockwell, Joseph Tung, Tsung-Yi Lin, Ming-Yu Liu, David F. Fouhey, and Chen-Hsuan Lin. Dynamic camera poses and where to find them. In CVPR, 2025. 8

  56. [63]

    Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bod- ies together. ACM Transactions on Graphics, (Proc. SIG- GRAPH Asia), 36(6), 2017. 2

  57. [64]

    Particle video: Long-range mo- tion estimation using point trajectories.International journal of computer vision, 80:72–91, 2008

    Peter Sand and Seth Teller. Particle video: Long-range mo- tion estimation using point trajectories.International journal of computer vision, 80:72–91, 2008. 2

  58. [65]

    Common pets in 3d: Dynamic new-view synthesis of real-life deformable categories

    Samarth Sinha, Roman Shapovalov, Jeremy Reizenstein, Ig- nacio Rocco, Natalia Neverova, Andrea Vedaldi, and David Novotn´y. Common pets in 3d: Dynamic new-view synthesis of real-life deformable categories. In CVPR, pages 4881–

  59. [66]

    A benchmark for the eval- uation of RGB-D SLAM systems

    J ¨urgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers. A benchmark for the eval- uation of RGB-D SLAM systems. In IROS, pages 573–580. IEEE, 2012. 6, 7, 8

  60. [67]

    Kalib: Markerless hand-eye calibration with keypoint track- ing

    Tutian Tang, Minghao Liu, Wenqiang Xu, and Cewu Lu. Kalib: Markerless hand-eye calibration with keypoint track- ing. CoRR, abs/2408.10562, 2024. 2

  61. [68]

    Raft: Recurrent all-pairs field transforms for optical flow

    Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part II 16, pages 402–419. Springer,

  62. [69]

    Bundle adjustment—a mod- 11 ern synthesis

    Bill Triggs, Philip F McLauchlan, Richard I Hartley, and Andrew W Fitzgibbon. Bundle adjustment—a mod- 11 ern synthesis. In Vision Algorithms: Theory and Prac- tice: International Workshop on Vision Algorithms Corfu, Greece, September 21–22, 1999 Proceedings , pages 298–

  63. [70]

    Marigold-dc: Zero-shot monocular depth completion with guided diffusion

    Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker, Konrad Schindler, and Anton Obukhov. Marigold-dc: Zero-shot monocular depth completion with guided diffusion. CoRR, abs/2412.13389, 2024. 3

  64. [71]

    Scenetracker: Long-term scene flow estimation network

    Bo Wang, Jian Li, Yang Yu, Li Liu, Zhenping Sun, and Dewen Hu. Scenetracker: Long-term scene flow estimation network. arXiv preprint arXiv:2403.19924, 2024. 3

  65. [72]

    Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment

    Jianyuan Wang, Christian Rupprecht, and David Novotny. Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 9773–9783,

  66. [73]

    Vggsfm: Visual geometry grounded deep structure from motion

    Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. Vggsfm: Visual geometry grounded deep structure from motion. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 21686–21697, 2024. 3

  67. [74]

    Vggt: Vi- sual geometry grounded transformer

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Vi- sual geometry grounded transformer. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 5294–5306, 2025. 2, 3, 6, 7, 8

  68. [75]

    Tracking everything everywhere all at once

    Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, and Noah Snavely. Tracking everything everywhere all at once. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 19795–19806, 2023. 2, 3

  69. [76]

    Shape of motion: 4d reconstruc- tion from a single video

    Qianqian Wang, Vickie Ye, Hang Gao, Jake Austin, Zhengqi Li, and Angjoo Kanazawa. Shape of motion: 4d reconstruc- tion from a single video. arXiv preprint arXiv:2407.13764,

  70. [77]

    Continuous 3d perception model with persistent state

    Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A Efros, and Angjoo Kanazawa. Continuous 3d perception model with persistent state. arXiv preprint arXiv:2501.12387, 2025. 6, 7, 8

  71. [79]

    Dust3r: Geometric 3d vi- sion made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20697– 20709, 2024. 6, 7, 8

  72. [80]

    Tartanair: A dataset to push the limits of visual slam

    Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Se- bastian Scherer. Tartanair: A dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4909–49...

  73. [81]

    Objctrl-2.5 d: Training-free object con- trol with camera poses

    Zhouxia Wang, Yushi Lan, Shangchen Zhou, and Chen Change Loy. Objctrl-2.5 d: Training-free object con- trol with camera poses. arXiv preprint arXiv:2412.07721 ,

  74. [82]

    Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation

    Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Jiawei Ren, Liang Pan, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, Dahua Lin, and Ziwei Liu. Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation. In CVPR, pages 803–814. IEEE, 2023. 5

  75. [83]

    RGBD objects in the wild: Scaling real-world 3d object learning from RGB-D videos

    Hongchi Xia, Yang Fu, Sifei Liu, and Xiaolong Wang. RGBD objects in the wild: Scaling real-world 3d object learning from RGB-D videos. In CVPR, pages 22378– 22389. IEEE, 2024. 5

  76. [84]

    Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes

    Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv preprint arXiv:1711.00199, 2017. 3

  77. [85]

    Level- S2fM: Structure From Motion on Neural Level Set of Im- plicit Surfaces

    Yuxi Xiao, Nan Xue, Tianfu Wu, and Gui-Song Xia. Level- S2fM: Structure From Motion on Neural Level Set of Im- plicit Surfaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17205– 17214, 2023. 3

  78. [86]

    Spatialtracker: Tracking any 2d pixels in 3d space

    Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20406–20417, 2024. 2, 3, 6

  79. [87]

    Depth anything: Unleashing the power of large-scale unlabeled data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10371–10381, 2024. 2, 3, 7

  80. [88]

    Depth any- thing v2

    Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2. Advances in Neural Information Processing Sys- tems, 37:21875–21911, 2025. 2, 3

  81. [89]

    Neural window fully-connected crfs for monocular depth estimation

    Weihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu, and Ping Tan. Neural window fully-connected crfs for monocular depth estimation. In CVPR, pages 3906–3915. IEEE, 2022. 3

  82. [90]

    Tapip3d: Tracking any point in persistent 3d geome- try

    Bowei Zhang, Lei Ke, Adam W Harley, and Katerina Fragki- adaki. Tapip3d: Tracking any point in persistent 3d geome- try. arXiv preprint arXiv:2504.14717, 2025. 3, 6

  83. [91]

    Monst3r: A simple approach for estimat- ing geometry in the presence of motion

    Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jam- pani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming- Hsuan Yang. Monst3r: A simple approach for estimat- ing geometry in the presence of motion. arXiv preprint arXiv:2410.03825, 2024. 6, 7, 8

  84. [92]

    Rel- pose: Predicting probabilistic relative rotation for single ob- jects in the wild

    Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani. Rel- pose: Predicting probabilistic relative rotation for single ob- jects in the wild. In ECCV, pages 592–611. Springer, 2022. 3

  85. [93]

    Cameras as rays: Pose estimation via ray diffusion

    Jason Y Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. arXiv preprint arXiv:2402.14817, 2024. 3

  86. [94]

    Zhoutong Zhang, Forrester Cole, Zhengqi Li, Michael Ru- binstein, Noah Snavely, and William T. Freeman. Struc- ture and motion from casual videos. In ECCV, pages 20–37. Springer, 2022. 8 12

  87. [95]

    Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild

    Wang Zhao, Shaohui Liu, Hengkai Guo, Wenping Wang, and Yong-Jin Liu. Particlesfm: Exploiting dense point trajecto- ries for localizing moving cameras in the wild. In ECCV. Springer, 2022. 2, 8

  88. [96]

    Pointodyssey: A large-scale synthetic dataset for long-term point tracking

    Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wet- zstein, and Leonidas J Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 19855–19865, 2023. 5, 9 13

  89. [2011]

    Computer Vision Foundation / IEEE Computer Soci- ety, 2018. 3

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.