Pith. sign in

REVIEW 2 major objections 6 minor 72 references

VidMap recovers metric camera poses and calibration from arbitrary long uncalibrated videos by fusing SLAM-style temporal trust with offline global SfM.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-07-30 13:49 UTC pith:XJ7Q7LGN

load-bearing objection Solid systems paper: provenance-aware tracks plus depth-regularized global SfM actually move the needle on long, messy, uncalibrated video. the 2 major comments →

arxiv 2607.27194 v1 pith:XJ7Q7LGN submitted 2026-07-29 cs.CV cs.RO

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

classification cs.CV cs.RO
keywords structure-from-motionvisual SLAMvideo reconstructionmetric monocular depthloop closuredense image matchingself-calibrationglobal positioning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Unconstrained video is the main way the world is recorded, yet turning it into accurate camera poses and calibration remains hard. Causal SLAM locks in early errors and often needs known intrinsics; classical SfM ignores time order and collapses under visual symmetries and extreme motion. VidMap keeps the temporal chain as a first-class signal—building long dense tracks, treating loop closures as soft provenance-labeled links, and injecting metric monocular depth into global positioning—while still solving everything offline with global optimization. On long phone, robot, and disaster-site videos the system is substantially more accurate and complete than both families of baselines, calibrated or not. If the claim holds, ordinary video becomes a scalable source of metric 3D training data for navigation and scene understanding.

Core claim

The paper shows that treating temporal order as provenance inside a global SfM pipeline—trusted sequential tracks, downweighted loop-closure edges, and per-image metric depth scales—yields metric reconstructions of long uncalibrated videos that are markedly more robust and accurate than state-of-the-art SLAM or SfM under extreme motion and visual aliasing.

What carries the argument

Provenance-aware mapping: sequential tracks are built by multi-flow dense matching and kept separate from loop-closure observations; rotation averaging and global positioning then apply tighter robust losses to sequential edges and softer losses to loop closures, while monocular depth priors with free per-image scales regularize degenerate geometry.

Load-bearing premise

Learned dense matching, retrieval, calibration, and monocular depth priors stay reliable enough that soft depth residuals and provenance losses can fix degenerate motion and reject false loop closures.

What would settle it

Run the full pipeline with depth priors and loop-closure edges ablated on the long LaMAR and CroCoDL robot sequences; if windowed translation AUC no longer beats strong SLAM and SfM baselines by a large margin, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X LinkedIn Reddit HN

If this is right

  • Ordinary phone and wearable video can be turned into metric posed training data without expert capture or known intrinsics.
  • Global SfM pipelines gain a practical way to use dense matchers and monocular depth without losing geometric precision on long sequences.
  • Causal SLAM’s structural failure modes—unrecoverable tracking loss and locked-in early drift—are avoidable when the same temporal cues are used non-causally.
  • Visual symmetries that break orderless SfM become manageable once sequential and loop-closure evidence are scored differently.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same provenance split could be ported to other global optimizers (incremental SfM, pose-graph SLAM) without adopting the full VidMap stack.
  • As monocular metric depth models improve on fisheye, grayscale, and disaster imagery, the calibrated–uncalibrated gap the paper reports should shrink further.
  • Failure modes listed by the authors—forward-motion under-keyframing and mutually consistent false loops—suggest a natural next test: active keyframe insertion driven by residual uncertainty rather than image displacement alone.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. VidMap proposes an offline video reconstruction pipeline that combines SLAM-style temporal structure with global SfM optimization. It selects motion-based keyframes, builds long sparse tracks via dense matching (RoMa) with multi-flow drift correction, separates sequential from loop-closure observations (non-transitive LC links), estimates or refines intrinsics, and runs provenance-aware rotation averaging plus global positioning and BA regularized by monocular metric depth with per-image scales (Eqs. 1–4). Evaluations on LaMAR, CroCoDL, ETH3D-SLAM, and EuRoC, in calibrated and uncalibrated settings, report large gains over classical SfM (GLOMAP variants), SLAM (DPV-SLAM, MegaSaM, ViPE, MASt3R-SLAM), and feedforward models (DA3-Long, LoGeR, etc.), with ablations (Table 4, Fig. 8), runtime analysis (Fig. 7), and public code.

Significance. If the empirical claims hold, this is a substantial systems contribution: metric poses and calibration from long, unconstrained, often uncalibrated video is a bottleneck for large-scale 3D and navigation data. The design choices—provenance-aware tracks/losses and depth-regularized global positioning—are concrete and well isolated. Strengths include multi-dataset evaluation with strong classical and learned baselines, fixed hyperparameters across datasets, calibrated/uncalibrated splits, ablations that attribute gains to depth and provenance, qualitative failure analysis, runtime/memory characterization, and a public implementation. The work is engineering-heavy rather than theoretically novel, but the integration is careful and the reported robustness gap on LaMAR/CroCoDL is large enough to matter in practice.

major comments (2)
  1. [Abstract; §1; §A.1] Abstract and §1 claim reconstruction of “arbitrary, long, uncalibrated videos,” while §A.1 lists load-bearing failure modes (forward-motion under-keyframing, mutually consistent false LCs overwhelming Cauchy downweighting, track loss, OOD depth/calibration priors, long-range drift without LC). Tables 1–2 support strong average superiority, but the central claim should be scoped to the evaluated regimes. Please either (i) qualify “arbitrary” in the abstract/intro to match the limitations, or (ii) add a short failure-rate breakdown (e.g., fraction of sequences with catastrophic collapse vs. baselines on LaMAR/CroCoDL) so readers can judge coverage of the stated operating envelope.
  2. [§4.3–4.4; Table 4; Fig. 8] Table 4 shows depth-in-GP and provenance losses are critical on LaMAR, and no-LC hurts ETH3D; Fig. 8 is qualitative. For the systems claim that the *interplay* of video-aware extraction and depth-augmented global optimization closes the gap (§4.3), a per-scene or per-failure-mode split (symmetry-heavy vs. degenerate-motion sequences) would better show that provenance and depth address distinct failure modes rather than correlated gains. This need not be exhaustive, but one stratified table or appendix breakdown on LaMAR would make the causal attribution load-bearing rather than aggregate-only.
minor comments (6)
  1. [§3.4] Eq. (1)–(3): define Σ^v_ik and the diagonal approximation earlier in the main text or point more explicitly to the appendix derivation; the main text defers details that affect how depth and bearing terms are weighted.
  2. [§4.2] W-AUC protocol (5% of window length) is justified in the supplement; a one-sentence reminder in §4.2 would help readers interpret Tables 1–2 without leaving the main paper.
  3. [Fig. 7] Fig. 7 runtime comparison mixes different keyframing densities; labeling approximate keyframe counts or FPS-normalized cost next to each method would reduce ambiguity when comparing to COLMAP/GLOMAP/DA3-Long.
  4. [Throughout] Typos/style: “Patakiet al.” spacing in running heads; occasional missing spaces in compound phrases (“video-awaretracking”, “toglobalSfM”); ensure consistent notation for keyframes K_i vs. indices.
  5. [§2] Related work could briefly situate Doppelgangers/Doppelgangers++ relative to provenance-aware LC rejection, since visual aliasing is a core motivation.
  6. [Table 3; §4.3] EuRoC uncalibrated results (Table 3) show a large calibrated–uncalibrated gap; a short note on fisheye/grayscale prior mismatch would clarify that this is expected rather than a silent regression.

Circularity Check

0 steps flagged

No significant circularity: empirical SfM/SLAM systems paper judged against external GT, not self-defined targets.

full rationale

VidMap is an engineering systems paper. Its central claim is empirical superiority on LaMAR, CroCoDL, ETH3D-SLAM, and EuRoC under fixed hyperparameters, measured by translation AUC/W-AUC after alignment to external ground truth (LiDAR-aligned SfM, motion capture, Vicon). The optimization objectives (Eqs. 1–4: bearing, depth, scale, and reprojection residuals with provenance-dependent robust losses) are standard MAP-style losses; they do not define the reported metrics by construction, and ablations (Table 4) isolate components against the same external GT. Self-citations (GLOMAP, MP-SfM, GeoCalib, RoMa line) supply reusable front-end and backend blocks with independent public artifacts; none is invoked as a uniqueness theorem that forces the result. Depth scales and loss annealing are ordinary fitted engineering knobs, not renamed predictions of the evaluation metric. No step reduces a claimed derivation or prediction to its own inputs. Score 0 is appropriate.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

Load-bearing content is engineering assumptions and off-the-shelf learned modules, not new physical entities. The claim rests on standard multi-view geometry, fixed hyperparameters, and the reliability of external networks (RoMa, MegaLoc, monocular depth, GeoCalib) plus the modeling choice that sequential observations deserve tighter robust losses than loop closures.

free parameters (5)
  • Keyframe motion/coverage thresholds (τc, τ0, τf) and track budget N = τc=0.01, τ0=0.12, τf=0.4, N=1200
    Hand-chosen triggers for keyframe insertion and track density; fixed across datasets after validation tuning (Appendix E.2 / Table 6).
  • Multi-flow acceptance tolerance τ and sliding window W = W=9 (stride 2); τ fixed tolerance factor
    Controls when non-adjacent dense predictions replace sequential propagation; structural to track quality.
  • Provenance-dependent robust loss schedules and depth Cauchy annealing
    Huber vs Cauchy by edge type, graduated non-convexity, depth-ratio outlier flags; chosen to stabilize GP/BA rather than derived.
  • Per-image depth scale prior uncertainty σs,i and monocular depth uncertainties σik
    Soft metric anchoring strength in Eqs. 1 and 4; affects long-horizon drift behavior in ablations.
  • Retrieval top-k and similarity threshold for loop closure
    Determines which non-sequential pairs enter as LC edges; directly affects aliasing exposure.
axioms (5)
  • standard math Multi-view geometry: relative rotations/poses from correspondences (and optional depth), rotation averaging, bearing-based global positioning, and bundle adjustment recover metric structure up to the usual gauge when constraints are sufficient.
    Core of Sec. 3.4 building on GLOMAP; standard SfM theory.
  • domain assumption Consecutive video frames are locally unambiguous enough that dense matching plus temporal chaining yields trustworthy sequential tracks more often than orderless matching.
    Stated in Sec. 3.1 as the key structural property enabling video-aware extraction.
  • domain assumption Incorrect loop closures can be treated as statistical outliers via heavier-tailed losses while sequential edges remain inliers under the same optimization.
    Provenance-aware RA/GP design (Sec. 3.2–3.4, Figs. 4–5); fails if many mutually consistent false LCs dominate (Limitations).
  • domain assumption Off-the-shelf monocular metric depth and calibration networks provide useful soft constraints after per-image scale (and shared focal) refinement.
    Depth terms in GP/BA and GeoCalib/view-graph calibration path; ablations show large drops without them on LaMAR.
  • ad hoc to paper Additive independent displacement error model for chaining match covariances along tracks.
    Sec. 3.1 covariance accumulation Σseq; convenient noise model, not independently validated as the true error process.
invented entities (2)
  • Provenance-aware sequential tracks with soft loop-closure observations (non-transitive LC links) independent evidence
    purpose: Keep sequential chains intact under wrong LC so robust losses can downweight aliasing without collapsing tracks.
    Central systems invention vs standard transitive track merging in GLOMAP-style SfM (Fig. 4).
  • VidMap end-to-end pipeline (video extraction + depth-regularized provenance-aware global mapping) independent evidence
    purpose: Package temporal extraction and global SfM extensions into one metric video reconstruction system.
    Named system contribution; evidence is empirical benchmarks and public code, not a new physical object.

reviewed 2026-07-30 · how reviews work

0 comments
Cite this review

Pith. "Pith review of VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion." pith.science (2026). https://pith.science/paper/XJ7Q7LGN

@misc{pith2026260727194,
  author       = {Pith},
  title        = {Pith review of: VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XJ7Q7LGN}},
  note         = {Machine review of arXiv:2607.27194}
}
Share X LinkedIn Reddit HN
read the original abstract

Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos. This system leverages recent advances in wide-baseline dense image matching, treats temporal ordering as a first-class citizen for reliable loop closure, and augments global optimization with metric monocular depth priors. As a result, thorough evaluations on diverse, challenging datasets that exhibit extreme motion and visual symmetries reveal that our approach is significantly more robust and accurate than both state-of-the-art SLAM and SfM, classical or learned, with given or unknown camera calibration. The code is publicly available at https://github.com/cvg/vidmap.

Figures

Figures reproduced from arXiv: 2607.27194 by Marc Pollefeys, Paul-Edouard Sarlin, Zador Pataki.

Figure 1
Figure 1. Figure 1: VidMap reconstructs long, unconstrained videos with complex mo [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: System overview. We select keyframes using optical flow and extract monoc￾ular depth for each keyframe. We sample and propagate sparse tracks across keyframes using dense matching and across loop-closures detected using image retrieval. We esti￾mate initial camera intrinsics using view-graph calibration and monocular calibration priors. Mapping preserves provenance for pairwise relative poses and track obs… view at source ↗
Figure 3
Figure 3. Figure 3: Multi-flow drift correction. For each new keyframe Ki, sequential tracks are propagated through the latest match Ki−1 → Ki, yielding prediction xˆ seq i with accumulated covariance Σ seq i . To reduce drift, VidMap also evaluates direct predictions from earlier anchors Ki−n→Ki, yielding xˆ (i−n) i with covariance Σ (i−n) i . We select the direct candidate with minimum covariance trace and accept it only if… view at source ↗
Figure 4
Figure 4. Figure 4: Provenance Aware Track Establishment. Left: VidMap constructs tracks from sequential matching and treats loop closure (LC) as a soft link between sequential tracks, which remain reliable when the LC is incorrect. Right: GLOMAP merges LC matches by transitivity, so provenance is lost and an incorrect closure can combine separate sequential chains into an inseparable track. incorrect loop closure reference r… view at source ↗
Figure 5
Figure 5. Figure 5: Provenance Aware Rotation Averaging and Global Positioning. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Qualitative comparison to existing approaches on LaMAR (top) [PITH_FULL_IMAGE:figures/full_fig_p014_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Analysis of the runtime. Top left: Duration of each component with increas￾ing number of keyframes. Matching is the most expensive step. Top right: Comparison of the total runtime against existing approaches. Bottom left: Impact of keyframing on the speed and accuracy. Motion-aware keyframing retains the accuracy of the dense keyframing but at lower cost, while aggressive keyframing reduces runtime at the … view at source ↗
Figure 8
Figure 8. Figure 8: Visual ablations. Each column shows trajectory images above the correspond￾ing reconstructions for VidMap (top) and the ablation (bottom); GT trajectories are yellow and estimates are blue. Left: red loop-closure edges are rejected as outliers while green edges are kept as inliers. Without provenance-aware losses, VidMap cannot down￾weight symmetries and the reconstruction collapses. Right: without depth p… view at source ↗
Figure 9
Figure 9. Figure 9: Geometric priors resolve degenerate camera motion. [PITH_FULL_IMAGE:figures/full_fig_p017_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Establishing tracks through textureless regions. [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Difficult transition. The camera moves through a doorway into a corridor, combining pure rotation with short tracks. Left: our approach recovers the trajectory. Middle: without geometric priors, the degenerate segment causes loss of tracking and collapse of the scale. Right: VidMap-ALIKED+LG creates two disjoint reconstructions. Both components are needed: geometric priors alone cannot compensate for miss… view at source ↗
Figure 12
Figure 12. Figure 12: Challenges of local symmetries. For three sequences, we show represen￾tative images that exhibit symmetries and visualize the camera poses and 3D point cloud estimated by standard SfM (left) and our approach (right). SfM treats all corre￾spondences equally via transitivity and does not distinguish sequential and loop closure observations. Incorrect loop closures, visualized as red edges in the view graph,… view at source ↗
Figure 13
Figure 13. Figure 13: Out-of-domain generalization on CroCoDL. [PITH_FULL_IMAGE:figures/full_fig_p020_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Qualitative comparison on LaMAR. Three scenes (CAB, HGE, LIN) with sample frames and bird’s-eye trajectory plots (left) and 3D reconstructions with overlaid trajectories (right; red: GT, blue: estimated). VidMap maintains consistent scale and topology; DA3-Long and ViPE exhibit progressive drift due to accumulated errors on longer sequences. lack repeatability on textureless regions, producing few or no m… view at source ↗
Figure 15
Figure 15. Figure 15: Qualitative comparison on CroCoDL. Two disaster-response scenes (iOS phone and quadruped robot). Bird’s-eye trajectory plots (left) comparing all baselines; 3D reconstructions from top and side views (right; red: GT, blue: estimated). VidMap generalizes to these out-of-distribution environments while baselines exhibit significant drift or fail to reconstruct. 1 2 3 4 5 Hop number 60 65 70 75 80 85 Recall@… view at source ↗
Figure 16
Figure 16. Figure 16: Track quality across keyframe hops. Recall@5px on ETH3D SLAM (left, by hop count) and LaMAR (right, by GT keyframe distance). RoMa already outper￾forms all flow-based trackers when used as a simple chained matcher; multi-flow prop￾agation further widens the gap, confirming that the correction step contributes beyond the choice of backbone. matched with LightGlue on consecutive pairs; tracks are chained th… view at source ↗
Figure 17
Figure 17. Figure 17: Epipolar error CDF per dataset. Cumulative distribution of epipolar error for ALIKED keypoints. Top: ETH3D SLAM, by keyframe hop count. Bottom: LaMAR, by GT keyframe distance bin (Dist 1–3, increasing displacement). All methods perform comparably at short range; as hops or distance increase, flow-based trackers shift rightward as drift compounds, while RoMa with multi-flow maintains the tightest distribut… view at source ↗
Figure 18
Figure 18. Figure 18: Tracker backbone comparison on LaMAR. Median WTE (m) averaged across scenes at increasing trajectory-length windows. Replacing RoMa with flow-based trackers degrades reconstruction quality at all scales. VidMap-ALIKED+LG, relying on keypoint repeatability, loses coverage on textureless regions where dense propagation maintains correspondences. LaMAR [47]. Three scenes in Zurich (CAB, HGE: indoor + outdoor… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

72 extracted references · 8 linked inside Pith

  1. [1]

    In: CVPRW (2025)

    Berton, G., Masone, C.: MegaLoc: One Retrieval to Place Them All. In: CVPRW (2025)

  2. [2]

    In: CVPR (2025)

    Blum, H., Mercurio, A., O’Reilly, J., Engelbracht, T., Dusmanu, M., Pollefeys, M., Bauer, Z.: CroCoDL: Cross-device Collaborative Dataset for Localization. In: CVPR (2025)

  3. [3]

    In: ICLR (2025)

    Bochkovskiy, A., Delaunoy, A., Germain, H., Santos, M., Zhou, Y., Richter, S., Koltun, V.: Depth Pro: Sharp Monocular Metric Depth in Less Than a Second. In: ICLR (2025)

  4. [4]

    IJRR (2016)

    Burri, M., Nikolic, J., Gohl, P., Schneider, T., Rehder, J., Omari, S., Achtelik, M.W., Siegwart, R.: The EuRoC micro aerial vehicle datasets. IJRR (2016)

  5. [5]

    In: ICCV (2023)

    Cai, R., Tung, J., Wang, Q., Averbuch-Elor, H., Hariharan, B., Snavely, N.: Doppel- gangers: Learning to Disambiguate Images of Similar Structures. In: ICCV (2023)

  6. [6]

    IEEE TRO (2021)

    Campos, C., Elvira, R., Rodríguez, J.J.G., Montiel, J.M.M., Tardós, J.D.: ORB- SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial, and Mul- timap SLAM. IEEE TRO (2021)

  7. [7]

    arXiv preprint arXiv:2604.14141 (2026)

    Chen, L.Z., Gao, J., Chen, Y., Cheng, K.L., Sun, Y., Hu, L., Xue, N., Zhu, X., Shen, Y., Yao, Y., Xu, Y.: Geometric Context Transformer for Streaming 3D Re- construction. arXiv preprint arXiv:2604.14141 (2026)

  8. [8]

    In: ICLR (2026)

    Chen, X., Chen, Y., Xiu, Y., Geiger, A., Chen, A.: TTT3R: 3D Reconstruction as Test-Time Training. In: ICLR (2026)

  9. [9]

    In: CVPR (2011)

    Crandall, D., Owens, A., Snavely, N., Huttenlocher, D.: Discrete-Continuous Op- timization for Large-Scale Structure from Motion. In: CVPR (2011)

  10. [10]

    IEEE RAL5(2), 721–728 (2020)

    Czarnowski, J., Laidlow, T., Clark, R., Davison, A.J.: DeepFactors: Real-Time Probabilistic Dense Monocular SLAM. IEEE RAL5(2), 721–728 (2020)

  11. [11]

    IEEE TPAMI (2007) VidMap 29

    Davison, A.J., Reid, I.D., Molton, N.D., Stasse, O.: MonoSLAM: Real-Time Single Camera SLAM. IEEE TPAMI (2007) VidMap 29

  12. [12]

    arXiv:2507.16443 (2025)

    Deng, K., Ti, Z., Xu, J., Yang, J., Xie, J.: VGGT-Long: Chunk it, loop it, align it – Pushing VGGT’s Limits on Kilometer-scale Long RGB Sequences. arXiv:2507.16443 (2025)

  13. [13]

    In: CVPRW (2018)

    DeTone, D., Malisiewicz, T., Rabinovich, A.: SuperPoint: Self-supervised interest point detection and description. In: CVPRW (2018)

  14. [14]

    In: ECCV (2024)

    Dexheimer, E., Davison, A.J.: COMO: Compact Mapping and Odometry. In: ECCV (2024)

  15. [15]

    In: ICCV (2025)

    Ding, Y., Kocur, V., Vávra, V., Berger Haladóvá, Z., Yang, J., Sattler, T., Kukelova, Z.: RePoseD: Efficient Relative Pose Estimation With Known Depth Information. In: ICCV (2025)

  16. [16]

    In: 3DV (2025)

    Duisterhof, B.P., Zust, L., Weinzaepfel, P., Leroy, V., Cabon, Y., Revaud, J.: MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion. In: 3DV (2025)

  17. [17]

    In: CVPR (2023)

    Edstedt, J., Athanasiadis, I., Wadenbäck, M., Felsberg, M.: DKM: Dense Kernel- ized Feature Matching for Geometry Estimation. In: CVPR (2023)

  18. [18]

    arXiv:2511.15706 (2025)

    Edstedt, J., Nordström, D., Zhang, Y., Bökman, G., Astermark, J., Larsson, V., Heyden,A.,Kahl,F.,Wadenbäck,M.,Felsberg,M.:RoMav2:HarderBetterFaster Denser Feature Matching. arXiv:2511.15706 (2025)

  19. [19]

    In: CVPR (2024)

    Edstedt, J., Sun, Q., Bökman, G., Wadenbäck, M., Felsberg, M.: RoMa: Robust Dense Feature Matching. In: CVPR (2024)

  20. [20]

    IEEE TPAMI (2018)

    Engel, J., Koltun, V., Cremers, D.: Direct sparse odometry. IEEE TPAMI (2018)

  21. [21]

    In: ECCV (2014)

    Engel, J., Schöps, T., Cremers, D.: LSD-SLAM: Large-Scale Direct Monocular SLAM. In: ECCV (2014)

  22. [22]

    In: ICCV (2023)

    Hagemann, A., Knorr, M., Stiller, C.: Deep Geometry-Aware Camera Self- Calibration from Video. In: ICCV (2023)

  23. [23]

    In: ICCV (2025)

    Harley, A.W., You, Y., Sun, X., Zheng, Y., Raghuraman, N., Gu, Y., Liang, S., Chu, W.H., Dave, A., You, S., Ambrus, R., Fragkiadaki, K., Guibas, L.: AllTracker: Efficient Dense Point Tracking at High Resolution. In: ICCV (2025)

  24. [24]

    IEEE TPAMI (2024)

    Hu, M., Yin, W., Zhang, C., Cai, Z., Long, X., Wang, K., Chen, H., Yu, G., Shen, C., Shen, S.: Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation. IEEE TPAMI (2024)

  25. [25]

    arXiv:2508.10934 (2025)

    Huang, J., Zhou, Q., Rabeti, H., Korovko, A., Ling, H., Ren, X., Shen, T., Gao, J., Slepichev, D., Lin, C.H., Ren, J., Xie, K., Biswas, J., Leal-Taixé, L., Fidler, S.: ViPE: Video Pose Engine for 3D Geometric Perception. arXiv:2508.10934 (2025)

  26. [26]

    In: ECCV (2024)

    Karaev, N., Rocco, I., Graham, B., Neverova, N., Vedaldi, A., Rupprecht, C.: CoTracker: It is better to track together. In: ECCV (2024)

  27. [27]

    In: IEEE Int

    Klein, G., Murray, D.: Parallel tracking and mapping for small AR workspaces. In: IEEE Int. Symp. Mixed and Augmented Reality. pp. 225–234 (2007)

  28. [28]

    In: ECCV (2024)

    Leroy, V., Cabon, Y., Revaud, J.: Grounding Image Matching in 3D with MASt3R. In: ECCV (2024)

  29. [29]

    arXiv preprint arXiv:2202.09199 (2022)

    Leutenegger, S.: OKVIS2: Realtime scalable visual-inertial SLAM with loop clo- sure. arXiv preprint arXiv:2202.09199 (2022)

  30. [30]

    In: CVPR (2026)

    Li, M., Zhu, Z., Pollefeys, M., Barath, D.: DROID-SLAM in the wild. In: CVPR (2026)

  31. [31]

    In: CVPR (2025)

    Li, Z., Tucker, R., Cole, F., Wang, Q., Jin, L., Ye, V., Kanazawa, A., Holynski, A., Snavely, N.: MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos. In: CVPR (2025)

  32. [32]

    arXiv preprint arXiv:2511.10647 (2025) 30 Patakiet al

    Lin, H., Chen, S., Liew, J., Chen, D.Y., Li, Z., Shi, G., Feng, J., Kang, B.: Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647 (2025) 30 Patakiet al

  33. [33]

    In: ICCV (2023)

    Lindenberger, P., Sarlin, P.E., Pollefeys, M.: LightGlue: Local Feature Matching at Light Speed. In: ICCV (2023)

  34. [34]

    In: ECCV (2024)

    Lipson, L., Teed, Z., Deng, J.: Deep patch visual SLAM. In: ECCV (2024)

  35. [35]

    In: ECCV (2024)

    Liu, S., Gao, Y., Zhang, T., Pautrat, R., Schönberger, J.L., Larsson, V., Pollefeys, M.: Robust incremental structure-from-motion with hybrid features. In: ECCV (2024)

  36. [36]

    In: CVPR (2022)

    Liu, S., Nie, X., Hamid, R.: Depth-Guided Sparse Structure-from-Motion for Movies and TV Shows. In: CVPR (2022)

  37. [37]

    In: NeurIPS (2025)

    Maggio, D., Lim, H., Carlone, L.: VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold. In: NeurIPS (2025)

  38. [38]

    In: ICCV (2013)

    Moulon, P., Monasse, P., Marlet, R.: Global fusion of relative motions for robust, accurate and scalable structure from motion. In: ICCV (2013)

  39. [39]

    IEEE TRO (2015)

    Mur-Artal, R., Montiel, J.M.M., Tardós, J.D.: ORB-SLAM: A Versatile and Ac- curate Monocular SLAM System. IEEE TRO (2015)

  40. [40]

    In: CVPR (2025)

    Murai,R.,Dexheimer,E.,Davison,A.J.:MASt3R-SLAM:Real-TimeDenseSLAM with 3D Reconstruction Priors. In: CVPR (2025)

  41. [41]

    In: WACV (2024)

    Neoral, M., Šerých, J., Matas, J.: MFT: Long-Term Tracking of Every Pixel. In: WACV (2024)

  42. [42]

    In: ECCV (2024)

    Pan, L., Barath, D., Pollefeys, M., Schönberger, J.L.: Global structure-from-motion revisited. In: ECCV (2024)

  43. [43]

    In: CVPR (2025)

    Pataki, Z., Sarlin, P.E., Schönberger, J.L., Pollefeys, M.: MP-SfM: Monocular Sur- face Priors for Robust Structure-from-Motion. In: CVPR (2025)

  44. [44]

    In: CVPR (2024)

    Piccinelli, L., Yang, Y.H., Sakaridis, C., Segu, M., Li, S., Van Gool, L., Yu, F.: UniDepth: Universal Monocular Metric Depth Estimation. In: CVPR (2024)

  45. [45]

    IJCV (2004)

    Pollefeys, M., Van Gool, L., Vergauwen, M., Verbiest, F., Cornelis, K., Tops, J., Koch, R.: Visual modeling with a hand-held camera. IJCV (2004)

  46. [46]

    In: CVPR (2020)

    Sarlin, P.E., DeTone, D., Malisiewicz, T., Rabinovich, A.: SuperGlue: Learning Feature Matching with Graph Neural Networks. In: CVPR (2020)

  47. [47]

    In: ECCV (2022)

    Sarlin, P.E., Dusmanu, M., Schönberger, J.L., Speciale, P., Gruber, L., Larsson, V., Miksik, O., Pollefeys, M.: LaMAR: Benchmarking Localization and Mapping for Augmented Reality. In: ECCV (2022)

  48. [48]

    In: CVPR (2016)

    Schönberger, J.L., Frahm, J.M.: Structure-from-Motion Revisited. In: CVPR (2016)

  49. [49]

    In: CVPR (2019)

    Schöps, T., Sattler, T., Pollefeys, M.: BAD SLAM: Bundle adjusted direct RGB-D SLAM. In: CVPR (2019)

  50. [50]

    In: 3DV (2025)

    Smith, C., Charatan, D., Tewari, A., Sitzmann, V.: FlowMap: High-Quality Cam- era Poses, Intrinsics, and Depth via Gradient Descent. In: 3DV (2025)

  51. [51]

    In: ACM SIGGRAPH (2006)

    Snavely, N., Seitz, S.M., Szeliski, R.: Photo tourism: Exploring photo collections in 3D. In: ACM SIGGRAPH (2006)

  52. [52]

    In: CVPR (2021)

    Sun, J., Shen, Z., Wang, Y., Bao, H., Zhou, X.: LoFTR: Detector-Free Local Fea- ture Matching With Transformers. In: CVPR (2021)

  53. [53]

    In: ICCV (2015)

    Sweeney, C., Sattler, T., Hollerer, T., Turk, M., Pollefeys, M.: Optimizing the Viewing Graph for Structure-from-Motion. In: ICCV (2015)

  54. [54]

    In: ECCV (2020)

    Teed, Z., Deng, J.: RAFT: Recurrent all-pairs field transforms for optical flow. In: ECCV (2020)

  55. [55]

    In: NeurIPS (2021)

    Teed, Z., Deng, J.: DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. In: NeurIPS (2021)

  56. [56]

    In: NeurIPS (2023)

    Teed, Z., Lipson, L., Deng, J.: Deep patch visual odometry. In: NeurIPS (2023)

  57. [57]

    In: ECCV (2024) VidMap 31

    Veicht, A., Sarlin, P.E., Lindenberger, P., Pollefeys, M.: GeoCalib: Learning Single- image Calibration with Geometric Optimization. In: ECCV (2024) VidMap 31

  58. [58]

    In: CVPR (2025)

    Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: VGGT: Visual geometry grounded transformer. In: CVPR (2025)

  59. [59]

    In: CVPR (2024)

    Wang, J., Karaev, N., Rupprecht, C., Novotny, D.: VGGSfM: Visual geometry grounded deep structure from motion. In: CVPR (2024)

  60. [60]

    In: CVPR (2025)

    Wang, Q., Zhang, Y., Holynski, A., Efros, A.A., Kanazawa, A.: Continuous 3D Perception Model with Persistent State. In: CVPR (2025)

  61. [61]

    In: CVPR (2025)

    Wang,R.,Xu,S.,Dai,C.,Xiang,J.,Deng,Y.,Tong,X.,Yang,J.:MoGe:Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision. In: CVPR (2025)

  62. [62]

    In: CVPR (2024)

    Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., Revaud, J.: DUSt3R: Geometric 3D vision made easy. In: CVPR (2024)

  63. [63]

    arXiv preprint arXiv:2507.13347 (2025)

    Wang, Y., Zhou, J., Zhu, H., Chang, W., Zhou, Y., Li, Z., Chen, J., Pang, J., Shen, C., He, T.:π 3: Permutation-equivariant visual geometry learning. arXiv preprint arXiv:2507.13347 (2025)

  64. [64]

    In: ECCV (2014)

    Wilson, K., Snavely, N.: Robust global translations with 1DSfM. In: ECCV (2014)

  65. [65]

    In: CVPR (2025)

    Xiangli, Y., Cai, R., Chen, H., Byrne, J., Snavely, N.: Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features. In: CVPR (2025)

  66. [66]

    IEEE RAL5(2), 1127–1134 (2020)

    Yang, H., Antonante, P., Tzoumas, V., Carlone, L.: Graduated non-convexity for robust spatial perception: From non-minimal solvers to global outlier rejection. IEEE RAL5(2), 1127–1134 (2020)

  67. [67]

    In: CVPR (2025)

    Yang, J., Sax, A., Liang, K.J., Henaff, M., Tang, H., Cao, A., Chai, J., Meier, F., Feiszli, M.: Fast3R: Towards 3D reconstruction of 1000+ images in one forward pass. In: CVPR (2025)

  68. [68]

    arXiv preprint arXiv:2603.03269 (2026)

    Zhang, J., Herrmann, C., Hur, J., Sun, C., Yang, M.H., Cole, F., Darrell, T., Sun, D.: LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory. arXiv preprint arXiv:2603.03269 (2026)

  69. [69]

    In: ICCV

    Zhang, Y., Tosi, F., Mattoccia, S., Poggi, M.: GO-SLAM: Global optimization for consistent 3D instant reconstruction. In: ICCV. pp. 3727–3737 (2023)

  70. [70]

    In: ECCV (2022)

    Zhao, W., Liu, S., Guo, H., Wang, W., Liu, Y.J.: ParticleSfM: Exploiting dense point trajectories for localizing moving cameras in the wild. In: ECCV (2022)

  71. [71]

    IEEE Transactions on Instrumentation and Measurement72, 1–16 (2023)

    Zhao, X., Wu, X., Chen, W., Chen, P.C.Y., Xu, Q., Li, Z.: ALIKED: A lighter keypoint and descriptor extraction network via deformable transformation. IEEE Transactions on Instrumentation and Measurement72, 1–16 (2023)

  72. [72]

    In: CVPR (2018)

    Zhuang, B., Cheong, L.F., Lee, G.H.: Baseline Desensitizing in Translation Aver- aging. In: CVPR (2018)

This paper was first reviewed by grok-4.5 on July 30, 2026.