REVIEW 3 major objections 5 minor 1 cited by
MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Monocular depth and normal priors let incremental SfM reconstruct scenes from two-view tracks alone.
desk verdict MP-SfM is a real advance in incremental SfM for low-overlap scenes; the main soft spot—untested robustness to non-scale depth bias—is worth a revision experiment but not a rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an uncertainty-weighted fusion of single-view and multi-view geometric constraints, solved by alternating optimization. The objective $C_{BA} + C_{reg} + C_{int}$ couples (i) standard bundle adjustment over sparse 3D points, (ii) a depth-regularization term $C_{reg}$ pulling scene points toward refined per-image depth maps $D^*_i$, and (iii) a bilateral normal-integration term $C_{int}$ that conditions $D^*_i$ on the monocular depth and normal priors with Mahalanobis weighting by predicted covariances. Because the Hessian of the joint cost loses the block-diagonal structure needed for Schur-complement elimination, the authors alternate: refine each depth map independently with $C_{reg}+C_{int}$ at fixed poses and points, then optimize poses and points with $C_{BA}+C_{reg}$ at fixed depth maps. A final dense depth-consistency check, comparing each image's refined depth against a reprojected min-depth buffer of overlapping views, rejects images that contradict free space.
What would settle it
Take a scene with strictly zero three-view overlap and run the full pipeline with the monocular depth term disabled (the paper's 'no lifting' ablation); the reported minimal-overlap AUC at 5 degrees falls from about 56 to 16 on the indoor benchmark, while the pose error with priors enabled is close to ground truth. A reader can reproduce that contrast on a held-out set of such triplets: if accurate poses persist without the depth prior, the three-view requirement was never the bottleneck; if they collapse, the two-view claim rests on the priors, exactly as stated.
Extended reading notes
Core claim
The paper's central claim is that the scale information incremental SfM normally obtains from three-view tracks can instead be supplied by per-image monocular priors, up to one unknown scale per image. Each view's predicted depth and surface normals act as soft constraints: 3D points lifted from a single view serve as 2D–3D correspondences for pose estimation, and a combined objective $C_{BA}+C_{reg}+C_{int}$ couples sparse bundle adjustment with depth regularization and bilateral normal integration. The predicted uncertainties are propagated and calibrated, so bad depth estimates are down-weighted rather than trusted. A dense forward–backward depth-consistency check then de-registers any image whose refined depth contradicts overlapping views, removing symmetry-induced false positives. On low-overlap subsets of standard benchmarks, the authors report accurate pose estimates for scenes with zero three-view overlap where existing incremental, structure-less, and learned two-view pipelines fail, and they maintain competitive accuracy in dense high-overlap settings; they further state that this makes the approach the first to reliably reconstruct challenging indoor scenes from few images.
Load-bearing premise
The method's load-bearing premise is that monocular depth and normal predictions are accurate enough up to a per-image scale that their uncertainty-weighted fusion improves multi-view geometry; if the predicted uncertainties are systematically overconfident or miscalibrated, the joint optimization can pull poses toward wrong depth rather than toward consistent multi-view structure.
Editorial extensions
If this is right
- A reconstruction can be built from image pairs with no triple overlap, so sparse casual captures no longer need careful planning to guarantee three-view coverage.
- Dense two-view correspondences in texture-poor regions become usable directly, improving completeness where sparse keypoints are scarce.
- Because the priors are treated as uncertain soft constraints, swapping in a different monocular depth or normal estimator requires little retuning, so future improvements in single-image geometry transfer to SfM.
- Symmetry-induced wrong registrations can be detected and removed by dense depth consistency, even when sparse geometric verification accepts them.
- In low-parallax configurations, incremental reconstruction approaches the accuracy of global methods, which previously did not suffer from the same failure.
Reading between the lines
- Beyond the paper, the same uncertainty-weighted fusion should apply to camera relocalization and dense SLAM, where pure rotation, low texture, or repetitive structure make epipolar constraints degenerate; a monocular depth prior could supply scale and depth hypotheses there too.
- The per-image scale formulation implies that a monocular model with reliable relative depth but no metric scale could be substituted if its scale is recovered from the first verified two-view pair; that would decouple the method from metric depth models.
- The depth-consistency check could be inverted into an active-capture signal: images that repeatedly fail it flag symmetric or ambiguous regions, telling a non-expert user exactly which additional views would disambiguate the scene.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MP-SfM, an incremental Structure-from-Motion pipeline that augments COLMAP with monocular depth and surface-normal priors, including their predicted uncertainties. The key idea is to lift the classical requirement for three-view tracks: depth-lifted 2D-3D correspondences allow registration of new views with only two-view overlap, while an alternating optimization of bundle adjustment, depth refinement, and normal integration fuses the priors into the reconstruction. A dense depth-consistency check rejects incorrectly registered views, particularly in symmetric scenes. The method is evaluated on ETH3D, SMERF, Tanks and Temples, and RealEstate10k under varying overlap and parallax conditions, showing consistent improvements over COLMAP, GLOMAP, SLR, DF-SfM, VGGSfM, StudioSfM, and MASt3R-SfM. The paper claims that this is the first approach capable of reliably reconstructing challenging indoor environments from few images, while requiring little tuning.
Significance. If the results hold, this is a meaningful advance for incremental SfM: it directly attacks the three-view-track requirement, a known practical bottleneck for non-expert capture, and demonstrates large gains on low-overlap and low-parallax benchmarks. The evaluation is commendably broad: external benchmarks, multiple sparse and dense matchers, several monocular depth models, and component-wise ablations. The public code release is a concrete strength that supports reproducibility. The main robustness claim is credible, but the evidence does not yet cover all the error modes that the paper claims to handle: the per-image scale correction assumes scale-only prior errors, and the calibrated uncertainty machinery relies on several unreported tuning constants. These gaps are fixable but should be addressed before publication.
major comments (3)
- [Eq. (1) and Sec. 3.1] The median-ratio scaling in Eq. (1) corrects the monocular depth prior only up to a single per-image scale factor. The paper's central claim of robustness to errors in the priors (Abstract; Sec. 1) therefore depends on the unstated assumption that prior depth errors are predominantly scale-only. This assumption is not tested: no experiment perturbs the priors with additive offsets, depth-dependent scale drift, or spatially varying bias. The ground-truth-depth ablation in Table 5 (ETH3D minimal overlap: AUC@1° improves from 27.3 to 42.9) shows that prior bias, not just its scale, limits fine-grained accuracy. Please add a synthetic-bias ablation or explicitly scope the robustness claim to scale-correct priors.
- [Sec. 3.4, Eq. (6)] The depth consistency check is a central safeguard against symmetry failures, yet the two decision parameters—gamma in Eq. (6) and the ratio beta_hat mentioned in the text—are never given numeric values. Since Table 6 shows this check is crucial in the SMERF scenes, leaving these thresholds unreported prevents reproduction and makes it impossible to judge how much tuning the method requires.
- [Appendix C and Sec. 4.3] The uncertainty calibration and robust-loss configuration involve several tuned quantities—the constant scaling factor for predicted uncertainties, the 2 cm standard-deviation clip, the depth-proportional uncertainty factor, and the robust loss scales for Creg and Cint—but the final values are not reported. Because the claim of 'principled uncertainty propagation' and little tuning (Abstract; Sec. 1) is part of the contribution, the paper should list all free parameters and the data splits used to select them.
minor comments (5)
- [Sec. 3.4] The sentence 'We consider a view c as inconsistent if any of the overlapping views' beta_i exceeds a ratio beta_hat of occluded pixels' is ambiguous: Eq. (6) defines beta_i as a ratio of inconsistent pixels, so the phrase 'ratio of occluded pixels' should be clarified or removed.
- [Eq. (7) / Appendix B] In the definition of Sigma_r, the last diagonal entry is written as sigma^2_{N-_u}; it should presumably be sigma^2_{N-_v}.
- [Fig. 7 caption] There are several typos in the caption: 'estiamtes', 'yileded', and 'uncertianties' should be corrected.
- [Table 1] The header of the right block is garbled in the manuscript ('max overlapminimal, 0% <5% <10% <30%'); please fix the column labels to make the overlap buckets unambiguous.
- [Sec. 4.1] The sentence 'The GT camera poses were estimated with COLMAP – achieving sufficient accuracy by using 10 to 100 times more images' is missing a subject; it should read 'The GT camera poses were estimated with COLMAP using 10 to 100 times more images, which we assume achieves sufficient accuracy.'
Circularity Check
No significant circularity: the two-view-only SfM pipeline is validated on external benchmarks against external baselines, and no predicted quantity reduces to a fitted input by construction.
full rationale
The derivation chain is self-contained. The central claim, that monocular depth and normal priors enable accurate incremental SfM from two-view tracks only, is implemented through the objective in Eqs. (2)-(5): a standard bundle adjustment term C_BA, a depth-regularization term C_reg, and an uncertainty-weighted prior term C_int. No output quantity is defined in terms of an input, and no evaluation metric is fitted by the method. The per-image scale in Eq. (1) is an internal median alignment of each monocular depth map to the current sparse 3D points; it is a normalization step, not a prediction of the pose or reconstruction quality that is later scored. Appendix C calibrates uncertainty scaling factors on a separate dataset or validation split, and the paper's Limitations section explicitly concedes the dependency: 'Our system depends on reliable uncertainties for the monocular priors. State-of-the-art depth models rarely estimate uncertainties and those that do are often over-confident.' That is an acknowledged assumption, not a circular reduction. Evaluations on ETH3D, SMERF, Tanks and Temples, and RealEstate10k compare against external baselines (COLMAP, GLOMAP, MASt3R-SfM, StudioSfM, etc.), and the ground-truth poses are derived from much larger reconstructions or benchmark data. Self-citations to COLMAP and GLOMAP point to publicly available, code-reproduced frameworks used both as a base and as baselines; they are not invoked as unverified uniqueness theorems. The unreported depth-check thresholds (gamma, beta_hat) and the load-bearing assumption that off-the-shelf priors are correct up to a per-image scale are legitimate correctness risks, but they are not circularity because the paper does not derive its headline result from those thresholds or from a fitted parameter renamed as a prediction.
Assumptions & free parameters
free parameters (5)
- depth uncertainty scaling factor =
not reported
- depth-proportional uncertainty factor =
not reported
- minimum depth stddev clip =
2 cm
- depth consistency thresholds gamma and beta_hat =
not reported
- robust loss scale for depth regularization and integration =
not reported
assumptions (4)
- domain assumption Monocular depth and normal priors are sufficiently accurate and their uncertainties calibratable to support fused SfM.
- domain assumption A global per-image scale factor can align each monocular depth map to the multi-view structure (Eq. 1).
- domain assumption Ground-truth poses for SMERF and Tanks and Temples, estimated with COLMAP from 10 to 100 times more images, are accurate enough for evaluation.
- standard math Bilateral normal integration with uncertainty weighting (Cao et al.) is a valid model for refining depth maps.
Cite this review
Pith. "Pith review of MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion." pith.science (2026). https://pith.science/paper/Z2FEURJH
@misc{pith2026250420040,
author = {Pith},
title = {Pith review of: MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z2FEURJH}},
note = {Machine review of arXiv:2504.20040}
}
read the original abstract
While Structure-from-Motion (SfM) has seen much progress over the years, state-of-the-art systems are prone to failure when facing extreme viewpoint changes in low-overlap, low-parallax or high-symmetry scenarios. Because capturing images that avoid these pitfalls is challenging, this severely limits the wider use of SfM, especially by non-expert users. We overcome these limitations by augmenting the classical SfM paradigm with monocular depth and normal priors inferred by deep neural networks. Thanks to a tight integration of monocular and multi-view constraints, our approach significantly outperforms existing ones under extreme viewpoint changes, while maintaining strong performance in standard conditions. We also show that monocular priors can help reject faulty associations due to symmetries, which is a long-standing problem for SfM. This makes our approach the first capable of reliably reconstructing challenging indoor environments from few images. Through principled uncertainty propagation, it is robust to errors in the priors, can handle priors inferred by different models with little tuning, and will thus easily benefit from future progress in monocular depth and normal estimation. Our code is publicly available at https://github.com/cvg/mpsfm.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
A Hybrid Neural-Microfacet BRDF Model for Real-Time Rendering
A hybrid BRDF model, combining a GGX analytical term with a tiny learned residual and gating network, fits measured materials more accurately than fully neural models at equal memory cost.
Reference graph
Works this paper leans on
-
[1]
Sameer Agarwal, Yasutaka Furukawa, Noah Snavely, Ian Simon, Brian Curless, Steven M Seitz, and Richard Szeliski. Building Rome in a day. TOG, 54(10):105–112, 2011. 1, 2
work page 2011
-
[2]
Sameer Agarwal, Keir Mierle, and Others. Ceres Solver. http://ceres-solver.org, 2024. 6
work page 2024
-
[3]
NetVLAD: CNN architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. NetVLAD: CNN architecture for weakly supervised place recognition. In CVPR, 2016. 6
work page 2016
-
[4]
Gwangbin Bae and Andrew J. Davison. Rethinking Inductive Biases for Surface Normal Estimation. In CVPR, 2024. 2, 6, 8, 10
work page 2024
-
[5]
Sequential Updating of Projective and Affine Struc- ture from Motion
Paul A Beardsley, Andrew Zisserman, and David William Murray. Sequential Updating of Projective and Affine Struc- ture from Motion. IJCV, 1997. 2
work page 1997
-
[6]
Aleksei Bochkovskii, Ama¨el Delaunoy, Hugo Germain, Mar- cel Santos, Yichao Zhou, Stephan R. Richter, and Vladlen Koltun. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second. arXiv:2410.02073, 2024. 2, 6, 8, 12
arXiv 2024
-
[7]
Visual Camera Re- Localization from RGB and RGB-D Images Using DSAC
Eric Brachmann and Carsten Rother. Visual Camera Re- Localization from RGB and RGB-D Images Using DSAC. IEEE TPAMI, 2021. 2
work page 2021
-
[8]
Eric Brachmann, Jamie Wynn, Shuai Chen, Tommaso Caval- lari, ´Aron Monszpart, Daniyar Turmukhambetov, and Vic- tor Adrian Prisacariu. Scene Coordinate Reconstruction: Posing of Image Collections via Incremental Learning of a Relocalizer. In ECCV, 2024. 2
work page 2024
Show all 74 references
-
[9]
Doppelgangers: Learning to Disambiguate Images of Similar Structures
Ruojin Cai, Joseph Tung, Qianqian Wang, Hadar Averbuch- Elor, Bharath Hariharan, and Noah Snavely. Doppelgangers: Learning to Disambiguate Images of Similar Structures. In ICCV, 2023. 1, 8 16 AUC(%): 49.64/85.67/96.05 AUC(%): 80.79/96.13/99.03 AUC(%): 24.51/53.14/76.65 AUC(%):...
2023
-
[10]
Hybrid camera pose estimation
Federico Camposeco, Andrea Cohen, Marc Pollefeys, and Torsten Sattler. Hybrid camera pose estimation. In CVPR,
-
[11]
Bilateral normal integration
Xu Cao, Hiroaki Santo, Boxin Shi, Fumio Okura, and Ya- suyuki Matsushita. Bilateral normal integration. In European Conference on Computer Vision, pages 552–567. Springer,
-
[12]
Locally Optimized RANSAC
Ondˇrej Chum, Jiˇr´ı Matas, and Josef Kittler. Locally Optimized RANSAC. In GCPR, 2003. 4
2003
-
[13]
SuperPoint: Self-Supervised Interest Point Detection and Description
Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. SuperPoint: Self-Supervised Interest Point Detection and Description. In CVPR Workshops, 2018. 1, 2, 6, 7, 13, 15
2018
-
[14]
Eric Dexheimer and Andrew J. Davison. COMO: Compact Mapping and Odometry. In ECCV, 2024. 3
2024
-
[15]
Daniel Duckworth, Peter Hedman, Christian Reiser, Pe- ter Zhizhin, Jean-Fran c ¸ois Thibert, Mario Lu ˇci´c, Richard Szeliski, and Jonathan T. Barron. SMERF: Streamable Mem- ory Efficient Radiance Fields for Real-Time Large-Scene Exploration. arXiv:2312.07541, 2023. 7, 9, 14
2023 arXiv
-
[16]
MASt3R- SfM: a Fully-Integrated Solution for Unconstrained Structure- from-Motion
Bardienus Duisterhof, Lojze Zust, Philippe Weinzaepfel, Vin- cent Leroy, Yohann Cabon, and Jerome Revaud. MASt3R- SfM: a Fully-Integrated Solution for Unconstrained Structure- from-Motion. arXiv:2409.19152, 2024. 1, 2, 7, 9, 10, 12, 14
2024 arXiv
-
[17]
D2-Net: A Trainable CNN for Joint Detection and Description of Local Features
Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Polle- feys, Josef Sivic, Akihiko Torii, and Torsten Sattler. D2-Net: A Trainable CNN for Joint Detection and Description of Local Features. In CVPR, 2019. 1, 2
2019
-
[18]
Sch¨onberger, and Marc Polle- feys
Mihai Dusmanu, Johannes L. Sch¨onberger, and Marc Polle- feys. Multi-View Optimization of Local Feature Geometry. In ECCV, 2020. 2
2020
-
[19]
RoMa: Robust Dense Feature Match- ing
Johan Edstedt, Qiyu Sun, Georg B¨okman, M˚arten Wadenb¨ack, and Michael Felsberg. RoMa: Robust Dense Feature Match- ing. CVPR, 2024. 2, 6, 7, 16
2024
-
[20]
Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
David Eigen, Christian Puhrsch, and Rob Fergus. Depth Map Prediction from a Single Image using a Multi-Scale Deep Network. NeurIPS, 2014. 2
2014
-
[21]
Building Rome on a cloudless day
Jan-Michael Frahm, Pierre Fite-Georgel, David Gallup, Tim Johnson, Rahul Raguram, Changchang Wu, Yi-Hung Jen, Enrique Dunn, Brian Clipp, Svetlana Lazebnik, et al. Building Rome on a cloudless day. In ECCV, 2010. 1, 2
2010
-
[22]
Privacy Preserving Structure-from-Motion
Marcel Geppert, Viktor Larsson, Pablo Speciale, Johannes L Sch¨onberger, and Marc Pollefeys. Privacy Preserving Structure-from-Motion. In ECCV, 2020. 1
2020
-
[23]
Haralick, Chung-Nan Lee, Karsten Ottenberg, and Michael N¨olle
Bert M. Haralick, Chung-Nan Lee, Karsten Ottenberg, and Michael N¨olle. Review and Analysis of Solutions of the Three Point Perspective Pose Estimation Problem. IJCV, 1994. 4
1994
-
[24]
Detector-Free Struc- ture from Motion
Xingyi He, Jiaming Sun, Yifan Wang, Sida Peng, Qixing Huang, Hujun Bao, and Xiaowei Zhou. Detector-Free Struc- ture from Motion. In CVPR, 2024. 7
2024
-
[25]
Cor- recting for Duplicate Scene Structure in Sparse 3D Recon- struction
Jared Heinly, Enrique Dunn, and Jan-Michael Frahm. Cor- recting for Duplicate Scene Structure in Sparse 3D Recon- struction. In ECCV, 2014. 5
2014
-
[26]
Reconstructing the World* in Six Days *(as Captured by the Yahoo 100 Million Image Dataset)
Jared Heinly, Johannes L Schonberger, Enrique Dunn, and Jan-Michael Frahm. Reconstructing the World* in Six Days *(as Captured by the Yahoo 100 Million Image Dataset). In CVPR, 2015. 1, 2 17
2015
-
[27]
Geometric Context from a Single Image
Derek Hoiem, Alexei A Efros, and Martial Hebert. Geometric Context from a Single Image. In ICCV, 2005. 2
2005
-
[28]
Metric3D v2: A Versatile Monocular Geo- metric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation
Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. Metric3D v2: A Versatile Monocular Geo- metric Foundation Model for Zero-Shot Metric Depth and Surface Normal Estimation. IEEE TPAMI, 2024. 6, 8, 10, 12
2024
-
[29]
Image Match- ing across Wide Baselines: From Paper to Practice
Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk, Jiˇr´ı Matas, Pascal Fua, Kwang Moo Yi, and Eduard Trulls. Image Match- ing across Wide Baselines: From Paper to Practice. IJCV,
-
[30]
Image-based localization using hybrid feature corre- spondences
Klas Josephson, Martin Byrod, Fredrik Kahl, and Kalle As- trom. Image-based localization using hybrid feature corre- spondences. In CVPR, 2007. 2
2007
-
[31]
Repurpos- ing Diffusion-Based Image Generators for Monocular Depth Estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Repurpos- ing Diffusion-Based Image Generators for Monocular Depth Estimation. In CVPR, 2024. 2
2024
-
[32]
3D Gaussian Splatting for Real-Time Radi- ance Field Rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radi- ance Field Rendering. TOG, 2023. 1
2023
-
[33]
Tanks and temples: Benchmarking large-scale scene reconstruction
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. TOG, 2017. 7, 9, 14
2017
-
[34]
Ground- ing Image Matching in 3D with MASt3R
Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing Image Matching in 3D with MASt3R. In ECCV, 2024. 2, 6, 10, 12, 15
2024
-
[35]
Pixel-Perfect Structure-from-Motion with Featuremetric Refinement
Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, and Marc Pollefeys. Pixel-Perfect Structure-from-Motion with Featuremetric Refinement. In ICCV, 2021. 2
2021
-
[36]
LightGlue: Local Feature Matching at Light Speed
Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Polle- feys. LightGlue: Local Feature Matching at Light Speed. In ICCV, 2023. 2, 6, 7, 15
2023
-
[37]
Depth-Guided Sparse Structure-from-Motion for Movies and TV Shows
Sheng Liu, Xiaohan Nie, and Raffay Hamid. Depth-Guided Sparse Structure-from-Motion for Movies and TV Shows. In CVPR, 2022. 3, 7
2022
-
[38]
David G. Lowe. Distinctive Image Features from Scale- Invariant Keypoints. IJCV, 60(2):91–110, 2004. 7
2004
-
[39]
Real-Time Visibility-Based Fusion of Depth Maps
Paul Merrell, Amir Akbarzadeh, Liang Wang, Philippos Mor- dohai, Jan-Michael Frahm, Ruigang Yang, David Nist´er, and Marc Pollefeys. Real-Time Visibility-Based Fusion of Depth Maps. In ICCV, 2007. 5
2007
-
[40]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Syn- thesis. In ECCV, 2020. 1
2020
-
[41]
Working hard to know your neighbor’s margins: Local descriptor learning loss
Anastasiia Mishchuk, Dmytro Mishkin, Filip Radenovic, and Jiri Matas. Working hard to know your neighbor’s margins: Local descriptor learning loss. NeurIPS, 30, 2017. 2
2017
-
[42]
OpenMVG: Open multiple view geometry
Pierre Moulon, Pascal Monasse, Romuald Perrot, and Renaud Marlet. OpenMVG: Open multiple view geometry. In In- ternational Workshop on Reproducible Research in Pattern Recognition, pages 60–74. Springer, 2016. 2
2016
-
[43]
Global Structure-from-Motion Revisited
Linfei Pan, Daniel Barath, Marc Pollefeys, and Johannes Lutz Sch¨onberger. Global Structure-from-Motion Revisited. In ECCV, 2024. 1, 2, 7
2024
-
[44]
Visual Modeling with a Hand-held Camera
Marc Pollefeys, Luc Van Gool, Maarten Vergauwen, Frank Verbiest, Kurt Cornelis, Jan Tops, and Reinhard Koch. Visual Modeling with a Hand-held Camera. IJCV, 2004. 2
2004
-
[45]
SuperGlue: Learning Feature Match- ing with Graph Neural Networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning Feature Match- ing with Graph Neural Networks. In CVPR, 2020. 1, 2
2020
-
[46]
Sch¨onberger, Pablo Speciale, Lukas Gruber, Viktor Larsson, Ondrej Miksik, and Marc Pollefeys
Paul-Edouard Sarlin, Mihai Dusmanu, Johannes L. Sch¨onberger, Pablo Speciale, Lukas Gruber, Viktor Larsson, Ondrej Miksik, and Marc Pollefeys. LaMAR: Benchmarking Localization and Mapping for Augmented Reality. In ECCV,
-
[47]
Pixel-Perfect Structure-From-Motion With Featuremetric Refinement
Paul-Edouard Sarlin, Philipp Lindenberger, Viktor Larsson, and Marc Pollefeys. Pixel-Perfect Structure-From-Motion With Featuremetric Refinement. IEEE TPAMI, 2023. 7
2023
-
[48]
Learning Depth from Single Monocular Images
Ashutosh Saxena, Sung Chung, and Andrew Ng. Learning Depth from Single Monocular Images. NeurIPS, 2005. 2
2005
-
[49]
How Do I Organize My Holiday Snaps?
Frederik Schaffalitzky and Andrew Zisserman. Multi-view Matching for Unordered Image Sets, or “How Do I Organize My Holiday Snaps?”. In ECCV, 2002. 2
2002
-
[50]
Structure-from-Motion Revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-Motion Revisited. In CVPR, 2016. 1, 2, 3, 7, 9, 10, 13, 15
2016
-
[51]
Pixelwise View Selection for Un- structured Multi-View Stereo
Johannes Lutz Sch¨onberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise View Selection for Un- structured Multi-View Stereo. In ECCV, 2016. 1, 14
2016
-
[52]
Comparative Evaluation of Hand-Crafted and Learned Local Features
Johannes Lutz Sch¨onberger, Hans Hardmeier, Torsten Sattler, and Marc Pollefeys. Comparative Evaluation of Hand-Crafted and Learned Local Features. In CVPR, 2017. 2
2017
-
[53]
A multi-view stereo benchmark with high- resolution images and multi-camera videos
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and An- dreas Geiger. A multi-view stereo benchmark with high- resolution images and multi-camera videos. In CVPR, 2017. 7, 9, 10, 12, 13, 14
2017
-
[54]
Semi-Dense Feature Matching With Transformers and its Applications in Multiple-View Geome- try
Zehong Shen, Jiaming Sun, Yuang Wang, Xingyi He, Hujun Bao, and Xiaowei Zhou. Semi-Dense Feature Matching With Transformers and its Applications in Multiple-View Geome- try. IEEE TPAMI, 2023. 7
2023
-
[55]
Cam- era Network Calibration from Dynamic Silhouettes
Sudipta Sinha, Marc Pollefeys, and Leonard McMillan. Cam- era Network Calibration from Dynamic Silhouettes. In CVPR,
-
[56]
FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent
Cameron Smith, David Charatan, Ayush Tewari, and Vincent Sitzmann. FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent. ECCV, 2024. 2
2024
-
[57]
Photo Tourism: exploring photo collections in 3D
Noah Snavely, Steven M Seitz, and Richard Szeliski. Photo Tourism: exploring photo collections in 3D. In TOG, 2006. 1, 2
2006
-
[58]
Privacy preserving image-based localization
Pablo Speciale, Johannes L Schonberger, Sing Bing Kang, Sudipta N Sinha, and Marc Pollefeys. Privacy preserving image-based localization. In CVPR, 2019. 1
2019
-
[59]
LoFTR: Detector-Free Local Feature Matching with Transformers
Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xi- aowei Zhou. LoFTR: Detector-Free Local Feature Matching with Transformers. In CVPR, 2021. 2, 7
2021
-
[60]
Recovering 3D Shape and Motion from Image Streams Using Non-Linear Least Squares
Richard Szeliski and Sing Bing Kang. Recovering 3D Shape and Motion from Image Streams Using Non-Linear Least Squares. Journal of Visual Communication and Image Repre- sentation, 1994. 2 18
1994
-
[61]
Is this the right place? geometric-semantic pose verification for indoor visual localization
Hajime Taira, Ignacio Rocco, Jiri Sedlar, Masatoshi Okutomi, Josef Sivic, Tomas Pajdla, Torsten Sattler, and Akihiko Torii. Is this the right place? geometric-semantic pose verification for indoor visual localization. In CVPR, 2019. 5
2019
-
[62]
GeoCalib: Single-image Calibration with Geometric Optimization
Alexander Veicht, Paul-Edouard Sarlin, Philipp Lindenberger, and Marc Pollefeys. GeoCalib: Single-image Calibration with Geometric Optimization. In ECCV, 2024. 2
2024
-
[63]
PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment
Jianyuan Wang, Christian Rupprecht, and David Novotny. PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment. In ICCV, 2023. 7
2023
-
[64]
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. VGGSfM: Visual Geometry Grounded Deep Structure From Motion. In CVPR, 2024. 1, 2, 7, 14
2024
-
[65]
DUSt3R: Geometric 3D Vision Made Easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. DUSt3R: Geometric 3D Vision Made Easy. In CVPR, 2024. 1, 2
2024
-
[66]
Generalized differentiable RANSAC
Tong Wei, Yash Patel, Alexander Shekhovtsov, Jiri Matas, and Daniel Barath. Generalized differentiable RANSAC. In ICCV, 2023. 2
2023
-
[67]
DeepSFM: Structure From Motion Via Deep Bundle Adjustment
Xingkui Wei, Yinda Zhang, Zhuwen Li, Yanwei Fu, and Xi- angyang Xue. DeepSFM: Structure From Motion Via Deep Bundle Adjustment. In ECCV, 2020. 2
2020
-
[68]
VisualSFM : A Visual Structure from Motion System
Changchang Wu. VisualSFM : A Visual Structure from Motion System. http://www.cs.washington.edu/ homes/ccwu/vsfm, 2011. 1, 2
2011
-
[69]
Towards Linear-time Incremental Structure from Motion
Changchang Wu. Towards Linear-time Incremental Structure from Motion. In 3DV, 2013. 5
2013
-
[70]
Depth Anything V2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth Anything V2. arXiv:2406.09414, 2024. 2, 6, 8, 12, 17
2024 arXiv
-
[71]
Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Im- age
Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaix- uan Wang, Xiaozhi Chen, and Chunhua Shen. Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Im- age. In ICCV, 2023. 2
2023
-
[72]
Disambiguating Visual Relations Using Loop Constraints
Christopher Zach, Manfred Klopschitz, and Marc Pollefeys. Disambiguating Visual Relations Using Loop Constraints. In CVPR, 2010. 1
2010
-
[73]
Structure From Motion Using Structure-Less Resection
Enliang Zheng and Changchang Wu. Structure From Motion Using Structure-Less Resection. In ICCV, 2015. 2, 7
2015
-
[74]
Stereo Magnification: Learning View Synthesis using Multiplane Images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo Magnification: Learning View Synthesis using Multiplane Images. In TOG, 2018. 7, 14 19
2018
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.