REVIEW 3 major objections 5 minor 15 cited by
Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A single forward pass reconstructs 1,000+ unposed images into one 3D frame.
desk verdict Real speed/scalability advance for multi-view pointmap reconstruction, but the SOTA pose claim leans on a cherry-picked variant and the 1000+ view accuracy is not actually demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the fusion Transformer: a 24-layer ViT-L-style stack that takes the concatenated, encoder-produced patch tokens of all views plus per-image index position embeddings and runs all-to-all self-attention, so every image can attend to every other image in one pass. The mechanism that enables 'train short, test long' is randomized index-embedding sampling from a pool of N'=1,000 positions, adapted from position interpolation, which makes the network robust to view counts far above the 20 used in training. The objective is DUSt3R's confidence-weighted, normalized pointmap regression loss, applied to both the global pointmap (first-camera frame) and the local pointmap (per-camera frame), with two DPT (dense-prediction transformer) heads decoding pointmaps and confidences. Memory-efficient attention and sharded optimizer states let this single pass fit on GPUs, and tensor-parallel DPT heads push the ceiling to roughly 1,500 views.
What would settle it
Evaluate Fast3R on a 1,000+ frame trajectory of a large scene outside its training distribution, such as a full building interior or city block, and compare per-view rotation and translation errors plus pointmap drift against a high-quality COLMAP reference. If per-view pose error grows with the number of views, or if low-confidence views drift and dropping them leaves coverage holes, the train-short/test-long claim is not general.
Extended reading notes
Core claim
Fast3R's central discovery is that the pairwise bottleneck in DUSt3R is removable: concatenating the patch tokens of all N images into one fusion Transformer, with all-to-all attention across every view, lets the model regress all pointmaps simultaneously and jointly reason about all cameras. The global pointmap is anchored to the first image's coordinate frame; a second, local head predicts each view's geometry in its own frame, and the local maps are aligned to the global frame with ICP for dense reconstruction. Training uses only N=20 views per sample, but view-index position embeddings are randomly sampled from a pool of 1,000, a randomized form of position interpolation that makes the model treat inference-time masking as natural and generalize to 1,000-1,500 views. Experiments show per-view pose and reconstruction accuracy improve when more views are used at inference, even beyond the training count, and that the model scales with views, data, and fusion-transformer size.
Load-bearing premise
The load-bearing premise is the 'train short, test long' scheme of Section 3.3: training on samples of only 20 views, with each view's index label randomly drawn from a pool of 1,000, is assumed to teach the model to handle 1,000 or more real views at inference; if that extrapolation fails on larger or more varied scenes, the central scalability claim collapses.
Editorial extensions
If this is right
- Camera poses and dense geometry for 1,000+ images can be obtained in one pass, removing per-scene global alignment as a separate optimization step.
- Per-view accuracy improves as more views are fed in, so applications can trade inference time for reconstruction quality by simply adding frames.
- Training on more views, more data, or a larger fusion transformer each yield consistent gains, suggesting the architecture benefits from standard scaling.
- Because there is no sequential dependency, the same model parallelizes across devices; the paper reports up to 1,500 views in a single forward pass on one A100-class setup.
- Pairwise methods like DUSt3R run out of memory around 48 views in the same setting, so Fast3R's single-pass design removes that ceiling.
Reading between the lines
- If train-short/test-long transfers beyond the reported benchmarks, one-pass many-view reconstruction could replace per-scene structure-from-motion as the initialization step for large-scale mapping: the dominant cost becomes a single attention pass instead of O(N^2) pair reconstructions.
- The global/local head split points to a modular recipe for downstream rendering: use the global head to anchor the frame and the local head for metric detail, then feed aligned local pointmaps directly into splatting or meshing without bundle adjustment.
- A direct stress test would be training the same architecture with 100+ views per sample; if per-view accuracy no longer improves, data diversity rather than the attention mechanism will be the limiting factor.
- The paper's own limitation note about drift past roughly 300 views suggests long-range global consistency, not throughput, is the next bottleneck; video-aware position encodings are the natural repair.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Fast3R, a transformer-based extension of DUSt3R that takes N unordered, unposed images and predicts per-pixel local and global pointmaps in a single forward pass using all-to-all attention. To enable inference with more views than used in training, the model randomly samples image-index embeddings from a large pool (position interpolation). The authors report camera-pose results on CO3Dv2 and RealEstate10K, 3D-reconstruction results on 7-Scenes, NRGBD, and DTU, and system-level throughput/memory scaling up to 1500 views, claiming state-of-the-art pose accuracy, 251.1 FPS, and the ability to reconstruct 1000+ images in one pass without global alignment.
Significance. If the central claims hold, Fast3R is a meaningful advance: replacing pairwise reconstruction plus global alignment with a single-pass many-view transformer is a natural and potentially impactful direction, and the reported speedups over DUSt3R and Spann3R are large. The paper includes several commendable elements: careful ablations of training-view count, model and data scaling, an honest limitations section, and reproducible benchmark protocols. The main significance hinges on whether the model actually maintains reconstruction/pose accuracy at the 1000+ view scale advertised in the title, since the current evidence at that scale is limited to time and memory measurements, while the paper's own limitations admit drift beyond roughly 300 views.
major comments (3)
- [Table 1 and Section 4.2] The 'state-of-the-art' pose claim on CO3D is based on the Fast3R-no-outdoor variant (99.7% RRA@15), while the full Fast3R model, which is the one used for the reconstruction Tables 3-4, ties DUSt3R at 96.2% RRA@15 and is worse on RTA@15 (81.6 vs 86.8). The paper should make this distinction explicit in the text and justify why the no-outdoor model is the appropriate SOTA comparison. Presenting both rows as 'Fast3R (Ours)' in Table 1 without a clear statement in the body risks overstating the accuracy advantage of the full method.
- [Section 4 (Architecture Details) vs Appendix A] There is a direct contradiction about the fusion transformer size used in the main experiments. Section 4 states 'The Fusion Transformer is a ViT-Large model initialized from scratch,' while Appendix A states 'Note that the Fusion Transformer size used in the main text for all experiments is a ViT-base.' This inconsistency affects reproducibility and the interpretation of the model-scaling experiments; please correct one of the two statements and clarify which configuration produced Tables 1-4.
- [Abstract and Section 1, contribution 3] The claimed 'over 14x error reduction compared to DUSt3R' is not obviously supported by the numbers in Table 1. With DUSt3R at 96.2% RRA@15 and Fast3R-no-outdoor at 99.7% RRA@15, the error rates (3.8% vs 0.3%) imply roughly a 12.7x reduction, not 14x. The claim should be recomputed or reworded to specify which metric and which variant are used.
minor comments (5)
- [Section 3.3] The phrase 'This strategy enables Fast3R to handle N = 1000 images during inference, even if only trained with N = 20 images' is not backed by accuracy experiments at N=1000; consider changing 'handle' to 'process' or adding a qualifier.
- [Table 1 caption] The caption states 'Fast3R does not assume known camera intrinsics,' but it is unclear whether the same is true for the DUSt3R and MASt3R baselines in the pose comparison; please clarify the intrinsic-assumption setup for all methods.
- [Section 4.3 and Table 3] For the reconstruction evaluation, the paper aligns local pointmaps to global pointmaps using ICP. Since the local head is not trained with a global frame, the ICP alignment step is itself an inference-time geometric postprocess; the paper should state this clearly in the main text and discuss any sensitivity to ICP initialization.
- [Section 5.3] Figure 8 evaluates the position-interpolation ablation only on a 4-to-24 view transfer. Given that the main model is trained with 20 views and tested with 1000+, an ablation at a larger ratio (e.g., 20 to 100 or 320 views) would be more informative for the generalizability claim.
- [Appendix A] The model-scaling experiment reports results for ViT-base, ViT-large, and ViT-huge fusion transformers but does not specify the compute budget or number of training steps for each size beyond '60k steps'; please provide the full settings so the scaling trend is reproducible.
Circularity Check
No circularity: Fast3R's claims rest on supervised training against external ground-truth datasets and held-out evaluation; the DUSt3R-derived loss and initialization are external prior work, not the paper's own output.
full rationale
The paper's derivation chain is empirical, not analytical, so there is no input/output equivalence to exhibit. Fast3R is trained with a supervised pointmap regression loss (Eqs. 1-3) using ground-truth labels from external datasets (CO3D, ScanNet++, ARKitScenes, Habitat, BlendedMVS, MegaDepth), and its pose and reconstruction numbers are measured on held-out benchmarks (CO3Dv2, RealEstate10K, 7-Scenes, NRGBD, DTU). The single load-bearing methodological choice that could conceivably be circular—the train-short/test-long index-embedding scheme of Sec. 3.3—is not circular: it is a generalization hypothesis, supported in the paper only by a 4-view-to-24-view transfer experiment (Fig. 8) and by the observation that accuracy improves up to 50 test views (Fig. 5). The paper's Limitations section explicitly concedes that beyond roughly 300 views in large scenes, low-confidence pointmaps begin to drift; that is an evidence/scoping weakness for the '1000+ images' headline, not a definitional reduction. DUSt3R pretrained weights initialize the encoder and global head, but DUSt3R is external prior work from a different group, and using its loss and weights does not make Fast3R's held-out results equal to its training signal. The only self-citations ([14], [18], [50]) appear in related-work context and carry no load-bearing argument. No uniqueness theorem is imported, no ansatz is justified by self-citation, and no known result is simply renamed. Accordingly, no circular step can be identified by the required evidence standard.
Assumptions & free parameters
free parameters (3)
- N' (index embedding pool size) =
1000
- N_train (views per training sample) =
20
- alpha (confidence loss weight) =
unspecified
assumptions (3)
- domain assumption Ground-truth pointmaps from the six training datasets are sufficiently accurate to supervise metric 3D reconstruction.
- domain assumption Position interpolation with uniformly random index sampling transfers from 20 views to 1000+ views.
- domain assumption The first image I1 consistently defines the global coordinate frame via fixed index embedding.
Cite this review
Pith. "Pith review of Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass." pith.science (2026). https://pith.science/paper/UFH2BRMX
@misc{pith2026250113928,
author = {Pith},
title = {Pith review of: Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass},
year = {2026},
howpublished = {\url{https://pith.science/paper/UFH2BRMX}},
note = {Machine review of arXiv:2501.13928}
}
read the original abstract
Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessitating costly global alignment procedures to reconstruct from multiple views. In this work, we propose Fast 3D Reconstruction (Fast3R), a novel multi-view generalization to DUSt3R that achieves efficient and scalable 3D reconstruction by processing many views in parallel. Fast3R's Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment. Through extensive experiments on camera pose estimation and 3D reconstruction, Fast3R demonstrates state-of-the-art performance, with significant improvements in inference speed and reduced error accumulation. These results establish Fast3R as a robust alternative for multi-view applications, offering enhanced scalability without compromising reconstruction accuracy.
Figures
Figures from the paper (11 more)
Forward citations
Cited by 15 Pith papers
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.
-
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
WinT3R combines a sliding-window decoder with a camera token pool to achieve state-of-the-art online 3D reconstruction and camera pose estimation at 17.2 FPS.
-
SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization
SAIL-Recon adds an anchor-based neural scene representation and attention masking to a VGGT-style regression transformer, enabling feedforward SfM on thousands of images with strong pose and view synthesis accuracy.
-
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
An incremental 3D Gaussian Splatting pipeline that jointly optimizes camera poses and scene geometry using MASt3R priors and density-adaptive octree anchors achieves state-of-the-art novel view synthesis on casual lon...
-
Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction
Puzzles synthesizes posed video-depth clips from single images and keyframes, letting 3D reconstruction models match full-data accuracy using only 10% of the data.
-
Test3R: Learning to Reconstruct 3D at Test Time
Test3R improves 3D reconstruction by optimizing visual prompts at test time so that pointmaps from different image pairs are geometrically consistent.
-
PoseIDON: 6DoF Pose Estimation with Foundation Model Features for Marine Sediment Burial Mapping
PoseIDON estimates 6DoF object pose and seafloor plane from ROV video using DINOv2/FoundPose and photogrammetry, achieving about 10 cm mean burial-depth error on 54 buried objects.
-
X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
A large transformer with fixed-voxel Gaussian splatting reconstructs CT volumes from 6-10 X-ray projections in under a second, substantially beating prior sparse-view methods in simulation.
-
FastMap: Revisiting Structure from Motion through First-Order Optimization
FastMap uses fused first-order gradient kernels in a point-free global SfM pipeline to run up to 10x faster than COLMAP and GLOMAP on dense photo sets, with comparable accuracy at relaxed thresholds but weaker strict-...
-
RayZer: A Self-supervised Large View Synthesis Model
A self-supervised transformer model predicts camera poses and scene features from unposed images and renders novel views, reaching performance on par with pose-supervised baselines.
-
Quo Vadis, World Modeling?
An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.
-
DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring
A pose-free deblurring 3D Gaussian Splatting pipeline using DUSt3R point clouds, confidence-balanced sampling, and event-decoded latent image supervision.
-
SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
SceneCompleter jointly denoises RGB and depth latents, conditioned on projected depth and global scene features, yielding higher quality and more pose-consistent novel views than 2D-inpainting baselines.
-
Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction
Mono3R adds a monocular-guided refinement module to DUSt3R, aligning frozen MoGe pointmaps and features with pairwise predictions and iteratively updating them, yielding better pose estimates and denser point clouds i...
-
Reconstructing 4D Spatial Intelligence: A Survey
A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.
Reference graph
Works this paper leans on
-
[1]
Large-scale data for multiple-view stereopsis
Henrik Aanæs, Rasmus Ramsbøl Jensen, George V ogiatzis, Engin Tola, and Anders Bjorholm Dahl. Large-scale data for multiple-view stereopsis. International Journal of Computer Vision, 120:153–168, 2016. 6, 7
work page 2016
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report, 2024. 5
work page 2024
-
[3]
Neural rgb-d surface reconstruction
Dejan Azinovi ´c, Ricardo Martin-Brualla, Dan B Goldman, Matthias Nießner, and Justus Thies. Neural rgb-d surface reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 6290– 6301, 2022. 6, 7, 8
work page 2022
-
[4]
Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, and Elad Shulman. Ark- itscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. arXiv preprint arXiv:2111.08897, 2021. 2, 4, 5
arXiv 2021
-
[5]
Extending context window of large language models via positional interpolation
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuan- dong Tian. Extending context window of large language models via positional interpolation. InProceedings of the In- ternational Conference on Learning Representations (ICLR),
-
[6]
Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2017. 2
work page 2017
-
[7]
FlashAttention-2: Faster attention with better par- allelism and work partitioning
Tri Dao. FlashAttention-2: Faster attention with better par- allelism and work partitioning. In International Conference on Learning Representations (ICLR), 2024. 5, 8
work page 2024
-
[8]
Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Tri Dao, Daniel Y . Fu, Stefano Ermon, Atri Rudra, and Christopher R´e. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 5, 8
work page 2022
Show all 76 references
-
[9]
Superpoint: Self-supervised interest point detection and description
Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. Superpoint: Self-supervised interest point detection and description. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 337–33712, 2017. 6
2018
-
[10]
Tap-vid: A benchmark for tracking any point in a video, 2022
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adri `a Re- casens, Lucas Smaira, Yusuf Aytar, Jo ˜ao Carreira, Andrew Zisserman, and Yi Yang. Tap-vid: A benchmark for tracking any point in a video, 2022. 16
2022
-
[11]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[12]
The llama 3 herd of models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models, 2024. 5
2024
-
[13]
D2-net: A trainable cnn for joint description and detection of local features
Mihai Dusmanu, Ignacio Rocco, Tom ´as Pajdla, Marc Polle- feys, Josef Sivic, Akihiko Torii, and Torsten Sattler. D2-net: A trainable cnn for joint description and detection of local features. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 808...
2019
-
[14]
Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans
Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 10786–10796, 2021. 3
2021
-
[15]
Fairscale: A general purpose modular pytorch library for high performance and large scale train- ing
FairScale authors. Fairscale: A general purpose modular pytorch library for high performance and large scale train- ing. https://github.com/facebookresearch/ fairscale, 2021. 5
2021
-
[16]
Instantsplat: Unbounded sparse-view pose-free gaus- sian splatting in 40 seconds, 2024
Zhiwen Fan, Wenyan Cong, Kairun Wen, Kevin Wang, Jian Zhang, Xinghao Ding, Danfei Xu, Boris Ivanovic, Marco Pavone, Georgios Pavlakos, Zhangyang Wang, and Yue Wang. Instantsplat: Unbounded sparse-view pose-free gaus- sian splatting in 40 seconds, 2024. 12
2024
-
[17]
Massively parallel multiview stereopsis by surface normal diffusion
Silvano Galliani, Katrin Lasinger, and Konrad Schindler. Massively parallel multiview stereopsis by surface normal diffusion. In Proceedings of the IEEE international confer- ence on computer vision, pages 873–881, 2015. 1
2015
-
[18]
Silk: Sim- ple learned keypoints
Pierre Gleize, Weiyao Wang, and Matt Feiszli. Silk: Sim- ple learned keypoints. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , pages 10932–10942, 2023. 2
2023
-
[19]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings...
2024
-
[20]
Multiple View Ge- ometry in Computer Vision
Richard Hartley and Andrew Zisserman. Multiple View Ge- ometry in Computer Vision . Cambridge University Press, New York, NY , USA, 2 edition, 2003. 2
2003
-
[21]
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and zhifeng Chen. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Syst...
2019
-
[22]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Deven- dra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guil- laume Lample, L ´elio Renard Lavaud, Lucile Saulnier, Mari...
2024
-
[23]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing image matching in 3d with mast3r. arXiv preprint arXiv:2406.09756, 2024. 6
2024 arXiv
-
[24]
Ground- ing image matching in 3d with mast3r, 2024
Vincent Leroy, Yohann Cabon, and Jerome Revaud. Ground- ing image matching in 3d with mast3r, 2024. 2, 5
2024
-
[25]
Pytorch distributed: Experiences on accelerating data parallel train- ing
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. Pytorch distributed: Experiences on accelerating data parallel train- ing. CoRR, abs/2006.15704, 2020. 5
2006 arXiv
-
[26]
Megadepth: Learning single- view depth prediction from internet photos
Zhengqi Li and Noah Snavely. Megadepth: Learning single- view depth prediction from internet photos. In Computer Vision and Pattern Recognition (CVPR), 2018. 5
2018
-
[27]
Relpose++: Recovering 6d poses from sparse-view observations
Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tul- siani. Relpose++: Recovering 6d poses from sparse-view observations. arXiv preprint arXiv:2305.04926, 2023. 2, 6
2023 arXiv
-
[28]
Pixel-perfect structure-from-motion with featuremetric refinement
Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, and Marc Pollefeys. Pixel-perfect structure-from-motion with featuremetric refinement. 2021 IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), pages 5967– 5977, 2021. 6
2021
-
[29]
Fixing weight decay reg- ularization in adam
Ilya Loshchilov and Frank Hutter. Fixing weight decay reg- ularization in adam. CoRR, abs/1711.05101, 2017. 5
2017 arXiv
-
[30]
Tard ´os
Ra ´ul Mur-Artal and Juan D. Tard ´os. ORB-SLAM2: an open-source SLAM system for monocular, stereo and RGB- D cameras. IEEE Transactions on Robotics , 33(5):1255– 1262, 2017. 2
2017
-
[31]
Orb-slam: a versatile and accurate monocular slam system
Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos. Orb-slam: a versatile and accurate monocular slam system. IEEE transactions on robotics , 31(5):1147–1163,
-
[32]
Ef- ficient large-scale language model training on GPU clusters
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. Ef- ficient large-scale language model training on GPU clusters...
2021 arXiv
-
[33]
Maxime Oquab, Timoth ´ee Darcet, Theo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Rus- sell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang- Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Ni...
2023
-
[34]
Infinite photore- alistic worlds using procedural generation
Alexander Raistrick, Lahav Lipson, Zeyu Ma, Lingjie Mei, Mingzhe Wang, Yiming Zuo, Karhan Kayan, Hongyu Wen, Beining Han, Yihan Wang, Alejandro Newell, Hei Law, Ankit Goyal, Kaiyu Yang, and Jia Deng. Infinite photore- alistic worlds using procedural generation. In Proceedings ...
2023
-
[35]
Zero: Memory optimization towards training A trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimization towards training A trillion parameter models. CoRR, abs/1910.02054, 2019. 5
1910 arXiv
-
[36]
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Ren ´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. In IEEE Transactions on Pattern Analysis and Ma- chine Intelligence (TPAMI), 2020. 3
2020
-
[37]
Vision transformers for dense prediction
Ren ´e Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. arXiv preprint arXiv:2103.13413, 2021. 4, 5
2021 arXiv
-
[38]
Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pag...
2020
-
[39]
Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In Proceedings of the IEEE/CVF international conference on computer vi...
2021
-
[40]
Susskind
Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M. Susskind. Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding. In International Conference on Computer Vision (ICCV) 2021,
2021
-
[41]
Superglue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superglue: Learning feature matching with graph neural networks. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4937–4946, 2019. 6
2020
-
[42]
SuperGlue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. SuperGlue: Learning feature matching with graph neural networks. In CVPR, 2020. 2
2020
-
[43]
Habitat: A platform for embodied AI research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A platform for embodied AI research. CoRR, abs/1904.01201, 2019. 5
1904 arXiv
-
[44]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 1, 2
2016
-
[45]
Hechtman
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sep- assi, and Blake A. Hechtman. Mesh-tensorflow: Deep learn- ing for supercomputers. CoRR, abs/1811.02084, 2018. 5
2018 arXiv
-
[46]
Megatron- lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron- lm: Training multi-billion parameter language models using model parallelism. CoRR, abs/1909.08053, 2019. 5
1909 arXiv
-
[47]
Scene co- ordinate regression forests for camera relocalization in rgb-d images
Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, and Andrew Fitzgibbon. Scene co- ordinate regression forests for camera relocalization in rgb-d images. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 2930...
2013
-
[48]
Flowmap: High-quality camera poses, intrinsics, and depth via gradient descent
Cameron Smith, David Charatan, Ayush Kumar Tewari, and Vincent Sitzmann. Flowmap: High-quality camera poses, intrinsics, and depth via gradient descent. ArXiv, abs/2404.15259, 2024. 2
2024 arXiv
-
[49]
Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ra- mamoorthi, Jonathan T
Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ra- mamoorthi, Jonathan T. Barron, and Ren Ng. Fourier fea- tures let networks learn high frequency functions in low di- mensional domains. In Advances in Neural I...
2020
-
[50]
Aden: Adaptive density representations for sparse-view camera pose estima- tion
Hao Tang, Weiyao Wang, and Matt Feiszli. Aden: Adaptive density representations for sparse-view camera pose estima- tion. arXiv preprint arXiv:2408.09042, 2024. 2
2024 arXiv
-
[51]
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Zachary Teed and Jia Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. In Neural Infor- mation Processing Systems, 2021. 2
2021
-
[52]
DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
Zachary Teed and Jia Deng. DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras. Ad- vances in neural information processing systems, 2021. 2
2021
-
[53]
Probabilistic robotics
Sebastian Thrun. Probabilistic robotics. Communications of the ACM, 45(3):52–57, 2002. 1
2002
-
[54]
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models, 2023. 5
2023
-
[55]
EPIC Fields: Marrying 3D Geometry and Video Understanding
Vadim Tschernezki, Ahmad Darkhalil, Zhifan Zhu, David Fouhey, Iro Larina, Diane Larlus, Dima Damen, and Andrea Vedaldi. EPIC Fields: Marrying 3D Geometry and Video Understanding. In Proceedings of the Neural Information Processing Systems (NeurIPS), 2023. 2
2023
-
[56]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems, 2017. 2
2017
-
[57]
3d reconstruction with spatial memory
Hengyi Wang and Lourdes Agapito. 3d reconstruction with spatial memory. arXiv preprint arXiv:2408.16061, 2024. 3, 5, 6
2024 arXiv
-
[58]
Vggsfm: Visual geometry grounded deep structure from motion
Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. Vggsfm: Visual geometry grounded deep structure from motion. 2024 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 21686– 21697, 2023. 2
2024
-
[59]
Rupprecht, and David Novotn ´y
Jianyuan Wang, C. Rupprecht, and David Novotn ´y. Posed- iffusion: Solving pose estimation via diffusion-aided bundle adjustment. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9739–9749, 2023. 2, 6
2023
-
[60]
Dust3r: Geometric 3d vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and J ´erˆome Revaud. Dust3r: Geometric 3d vision made easy. 2024 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 20697– 20709, 2023. 6
2024
-
[61]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. In CVPR, 2024. 1, 2, 3, 4, 5, 6
2024
-
[62]
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Se- bastian Scherer. Tartanair: A dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020. 12
2020
-
[63]
Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion
Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Ro- main Br´egier, Yohann Cabon, Vaibhav Arora, Leonid Ants- feld, Boris Chidlovskii, Gabriela Csurka, and J ´erˆome Re- vaud. Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion. Advances in Neura...
2022
-
[64]
Hug- gingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chau- mond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R´emi Louf, Morgan Funtowicz, and Jamie Brew. Hug- gingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771, 2019. 5
1910 arXiv
-
[65]
Frozenrecon: Pose-free 3d scene reconstruction with frozen depth models
Guangkai Xu, Wei Yin, Hao Chen, Chunhua Shen, Kai Cheng, and Feng Zhao. Frozenrecon: Pose-free 3d scene reconstruction with frozen depth models. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 9310–9320, 2023. 6
2023
-
[66]
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In CVPR, 2024. 3, 4, 8
2024
-
[67]
Blendedmvs: A large-scale dataset for generalized multi-view stereo net- works
Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo net- works. Computer Vision and Pattern Recognition (CVPR) ,
-
[68]
Scannet++: A high-fidelity dataset of 3d in- door scenes
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d in- door scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12–22, 2023. 2, 4, 5
2023
-
[69]
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pas- cal V . Fua. Lift: Learned invariant feature transform. In European Conference on Computer Vision, 2016. 2
2016
-
[70]
Monst3r: A simple approach for estimat- ing geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jam- pani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming- Hsuan Yang. Monst3r: A simple approach for estimat- ing geometry in the presence of motion. arXiv preprint arxiv:2410.03825, 2024. 2, 12
-
[71]
Zhang, Deva Ramanan, and Shubham Tulsiani
Jason Y . Zhang, Deva Ramanan, and Shubham Tulsiani. Rel- pose: Predicting probabilistic relative rotation for single ob- jects in the wild. ArXiv, abs/2208.05963, 2022. 6
2022 arXiv
-
[72]
Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani
Jason Y . Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cam- eras as rays: Pose estimation via ray diffusion. ArXiv, abs/2402.14817, 2024. 2
2024 arXiv
-
[73]
Wang Zhao, Shaohui Liu, Hengkai Guo, Wenping Wang, and Y . Liu. Particlesfm: Exploiting dense point trajectories for localizing moving cameras in the wild. In European Confer- ence on Computer Vision, 2022. 2
2022
-
[74]
Pointodyssey: A large-scale synthetic dataset for long-term point tracking
Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wet- zstein, and Leonidas J Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 19855–19865, 2023. 12
2023
-
[75]
Stereo magnification: learning view syn- thesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: learning view syn- thesis using multiplane images. ACM Trans. Graph., 37(4),
-
[76]
loop closure
Dawei Zhu, Nan Yang, Liang Wang, Yifan Song, Wenhao Wu, Furu Wei, and Sujian Li. Pose: Efficient context win- dow extension of llms via positional skip-wise training. In The Twelfth International Conference on Learning Represen- tations. 8 Fast3R: Towards 3D Reconstruction of ...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.