REVIEW 4 major objections 6 minor 68 references
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Training a diffusion model on over a million 360° videos lets it synthesize free-camera views of real scenes from a single image and reconstruct their 3D geometry.
desk verdict The 360-1M dataset and correspondence pipeline are a genuine contribution, but the headline scene-geometry result rests on an evaluation loop that uses Dust3R for both the pseudo-ground truth and the reconstruction. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a scalable correspondence-mining pipeline for 360° video. Frames are sampled at one per second, each equirectangular frame is projected into four views at 90° yaw increments, and pairs within a 20-frame window are fed to Dust3R, which returns relative poses and confidence maps; a mean-confidence threshold filters out non-overlapping pairs. A graph-propagation step then links frames that share a common correspondent, recovering long-range pairs without exhaustive search, and the dimensionless Dust3R poses are anchored to metric scale by fitting a scale factor against monocular depth from Depth Anything. The generative side is a latent diffusion U-Net whose conditioning includes rotation and translation and whose output is multiplied by a learned motion mask with an auxiliary loss that keeps the mask from collapsing to zero; at inference, views are sampled along a smooth trajectory and the resulting image set is fed back through Dust3R to build the 3D scene.
What would settle it
On a held-out set of 360° scenes, recover camera trajectories with an independent metric-scale system such as LiDAR-equipped scanning or COLMAP with known calibration, generate ODIN views from one frame, reconstruct them with Dust3R, and measure the alignment error between the reconstructed point cloud and the independent geometry; a large misalignment alongside small Dust3R reconstruction error would falsify the claim that ODIN's views are geometrically accurate.
Extended reading notes
Core claim
ODIN, a latent diffusion model conditioned on both camera rotation and translation, is trained on 360-1M and is argued to be the first model that can reasonably synthesize real-world 3D scenes and reconstruct their geometry from a single input image with free camera movement. On the DTU and Mip-NeRF 360 novel-view-synthesis benchmarks it improves LPIPS over prior single-image methods without fine-tuning, and on Google Scanned Objects and a held-out 360-1M split it improves Chamfer distance and volumetric IoU for 3D reconstruction. The enabling observation is that a 360° video contains, in principle, many views of the same content from different positions: by rotating the equirectangular projection of nearby frames, one can align them to overlapping views and recover their relative pose.
Load-bearing premise
The evaluation of 3D reconstruction quality assumes Dust3R's pose and pointmap estimates are accurate enough to serve as ground truth, even though the same model is used to find training correspondences and to reconstruct ODIN's output.
Editorial extensions
If this is right
- A single image of a real scene becomes enough to generate a sequence of views that supports 3D reconstruction, without per-scene optimization or known camera poses.
- Novel-view-synthesis models can move the camera through an environment rather than only rotating around a central point, extending generative 3D from objects to scenes.
- The motion-masking loss allows training on in-the-wild, partially dynamic video, removing the need to manually filter or curate static scenes.
- ODIN improves LPIPS on Mip-NeRF 360 and Chamfer distance and IoU on Google Scanned Objects and a held-out 360-1M split relative to ZeroNVS and Zero-1-to-3.
- The released 360-1M dataset, with 363 million correspondences and poses, provides a resource for other multi-view and 3D learning tasks.
Reading between the lines
- The correspondence-mining recipe should transfer to any video source with wide fields of view or camera motion, not only 360° footage, potentially enlarging the pool of real-world multi-view data.
- Because Dust3R is used both to build the pseudo-ground truth and to reconstruct ODIN's generations, some of the reported geometric gains could reflect ODIN learning to produce images that Dust3R finds easy to align; an independent geometric check would separate true geometry from this feedback loop.
- The metric-scale anchoring inherits the bias of monocular depth estimation, so applications like robotics that need accurate absolute scale will likely require additional calibration.
- Extending motion masking from a soft filter to explicit modeling of moving objects would turn the static-scene assumption into a full 4D generator, which the paper itself identifies as the next step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces 360-1M, a dataset of over one million 360-degree videos from YouTube, together with a scalable pipeline for extracting multi-view frame correspondences using Dust3R-based pose estimation and confidence filtering, graph-based correspondence propagation, and metric scale anchoring via monocular depth. The authors train ODIN, a latent diffusion model for novel view synthesis conditioned on relative rotation and translation, with a motion-masking loss to handle dynamic content. They report improved LPIPS on DTU and MipNeRF360, improved Chamfer distance and IoU on Google Scanned Objects, and improved Chamfer distance and IoU on a held-out 360-1M set, claiming that ODIN can generate free-camera views of real-world scenes and enable 3D reconstruction from a single image.
Significance. If the claims hold, this is a significant contribution: 360-1M is the largest real-world multi-view dataset to date, the correspondence-search pipeline is practical and scalable, and ODIN demonstrates a new capability of long-range free-camera view synthesis from a single image. The motion-masking technique is a simple and useful idea for training on in-the-wild video. However, the quantitative evidence for the scene-geometry claim is currently undermined by an evaluation loop involving Dust3R and by missing statistical rigor. The dataset and code release, if provided, would be valuable resources.
major comments (4)
- [Section 6.3 / Table 4 (Appendix C)] The 360-1M reconstruction evaluation is circular. Dust3R is used (i) in Section 3.1 to select training correspondences by mean confidence threshold tau=4, (ii) in Section 6.3 to build pseudo-ground-truth point clouds from ground-truth views, and (iii) in Section 5.3 to reconstruct geometry from ODIN-generated images. Because ODIN is trained on pairs filtered by Dust3R confidence, it is incentivized to produce images that Dust3R can register; the reported Chamfer Distance and IoU then measure agreement with Dust3R's inductive biases rather than independent geometric accuracy. The authors should break this loop, for example by evaluating with an independent SfM/MVS pipeline such as COLMAP or against datasets with ground-truth 3D scans such as ScanNet or Matterport3D, and should report results with the reconstruction method fixed across all compared methods.
- [Section 6.3 / Table 4] There is a mismatch between the text and the table: Section 6.3 states 'We compare with ZeroNVS for scene reconstruction on a held-out set of 360-1M (Table 4 in Appendix),' but Table 4 is headed 'Comparison with Zero 1-to-3.' If the baseline is actually Zero-1-to-3, the comparison is not meaningful for scene-level reconstruction because that model is object-centric; if the baseline is ZeroNVS, the header must be corrected. This must be resolved before the scene-reconstruction claim can be assessed.
- [Section 6.2 / Tables 1-3] The non-circular evidence for the central scene-geometry claim is thin. Table 1 (DTU) shows only a 0.002 LPIPS improvement over ZeroNVS on an object-centric benchmark, Table 2 (MipNeRF360) reports image-quality metrics rather than geometry, and Table 3 (GSO) is object-level and shows only a small Chamfer improvement over Zero-1-to-3. The only scene-level geometric evaluation (Table 4) is the circular one discussed above. The paper would be substantially stronger if it added a scene-level geometric evaluation that does not use Dust3R at any stage, or if it explicitly qualified the claim to exclude scene geometry.
- [Tables 1-4] No error bars, confidence intervals, or significance tests are reported anywhere in the experimental section. Given the small margins (e.g., LPIPS 0.380 vs 0.378 in Table 1; Chamfer distance 0.0717 vs 0.0697 in Table 3), the reader cannot determine whether the improvements are statistically meaningful. The authors should report variances over evaluation scenes or runs, or at least per-scene results.
minor comments (6)
- [Abstract / throughout] The abstract and Section 1 use 'Odin' while the rest of the paper uses 'ODIN'; please standardize the capitalization.
- [Section 5.2, Eq. (2)] Equation (2) uses epsilon_theta but the text defines the denoiser as f_theta; please align the notation.
- [Section 6.1 / Table 2] The paper uses 'Mip-NeRF 360' in Section 6.1 and 'MipNeRF360' in Table 2; please use one consistent name.
- [Appendix C, Table 4] Table 4 is referenced as evaluating 360-1M, but the table caption does not state the evaluation set; add a clear caption that identifies the dataset and the baseline.
- [Section 4.2] The paper reports an average video length of 6.3 minutes while Figure 5 shows a long-tail distribution; clarify whether the mean is computed over all videos or only over videos that yielded correspondences.
- [Section 1 / Conclusion] The paper states that code, models, and dataset will be open-sourced, but no code or data access is provided in the submission; for reproducibility, please include a link or an appendix with dataset metadata details.
Circularity Check
360-1M geometry evaluation is partially circular: Dust3R filters training data, builds pseudo-ground-truth, and reconstructs ODIN's output.
-
fitted input called prediction
[Sections 3.1, 5.3, and 6.1 (360-1M reconstruction evaluation, Table 4)]
"We then pass all pairs within the time window to the Dust3r model [52] which outputs relative pose estimate, P and confidence map, C. ... filter out frames below threshold, τ = 4. ... For 360-1M we derive the pseudo-ground truth from a Dust3R model which is trained on all ground truth views of the scene given by the video. ... The 3D reconstructions for our model are created by generating images along trajectories then using Dust3r to reconstruct the scene."
Dust3R is the common estimator in every stage of the claimed scene-geometry result. Training correspondences are admitted only if Dust3R's confidence exceeds τ=4 (Sec. 3.1), so ODIN is optimized to produce frames that Dust3R can register. At evaluation, the pseudo-ground-truth point cloud is itself computed by Dust3R from the ground-truth views, and the reconstructed geometry scored is again Dust3R applied to ODIN's generated images (Sec. 6.1). Thus the 360-1M Chamfer/IoU comparison reduces to measuring how well ODIN's images match Dust3R's own prior, not independently verified scene geometry; the comparison against ZeroNVS is also uneven because ZeroNVS was not trained on Dust3R-filtered correspondences.
full rationale
The paper's core novelty—free-camera scene geometry from a single image—is quantitatively supported mainly by the held-out 360-1M reconstruction table. That table is a closed loop through Dust3R: the same model filters the training correspondences, produces the pseudo-ground-truth, and turns ODIN's generated images into the scored geometry. This is a genuine but partial circularity; it does not reduce the derivation to a pure tautology because ODIN still must synthesize images that Dust3R maps close to the pseudo-ground-truth, and independent benchmarks (DTU, GSO, MipNeRF360) provide some out-of-loop evidence. No load-bearing self-citation or ansatz-smuggling was found; citations to the authors' own Objaverse datasets are contextual. Overall score 6 reflects that the central scene-geometry claim rests on an evaluation loop even though other results are not circular.
Assumptions & free parameters
free parameters (6)
- Frame sampling rate r =
1 FPS
- Correspondence confidence threshold tau =
4.0
- Search window L =
20 frames
- Minimum translation =
0.25 m
- Motion masking weight lambda =
1.0
- Scale factor sigma =
per-correspondence optimized
assumptions (4)
- domain assumption Dust3R provides accurate relative pose and confidence maps between overlapping 360 frames
- domain assumption Depth Anything provides metric depth maps accurate enough for scale calibration
- domain assumption Rotating an equirectangular frame produces views that still contain enough overlap for Dust3R to find correspondences
- domain assumption Transitivity of correspondences holds for long-range pairs
Cite this review
Pith. "Pith review of From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos." pith.science (2026). https://pith.science/paper/YUHBC3YG
@misc{pith2026241207770,
author = {Pith},
title = {Pith review of: From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/YUHBC3YG}},
note = {Machine review of arXiv:2412.07770}
}
read the original abstract
Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and object-centric 3D datasets have shown to be effective in training models that have 3D understanding of objects. However, applying a similar approach to real-world objects and scenes is difficult due to a lack of large-scale data. Videos are a potential source for real-world 3D data, but finding diverse yet corresponding views of the same content has shown to be difficult at scale. Furthermore, standard videos come with fixed viewpoints, determined at the time of capture. This restricts the ability to access scenes from a variety of more diverse and potentially useful perspectives. We argue that large scale 360 videos can address these limitations to provide: scalable corresponding frames from diverse views. In this paper, we introduce 360-1M, a 360 video dataset, and a process for efficiently finding corresponding frames from diverse viewpoints at scale. We train our diffusion-based model, Odin, on 360-1M. Empowered by the largest real-world, multi-view dataset to date, Odin is able to freely generate novel views of real-world scenes. Unlike previous methods, Odin can move the camera through the environment, enabling the model to infer the geometry and layout of the scene. Additionally, we show improved performance on standard novel view synthesis and 3D reconstruction benchmarks.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
S. Agarwal, Y . Furukawa, N. Snavely, I. Simon, B. Curless, S. M. Seitz, and R. Szeliski. Building rome in a day. Communications of the ACM, 54:105–112, 2011. 3
work page 2011
-
[3]
J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021. 2 9
work page 2021
-
[4]
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman. Mip-NeRF 360: Unbounded anti-aliased neural radiance fields. In CVPR, 2022. 3, 8
work page 2022
-
[5]
E. Brachmann, J. Wynn, S. Chen, T. Cavallari, Á. Monszpart, D. Turmukhambetov, and V . A. Prisacariu. Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer. arXiv preprint arXiv:2404.14351, 2024. 2
arXiv 2024
-
[6]
E. R. Chan, K. Nagano, M. A. Chan, A. W. Bergman, J. J. Park, A. Levy, M. Aittala, S. D. Mello, T. Karras, and G. Wetzstein. GeNVS: Generative novel view synthesis with 3D-aware diffusion models. In ICCV, 2023. 3
work page 2023
-
[7]
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al. ShapeNet: An information-rich 3D model repository. arXiv preprint arXiv:1512.03012, 2015. 3
arXiv 2015
- [8]
Show all 68 references
-
[9]
Damen, H
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al. The epic-kitchens dataset: Collection, challenges and baselines. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11):4125–4141,
-
[10]
Deitke, D
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi. Objaverse: A universe of annotated 3D objects. arXiv preprint arXiv:2212.08051, 2022. 1, 3, 8
2022 arXiv
-
[11]
Deitke, R
M. Deitke, R. Liu, M. Wallingford, H. Ngo, O. Michel, A. Kusupati, A. Fan, C. Laforte, V . V oleti, S. Y . Gadre, E. VanderBilt, A. Kembhavi, C. V ondrick, G. Gkioxari, K. Ehsani, L. Schmidt, and A. Farhadi. Objaverse-XL: A universe of 10M+ 3D objects. arXiv preprint arXiv:2307.05663,
-
[12]
C. Deng, C. Jiang, C. R. Qi, X. Yan, Y . Zhou, L. Guibas, D. Anguelov, et al. NeRDi: Single-view NeRF synthesis with language-guided diffusion as general image priors. In CVPR, 2022. 8
2022
-
[13]
Downs, A
L. Downs, A. Francis, N. Koenig, B. Kinman, R. Hickman, K. Reymann, T. B. McHugh, and V . Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In 2022 International Conference on Robotics and Automation (ICRA), pages 2553–2560. IEEE,
2022
-
[14]
D. Fox, W. Burgard, F. Dellaert, and S. Thrun. Monte carlo localization: Efficient position estimation for mobile robots. AAAI, 1999. 1
1999
-
[15]
Geiger, P
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun. Vision meets robotics: The kitti dataset. International Journal of Robotics Research (IJRR), 2013. 3
2013
-
[16]
A. Jain, M. Tancik, and P. Abbeel. Putting NeRF on a diet: Semantically consistent few-shot view synthesis. In ICCV, 2021. 3, 8
2021
-
[17]
T. Jain, C. Lennan, Z. John, and D. Tran. Imagededup. https://github.com/idealo/ imagededup, 2019. 5
2019
-
[18]
Kerbl, G
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4):1–14, 2023. 3
2023
-
[19]
T. Kipf, G. F. Elsayed, A. Mahendran, A. Stone, S. Sabour, G. Heigold, R. Jonschkowski, A. Dosovitskiy, and K. Greff. Conditional object-centric learning from video. arXiv preprint arXiv:2111.12594, 2021. 1
2021 arXiv
-
[20]
Kolve, R
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, M. Deitke, K. Ehsani, D. Gordon, Y . Zhu, et al. Ai2-thor: An interactive 3d environment for visual ai.arXiv, 2017. 1
2017
-
[21]
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y . Liu, and T.-Y . Lin. Magic3D: High-resolution text-to-3D content creation. InCVPR, 2023. 3 10
2023
-
[22]
H. Lin. Robotic manipulation based on 3d vision: A survey. Proceedings of the 2020 International Conference on Pattern Recognition and Intelligent Systems , 2020. URL https://api.semanticscholar.org/CorpusID:221498989. 1
2020
-
[23]
A. Liu, R. Tucker, V . Jampani, A. Makadia, N. Snavely, and A. Kanazawa. Infinite nature: Perpetual view generation of natural scenes from a single image. In ICCV, 2021. 3, 8
2021
-
[24]
R. Liu, R. Wu, B. V . Hoorick, P. Tokmakov, S. Zakharov, and C. V ondrick. Zero-1-to-3: Zero-shot one image to 3D object. In CVPR, 2023. 1, 6, 7, 8, 9, 15
2023
-
[25]
W.-C. Ma, A. J. Yang, S. Wang, R. Urtasun, and A. Torralba. Virtual correspondence: Humans as a cue for extreme-view geometry. In CVPR, 2022. 1
2022
-
[26]
Mildenhall, P
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 2, 6
2020
-
[27]
Mottaghi, C
R. Mottaghi, C. Schenck, D. Fox, and A. Farhadi. See the glass half full: Reasoning about liquid containers, their volume and content. ICCV, 2017. 1
2017
-
[28]
Mur-Artal, J
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos. Orb-slam: a versatile and accurate monocular slam system. IEEE transactions on robotics, 31(5):1147–1163, 2015. 3
2015
-
[29]
Mur-Artal, J
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardós. ORB-SLAM: A versatile and accurate monocular SLAM system. IEEE Transactions on Robotics, 31(5):1147–1163, 2015. 5
2015
-
[30]
Nichol, H
A. Nichol, H. Jun, P. Dhariwal, P. Mishkin, and M. Chen. Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 9
2022 arXiv
-
[31]
Poole, A
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall. DreamFusion: Text-to-3D using 2D diffusion. In ICLR, 2022. 1, 3, 7
2022
-
[32]
Raistrick, L
A. Raistrick, L. Lipson, Z. Ma, L. Mei, M. Wang, Y . Zuo, K. Kayan, H. Wen, B. Han, Y . Wang, et al. Infinite photorealistic worlds using procedural generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12630–12641, 2023. 3
2023
-
[33]
Ramesh, P
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125, 2022. 5
2022 arXiv
-
[34]
Reizenstein, R
J. Reizenstein, R. Shapovalov, P. Henzler, L. Sbordone, P. Labatut, and D. Novotny. Common objects in 3D: Large-scale learning and evaluation of real-life 3D category reconstruction. In ICCV, 2021. 2, 3, 5, 6, 8
2021
-
[35]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. arXiv, 2021. 3
2021
-
[36]
M. S. Sajjadi, D. Duckworth, A. Mahendran, S. Van Steenkiste, F. Pavetic, M. Lucic, L. J. Guibas, K. Greff, and T. Kipf. Object scene representation transformer. Advances in Neural Information Processing Systems, 35:9512–9524, 2022. 1
2022
-
[37]
M. S. Sajjadi, H. Meyer, E. Pot, U. Bergmann, K. Greff, N. Radwan, S. V ora, M. Lu ˇci´c, D. Duckworth, A. Dosovitskiy, et al. Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. In Proceedings of the IEEE/CVF Conferen...
2022
-
[38]
M. S. Sajjadi, A. Mahendran, T. Kipf, E. Pot, D. Duckworth, M. Lu ˇci´c, and K. Greff. Rust: Latent neural scene representations from unposed imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17297–17306, 2023. 1
2023
-
[39]
Sargent, Z
K. Sargent, Z. Li, T. Shah, C. Herrmann, H.-X. Yu, Y . Zhang, E. R. Chan, D. Lagun, L. Fei-Fei, D. Sun, et al. Zeronvs: Zero-shot 360-degree view synthesis from a single real image. arXiv preprint arXiv:2310.17994, 2023. 3, 6, 7, 8
2023 arXiv
-
[40]
Sarlin, C
P.-E. Sarlin, C. Cadena, R. Siegwart, and M. Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12716–12725, 2019. 2, 3 11
2019
-
[41]
J. L. Schönberger and J.-M. Frahm. Structure-from-motion revisited. In CVPR, 2016. 3, 5
2016
-
[42]
J. L. Schönberger and J.-M. Frahm. Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2, 3
2016
-
[43]
J. L. Schonberger and J.-M. Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016. 3
2016
-
[44]
Y . Shi, P. Wang, J. Ye, M. Long, K. Li, and X. Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023. 1
2023 arXiv
-
[45]
Shriram, A
J. Shriram, A. Trevithick, L. Liu, and R. Ramamoorthi. Realmdreamer: Text-driven 3d scene generation with inpainting and depth diffusion. arXiv preprint arXiv:2404.07199, 2024. 3
2024 arXiv
-
[46]
Teed and J
Z. Teed and J. Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neural information processing systems, 34:16558–16569, 2021. 3
2021
-
[47]
Tewari, T
A. Tewari, T. Yin, G. Cazenavette, S. Rezchikov, J. Tenenbaum, F. Durand, B. Freeman, and V . Sitzmann. Diffusion with forward models: Solving stochastic inverse problems without direct supervision. Advances in Neural Information Processing Systems, 36, 2024. 3
2024
-
[48]
J. T. Todd. The visual perception of 3d shape. Trends in cognitive sciences, 8(3):115–121, 2004. 1
2004
-
[49]
Tschernezki, A
V . Tschernezki, A. Darkhalil, Z. Zhu, D. Fouhey, I. Laina, D. Larlus, D. Damen, and A. Vedaldi. Epic fields: Marrying 3d geometry and video understanding. Advances in Neural Information Processing Systems, 36, 2024. 3
2024
-
[50]
S. Ullman. The interpretation of structure from motion. Proceedings of the Royal Society of London. Series B. Biological Sciences, 203(1153):405–426, 1979. 3
1979
-
[51]
H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich. Score Jacobian chaining: Lifting pretrained 2D diffusion models for 3D generation. arXiv preprint arXiv:2212.00774, 2022. 9
2022 arXiv
-
[52]
S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud. Dust3r: Geometric 3d vision made easy. arXiv preprint arXiv:2312.14132, 2023. 3, 4, 7
2023 arXiv
-
[53]
Z. Wang, C. Lu, Y . Wang, F. Bao, C. Li, H. Su, and J. Zhu. ProlificDreamer: High-fidelity and di- verse text-to-3D generation with variational score distillation. arXiv preprint arXiv:2305.16213,
-
[54]
C.-Y . Wu, J. Johnson, J. Malik, C. Feichtenhofer, and G. Gkioxari. Multiview compressive coding for 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9065–9075, 2023. 9
2023
-
[55]
L. Wu, J. Y . Lee, A. Bhattad, Y .-X. Wang, and D. Forsyth. Diver: Real-time and accurate neural radiance fields with deterministic integration for volume rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16200–16209,
-
[56]
R. Wu, B. Mildenhall, P. Henzler, K. Park, R. Gao, D. Watson, P. P. Srinivasan, D. Verbin, J. T. Barron, B. Poole, and A. Holynski. Reconfusion: 3d reconstruction with diffusion priors. arXiv,
-
[57]
Xia, Z.-H
H. Xia, Z.-H. Lin, W.-C. Ma, and S. Wang. Video2game: Real-time, interactive, realistic and browser-compatible environment from a single video, 2024. 1
2024
-
[58]
Xiang, T
Y . Xiang, T. Schmidt, V . Narayanan, and D. Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv, 2017. 1
2017
-
[59]
D. Xu, Y . Jiang, P. Wang, Z. Fan, H. Shi, and Z. Wang. SinNeRF: Training neural radiance fields on complex scenes from a single image. In ECCV, 2022. 8
2022
-
[60]
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. arXiv preprint arXiv:2401.10891, 2024. 5 12
2024 arXiv
-
[61]
Z. Yang, Y . Chen, J. Wang, S. Manivasagam, W.-C. Ma, A. J. Yang, and R. Urtasun. Unisim: A neural closed-loop sensor simulator. CVPR, 2023. 1
2023
-
[62]
Yen-Chen, P
L. Yen-Chen, P. Florence, A. Zeng, J. T. Barron, Y . Du, W.-C. Ma, A. Simeonov, A. R. Garcia, and P. Isola. Mira: Mental imagery for robotic affordances, 2022. 1
2022
-
[63]
A. Yu, R. Li, M. Tancik, H. Li, R. Ng, and A. Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5752–5761, 2021. 2
2021
-
[64]
A. Yu, V . Ye, M. Tancik, and A. Kanazawa. pixelNeRF: Neural radiance fields from one or few images. In CVPR, 2021. 3, 8
2021
-
[65]
X. Yu, M. Xu, Y . Zhang, H. Liu, C. Ye, Y . Wu, Z. Yan, C. Zhu, Z. Xiong, T. Liang, et al. Mvimgnet: A large-scale dataset of multi-view images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9150–9161, 2023. 2, 3, 5, 6
2023
-
[66]
X. Zhao, A. Colburn, F. Ma, M. A. Bautista, J. M. Susskind, and A. G. Schwing. Is gener- alized dynamic novel view synthesis from monocular videos possible today? arXiv preprint arXiv:2310.08587, 2023. 1
2023 arXiv
-
[67]
B. Zhou, P. Krähenbühl, and V . Koltun. Does computer vision matter for action? Science Robotics, 2019. 1
2019
-
[68]
T. Zhou, R. Tucker, J. Flynn, G. Fyffe, and N. Snavely. Stereo magnification: Learning view synthesis using multiplane images. ACM Trans. Graph. (Proc. SIGGRAPH), 37, 2018. 3, 6, 8 13 A Dataset Statistics 103 104 105 106 Count 0 2000 4000 6000 8000 10000 12000Duration (s) Figu...
2018
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.