Pith. sign in

REVIEW 4 major objections 6 minor 68 references

From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Training a diffusion model on over a million 360° videos lets it synthesize free-camera views of real scenes from a single image and reconstruct their 3D geometry.

desk verdict The 360-1M dataset and correspondence pipeline are a genuine contribution, but the headline scene-geometry result rests on an evaluation loop that uses Dust3R for both the pseudo-ground truth and the reconstruction. read the letter →

arxiv 2412.07770 v1 pith:YUHBC3YG submitted 2024-12-10 cs.CV cs.LG

classification cs.CVcs.LG
keywords 360-degreevideodatasetnovelviewsynthesisdiffusionmodels3Dreconstructionfromasingleimagemulti-viewcorrespondencescameraposeestimationmotionmaskingscenegeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that 360° video, mined at scale, can supply the multi-view training data that real-world 3D generation has been missing, and that a diffusion model trained on it can turn a single photograph into new views taken from freely moved camera positions. To that end the authors build 360-1M, a dataset of over a million 360° videos with hundreds of millions of frame correspondences and relative poses. Their model ODIN learns to generate those views and, because the views are geometrically consistent, the scene can then be reconstructed in 3D. The claim matters because prior generative models were effectively limited to rotating around a single object, whereas ODIN can move through a scene, which is what robotics, AR, and graphics applications need.

What carries the argument

The load-bearing machinery is a scalable correspondence-mining pipeline for 360° video. Frames are sampled at one per second, each equirectangular frame is projected into four views at 90° yaw increments, and pairs within a 20-frame window are fed to Dust3R, which returns relative poses and confidence maps; a mean-confidence threshold filters out non-overlapping pairs. A graph-propagation step then links frames that share a common correspondent, recovering long-range pairs without exhaustive search, and the dimensionless Dust3R poses are anchored to metric scale by fitting a scale factor against monocular depth from Depth Anything. The generative side is a latent diffusion U-Net whose conditioning includes rotation and translation and whose output is multiplied by a learned motion mask with an auxiliary loss that keeps the mask from collapsing to zero; at inference, views are sampled along a smooth trajectory and the resulting image set is fed back through Dust3R to build the 3D scene.

What would settle it

On a held-out set of 360° scenes, recover camera trajectories with an independent metric-scale system such as LiDAR-equipped scanning or COLMAP with known calibration, generate ODIN views from one frame, reconstruct them with Dust3R, and measure the alignment error between the reconstructed point cloud and the independent geometry; a large misalignment alongside small Dust3R reconstruction error would falsify the claim that ODIN's views are geometrically accurate.

Watch

Extended reading notes

Core claim

ODIN, a latent diffusion model conditioned on both camera rotation and translation, is trained on 360-1M and is argued to be the first model that can reasonably synthesize real-world 3D scenes and reconstruct their geometry from a single input image with free camera movement. On the DTU and Mip-NeRF 360 novel-view-synthesis benchmarks it improves LPIPS over prior single-image methods without fine-tuning, and on Google Scanned Objects and a held-out 360-1M split it improves Chamfer distance and volumetric IoU for 3D reconstruction. The enabling observation is that a 360° video contains, in principle, many views of the same content from different positions: by rotating the equirectangular projection of nearby frames, one can align them to overlapping views and recover their relative pose.

Load-bearing premise

The evaluation of 3D reconstruction quality assumes Dust3R's pose and pointmap estimates are accurate enough to serve as ground truth, even though the same model is used to find training correspondences and to reconstruct ODIN's output.

Editorial extensions

If this is right

  • A single image of a real scene becomes enough to generate a sequence of views that supports 3D reconstruction, without per-scene optimization or known camera poses.
  • Novel-view-synthesis models can move the camera through an environment rather than only rotating around a central point, extending generative 3D from objects to scenes.
  • The motion-masking loss allows training on in-the-wild, partially dynamic video, removing the need to manually filter or curate static scenes.
  • ODIN improves LPIPS on Mip-NeRF 360 and Chamfer distance and IoU on Google Scanned Objects and a held-out 360-1M split relative to ZeroNVS and Zero-1-to-3.
  • The released 360-1M dataset, with 363 million correspondences and poses, provides a resource for other multi-view and 3D learning tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The correspondence-mining recipe should transfer to any video source with wide fields of view or camera motion, not only 360° footage, potentially enlarging the pool of real-world multi-view data.
  • Because Dust3R is used both to build the pseudo-ground truth and to reconstruct ODIN's generations, some of the reported geometric gains could reflect ODIN learning to produce images that Dust3R finds easy to align; an independent geometric check would separate true geometry from this feedback loop.
  • The metric-scale anchoring inherits the bias of monocular depth estimation, so applications like robotics that need accurate absolute scale will likely require additional calibration.
  • Extending motion masking from a soft filter to explicit modeling of moving objects would turn the static-scene assumption into a full 4D generator, which the paper itself identifies as the next step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces 360-1M, a dataset of over one million 360-degree videos from YouTube, together with a scalable pipeline for extracting multi-view frame correspondences using Dust3R-based pose estimation and confidence filtering, graph-based correspondence propagation, and metric scale anchoring via monocular depth. The authors train ODIN, a latent diffusion model for novel view synthesis conditioned on relative rotation and translation, with a motion-masking loss to handle dynamic content. They report improved LPIPS on DTU and MipNeRF360, improved Chamfer distance and IoU on Google Scanned Objects, and improved Chamfer distance and IoU on a held-out 360-1M set, claiming that ODIN can generate free-camera views of real-world scenes and enable 3D reconstruction from a single image.

Significance. If the claims hold, this is a significant contribution: 360-1M is the largest real-world multi-view dataset to date, the correspondence-search pipeline is practical and scalable, and ODIN demonstrates a new capability of long-range free-camera view synthesis from a single image. The motion-masking technique is a simple and useful idea for training on in-the-wild video. However, the quantitative evidence for the scene-geometry claim is currently undermined by an evaluation loop involving Dust3R and by missing statistical rigor. The dataset and code release, if provided, would be valuable resources.

major comments (4)
  1. [Section 6.3 / Table 4 (Appendix C)] The 360-1M reconstruction evaluation is circular. Dust3R is used (i) in Section 3.1 to select training correspondences by mean confidence threshold tau=4, (ii) in Section 6.3 to build pseudo-ground-truth point clouds from ground-truth views, and (iii) in Section 5.3 to reconstruct geometry from ODIN-generated images. Because ODIN is trained on pairs filtered by Dust3R confidence, it is incentivized to produce images that Dust3R can register; the reported Chamfer Distance and IoU then measure agreement with Dust3R's inductive biases rather than independent geometric accuracy. The authors should break this loop, for example by evaluating with an independent SfM/MVS pipeline such as COLMAP or against datasets with ground-truth 3D scans such as ScanNet or Matterport3D, and should report results with the reconstruction method fixed across all compared methods.
  2. [Section 6.3 / Table 4] There is a mismatch between the text and the table: Section 6.3 states 'We compare with ZeroNVS for scene reconstruction on a held-out set of 360-1M (Table 4 in Appendix),' but Table 4 is headed 'Comparison with Zero 1-to-3.' If the baseline is actually Zero-1-to-3, the comparison is not meaningful for scene-level reconstruction because that model is object-centric; if the baseline is ZeroNVS, the header must be corrected. This must be resolved before the scene-reconstruction claim can be assessed.
  3. [Section 6.2 / Tables 1-3] The non-circular evidence for the central scene-geometry claim is thin. Table 1 (DTU) shows only a 0.002 LPIPS improvement over ZeroNVS on an object-centric benchmark, Table 2 (MipNeRF360) reports image-quality metrics rather than geometry, and Table 3 (GSO) is object-level and shows only a small Chamfer improvement over Zero-1-to-3. The only scene-level geometric evaluation (Table 4) is the circular one discussed above. The paper would be substantially stronger if it added a scene-level geometric evaluation that does not use Dust3R at any stage, or if it explicitly qualified the claim to exclude scene geometry.
  4. [Tables 1-4] No error bars, confidence intervals, or significance tests are reported anywhere in the experimental section. Given the small margins (e.g., LPIPS 0.380 vs 0.378 in Table 1; Chamfer distance 0.0717 vs 0.0697 in Table 3), the reader cannot determine whether the improvements are statistically meaningful. The authors should report variances over evaluation scenes or runs, or at least per-scene results.
minor comments (6)
  1. [Abstract / throughout] The abstract and Section 1 use 'Odin' while the rest of the paper uses 'ODIN'; please standardize the capitalization.
  2. [Section 5.2, Eq. (2)] Equation (2) uses epsilon_theta but the text defines the denoiser as f_theta; please align the notation.
  3. [Section 6.1 / Table 2] The paper uses 'Mip-NeRF 360' in Section 6.1 and 'MipNeRF360' in Table 2; please use one consistent name.
  4. [Appendix C, Table 4] Table 4 is referenced as evaluating 360-1M, but the table caption does not state the evaluation set; add a clear caption that identifies the dataset and the baseline.
  5. [Section 4.2] The paper reports an average video length of 6.3 minutes while Figure 5 shows a long-tail distribution; clarify whether the mean is computed over all videos or only over videos that yielded correspondences.
  6. [Section 1 / Conclusion] The paper states that code, models, and dataset will be open-sourced, but no code or data access is provided in the submission; for reproducibility, please include a link or an appendix with dataset metadata details.

Circularity Check

1 steps flagged · score 6.0 of 10

360-1M geometry evaluation is partially circular: Dust3R filters training data, builds pseudo-ground-truth, and reconstructs ODIN's output.

  1. fitted input called prediction [Sections 3.1, 5.3, and 6.1 (360-1M reconstruction evaluation, Table 4)]
    "We then pass all pairs within the time window to the Dust3r model [52] which outputs relative pose estimate, P and confidence map, C. ... filter out frames below threshold, τ = 4. ... For 360-1M we derive the pseudo-ground truth from a Dust3R model which is trained on all ground truth views of the scene given by the video. ... The 3D reconstructions for our model are created by generating images along trajectories then using Dust3r to reconstruct the scene."

    Dust3R is the common estimator in every stage of the claimed scene-geometry result. Training correspondences are admitted only if Dust3R's confidence exceeds τ=4 (Sec. 3.1), so ODIN is optimized to produce frames that Dust3R can register. At evaluation, the pseudo-ground-truth point cloud is itself computed by Dust3R from the ground-truth views, and the reconstructed geometry scored is again Dust3R applied to ODIN's generated images (Sec. 6.1). Thus the 360-1M Chamfer/IoU comparison reduces to measuring how well ODIN's images match Dust3R's own prior, not independently verified scene geometry; the comparison against ZeroNVS is also uneven because ZeroNVS was not trained on Dust3R-filtered correspondences.

full rationale

The paper's core novelty—free-camera scene geometry from a single image—is quantitatively supported mainly by the held-out 360-1M reconstruction table. That table is a closed loop through Dust3R: the same model filters the training correspondences, produces the pseudo-ground-truth, and turns ODIN's generated images into the scored geometry. This is a genuine but partial circularity; it does not reduce the derivation to a pure tautology because ODIN still must synthesize images that Dust3R maps close to the pseudo-ground-truth, and independent benchmarks (DTU, GSO, MipNeRF360) provide some out-of-loop evidence. No load-bearing self-citation or ansatz-smuggling was found; citations to the authors' own Objaverse datasets are contextual. Overall score 6 reflects that the central scene-geometry claim rests on an evaluation loop even though other results are not circular.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The central contribution is a data curation pipeline, not a physical theory. The main free parameters are engineering choices in that pipeline and a per-pair scale factor. The key assumptions are about the reliability of Dust3R and Depth Anything, which are also used in evaluation, creating a circularity concern.

free parameters (6)
  • Frame sampling rate r = 1 FPS
    Chosen as a balance between compute and accuracy; Table 5 shows higher FPS slightly improves metrics but at quadratic computational cost.
  • Correspondence confidence threshold tau = 4.0
    Used to filter Dust3R predictions; value chosen empirically with no sensitivity analysis reported.
  • Search window L = 20 frames
    Limits pairwise comparisons in the initial pass; chosen for computational tractability, not tuned.
  • Minimum translation = 0.25 m
    Discards frame pairs with small baseline because they provide little supervision; threshold selected arbitrarily.
  • Motion masking weight lambda = 1.0
    Controls the auxiliary loss that keeps the mask non-zero; Table 6 ablates values and shows best results at 1.0.
  • Scale factor sigma = per-correspondence optimized
    Optimized per pair via Eq. 1 to anchor Dust3R pointmaps to metric depth from Depth Anything. This is a fitted value, not a learned constant.
assumptions (4)
  • domain assumption Dust3R provides accurate relative pose and confidence maps between overlapping 360 frames
    The entire correspondence search and dataset construction rely on Dust3R's output (Section 3.1).
  • domain assumption Depth Anything provides metric depth maps accurate enough for scale calibration
    The scale factor sigma is fit against Depth Anything's depth output (Section 3.3).
  • domain assumption Rotating an equirectangular frame produces views that still contain enough overlap for Dust3R to find correspondences
    The pipeline projects four yaw-rotated views per frame and assumes these are valid multi-view inputs (Section 3.1).
  • domain assumption Transitivity of correspondences holds for long-range pairs
    Correspondence propagation assumes if two frames share a corresponding frame they also correspond (Section 3.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos." pith.science (2026). https://pith.science/paper/YUHBC3YG

@misc{pith2026241207770,
  author       = {Pith},
  title        = {Pith review of: From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YUHBC3YG}},
  note         = {Machine review of arXiv:2412.07770}
}
read the original abstract

Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and object-centric 3D datasets have shown to be effective in training models that have 3D understanding of objects. However, applying a similar approach to real-world objects and scenes is difficult due to a lack of large-scale data. Videos are a potential source for real-world 3D data, but finding diverse yet corresponding views of the same content has shown to be difficult at scale. Furthermore, standard videos come with fixed viewpoints, determined at the time of capture. This restricts the ability to access scenes from a variety of more diverse and potentially useful perspectives. We argue that large scale 360 videos can address these limitations to provide: scalable corresponding frames from diverse views. In this paper, we introduce 360-1M, a 360 video dataset, and a process for efficiently finding corresponding frames from diverse viewpoints at scale. We train our diffusion-based model, Odin, on 360-1M. Empowered by the largest real-world, multi-view dataset to date, Odin is able to freely generate novel views of real-world scenes. Unlike previous methods, Odin can move the camera through the environment, enabling the model to infer the geometry and layout of the scene. Additionally, we show improved performance on standard novel view synthesis and 3D reconstruction benchmarks.

Figures

Figures reproduced from arXiv: 2412.07770 by the authors.

Figure 1
Figure 1. By learning from the largest real-world, multi-view dataset to date, our model [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Left: An illustrative trajectory of standard video with the view point fixed at the time of [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Qualitative comparison of novel view synthesis on real-world scenes. The left and right [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Examples of generated 3D scenes using ODIN. The blue dot indicates the location of the input image and the red lines indicate the trajectory of the camera which generated the images. ODIN is capable of long-range generation of geometrically consistent images. In the bo…
Figure 6
Figure 6. Figure 6: Video categories’ distribution in 360-1M. 0 10K 40K 100K Count en es ja ko ru pt fr nl de it Language [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Video language distribution in 360-1M. B Correspondence Examples C 3D Reconstruction Evaluation [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Example of long-range correspondence found automatically within 360-1M. [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Example of long-range correspondence found within 360-1M. [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: General example of correspondences from MVImageNet. Previously the largest multi [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 53 canonical work pages

  1. [1]

    Aanæs, R

    H. Aanæs, R. R. Jensen, G. V ogiatzis, E. Tola, and A. B. Dahl. Large-scale data for multiple-view stereopsis. International Journal of Computer Vision, pages 1–16, 2016. 3, 8

  2. [2]

    Agarwal, Y

    S. Agarwal, Y . Furukawa, N. Snavely, I. Simon, B. Curless, S. M. Seitz, and R. Szeliski. Building rome in a day. Communications of the ACM, 54:105–112, 2011. 3

  3. [3]

    J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021. 2 9

  4. [4]

    J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman. Mip-NeRF 360: Unbounded anti-aliased neural radiance fields. In CVPR, 2022. 3, 8

  5. [5]

    Brachmann, J

    E. Brachmann, J. Wynn, S. Chen, T. Cavallari, Á. Monszpart, D. Turmukhambetov, and V . A. Prisacariu. Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer. arXiv preprint arXiv:2404.14351, 2024. 2

  6. [6]

    E. R. Chan, K. Nagano, M. A. Chan, A. W. Bergman, J. J. Park, A. Levy, M. Aittala, S. D. Mello, T. Karras, and G. Wetzstein. GeNVS: Generative novel view synthesis with 3D-aware diffusion models. In ICCV, 2023. 3

  7. [7]

    A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al. ShapeNet: An information-rich 3D model repository. arXiv preprint arXiv:1512.03012, 2015. 3

  8. [8]

    Chung, S

    J. Chung, S. Lee, H. Nam, J. Lee, and K. M. Lee. Luciddreamer: Domain-free generation of 3d gaussian splatting scenes. arXiv preprint arXiv:2311.13384, 2023. 3

Show all 68 references
  1. [9]

    Damen, H

    D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al. The epic-kitchens dataset: Collection, challenges and baselines. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11):4125–4141,

  2. [10]

    Deitke, D

    M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi. Objaverse: A universe of annotated 3D objects. arXiv preprint arXiv:2212.08051, 2022. 1, 3, 8

  3. [11]

    Deitke, R

    M. Deitke, R. Liu, M. Wallingford, H. Ngo, O. Michel, A. Kusupati, A. Fan, C. Laforte, V . V oleti, S. Y . Gadre, E. VanderBilt, A. Kembhavi, C. V ondrick, G. Gkioxari, K. Ehsani, L. Schmidt, and A. Farhadi. Objaverse-XL: A universe of 10M+ 3D objects. arXiv preprint arXiv:2307.05663,

  4. [12]

    C. Deng, C. Jiang, C. R. Qi, X. Yan, Y . Zhou, L. Guibas, D. Anguelov, et al. NeRDi: Single-view NeRF synthesis with language-guided diffusion as general image priors. In CVPR, 2022. 8

  5. [13]

    Downs, A

    L. Downs, A. Francis, N. Koenig, B. Kinman, R. Hickman, K. Reymann, T. B. McHugh, and V . Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In 2022 International Conference on Robotics and Automation (ICRA), pages 2553–2560. IEEE,

  6. [14]

    D. Fox, W. Burgard, F. Dellaert, and S. Thrun. Monte carlo localization: Efficient position estimation for mobile robots. AAAI, 1999. 1

  7. [15]

    Geiger, P

    A. Geiger, P. Lenz, C. Stiller, and R. Urtasun. Vision meets robotics: The kitti dataset. International Journal of Robotics Research (IJRR), 2013. 3

  8. [16]

    A. Jain, M. Tancik, and P. Abbeel. Putting NeRF on a diet: Semantically consistent few-shot view synthesis. In ICCV, 2021. 3, 8

  9. [17]

    T. Jain, C. Lennan, Z. John, and D. Tran. Imagededup. https://github.com/idealo/ imagededup, 2019. 5

  10. [18]

    Kerbl, G

    B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4):1–14, 2023. 3

  11. [19]

    T. Kipf, G. F. Elsayed, A. Mahendran, A. Stone, S. Sabour, G. Heigold, R. Jonschkowski, A. Dosovitskiy, and K. Greff. Conditional object-centric learning from video. arXiv preprint arXiv:2111.12594, 2021. 1

  12. [20]

    Kolve, R

    E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, M. Deitke, K. Ehsani, D. Gordon, Y . Zhu, et al. Ai2-thor: An interactive 3d environment for visual ai.arXiv, 2017. 1

  13. [21]

    C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y . Liu, and T.-Y . Lin. Magic3D: High-resolution text-to-3D content creation. InCVPR, 2023. 3 10

  14. [22]

    H. Lin. Robotic manipulation based on 3d vision: A survey. Proceedings of the 2020 International Conference on Pattern Recognition and Intelligent Systems , 2020. URL https://api.semanticscholar.org/CorpusID:221498989. 1

  15. [23]

    A. Liu, R. Tucker, V . Jampani, A. Makadia, N. Snavely, and A. Kanazawa. Infinite nature: Perpetual view generation of natural scenes from a single image. In ICCV, 2021. 3, 8

  16. [24]

    R. Liu, R. Wu, B. V . Hoorick, P. Tokmakov, S. Zakharov, and C. V ondrick. Zero-1-to-3: Zero-shot one image to 3D object. In CVPR, 2023. 1, 6, 7, 8, 9, 15

  17. [25]

    W.-C. Ma, A. J. Yang, S. Wang, R. Urtasun, and A. Torralba. Virtual correspondence: Humans as a cue for extreme-view geometry. In CVPR, 2022. 1

  18. [26]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 2, 6

  19. [27]

    Mottaghi, C

    R. Mottaghi, C. Schenck, D. Fox, and A. Farhadi. See the glass half full: Reasoning about liquid containers, their volume and content. ICCV, 2017. 1

  20. [28]

    Mur-Artal, J

    R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos. Orb-slam: a versatile and accurate monocular slam system. IEEE transactions on robotics, 31(5):1147–1163, 2015. 3

  21. [29]

    Mur-Artal, J

    R. Mur-Artal, J. M. M. Montiel, and J. D. Tardós. ORB-SLAM: A versatile and accurate monocular SLAM system. IEEE Transactions on Robotics, 31(5):1147–1163, 2015. 5

  22. [30]

    Nichol, H

    A. Nichol, H. Jun, P. Dhariwal, P. Mishkin, and M. Chen. Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 9

  23. [31]

    Poole, A

    B. Poole, A. Jain, J. T. Barron, and B. Mildenhall. DreamFusion: Text-to-3D using 2D diffusion. In ICLR, 2022. 1, 3, 7

  24. [32]

    Raistrick, L

    A. Raistrick, L. Lipson, Z. Ma, L. Mei, M. Wang, Y . Zuo, K. Kayan, H. Wen, B. Han, Y . Wang, et al. Infinite photorealistic worlds using procedural generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12630–12641, 2023. 3

  25. [33]

    Ramesh, P

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125, 2022. 5

  26. [34]

    Reizenstein, R

    J. Reizenstein, R. Shapovalov, P. Henzler, L. Sbordone, P. Labatut, and D. Novotny. Common objects in 3D: Large-scale learning and evaluation of real-life 3D category reconstruction. In ICCV, 2021. 2, 3, 5, 6, 8

  27. [35]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. arXiv, 2021. 3

  28. [36]

    M. S. Sajjadi, D. Duckworth, A. Mahendran, S. Van Steenkiste, F. Pavetic, M. Lucic, L. J. Guibas, K. Greff, and T. Kipf. Object scene representation transformer. Advances in Neural Information Processing Systems, 35:9512–9524, 2022. 1

  29. [37]

    M. S. Sajjadi, H. Meyer, E. Pot, U. Bergmann, K. Greff, N. Radwan, S. V ora, M. Lu ˇci´c, D. Duckworth, A. Dosovitskiy, et al. Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. In Proceedings of the IEEE/CVF Conferen...

  30. [38]

    M. S. Sajjadi, A. Mahendran, T. Kipf, E. Pot, D. Duckworth, M. Lu ˇci´c, and K. Greff. Rust: Latent neural scene representations from unposed imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17297–17306, 2023. 1

  31. [39]

    Sargent, Z

    K. Sargent, Z. Li, T. Shah, C. Herrmann, H.-X. Yu, Y . Zhang, E. R. Chan, D. Lagun, L. Fei-Fei, D. Sun, et al. Zeronvs: Zero-shot 360-degree view synthesis from a single real image. arXiv preprint arXiv:2310.17994, 2023. 3, 6, 7, 8

  32. [40]

    Sarlin, C

    P.-E. Sarlin, C. Cadena, R. Siegwart, and M. Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12716–12725, 2019. 2, 3 11

  33. [41]

    J. L. Schönberger and J.-M. Frahm. Structure-from-motion revisited. In CVPR, 2016. 3, 5

  34. [42]

    J. L. Schönberger and J.-M. Frahm. Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2, 3

  35. [43]

    J. L. Schonberger and J.-M. Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016. 3

  36. [44]

    Y . Shi, P. Wang, J. Ye, M. Long, K. Li, and X. Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023. 1

  37. [45]

    Shriram, A

    J. Shriram, A. Trevithick, L. Liu, and R. Ramamoorthi. Realmdreamer: Text-driven 3d scene generation with inpainting and depth diffusion. arXiv preprint arXiv:2404.07199, 2024. 3

  38. [46]

    Teed and J

    Z. Teed and J. Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neural information processing systems, 34:16558–16569, 2021. 3

  39. [47]

    Tewari, T

    A. Tewari, T. Yin, G. Cazenavette, S. Rezchikov, J. Tenenbaum, F. Durand, B. Freeman, and V . Sitzmann. Diffusion with forward models: Solving stochastic inverse problems without direct supervision. Advances in Neural Information Processing Systems, 36, 2024. 3

  40. [48]

    J. T. Todd. The visual perception of 3d shape. Trends in cognitive sciences, 8(3):115–121, 2004. 1

  41. [49]

    Tschernezki, A

    V . Tschernezki, A. Darkhalil, Z. Zhu, D. Fouhey, I. Laina, D. Larlus, D. Damen, and A. Vedaldi. Epic fields: Marrying 3d geometry and video understanding. Advances in Neural Information Processing Systems, 36, 2024. 3

  42. [50]

    S. Ullman. The interpretation of structure from motion. Proceedings of the Royal Society of London. Series B. Biological Sciences, 203(1153):405–426, 1979. 3

  43. [51]

    H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich. Score Jacobian chaining: Lifting pretrained 2D diffusion models for 3D generation. arXiv preprint arXiv:2212.00774, 2022. 9

  44. [52]

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud. Dust3r: Geometric 3d vision made easy. arXiv preprint arXiv:2312.14132, 2023. 3, 4, 7

  45. [53]

    Z. Wang, C. Lu, Y . Wang, F. Bao, C. Li, H. Su, and J. Zhu. ProlificDreamer: High-fidelity and di- verse text-to-3D generation with variational score distillation. arXiv preprint arXiv:2305.16213,

  46. [54]

    C.-Y . Wu, J. Johnson, J. Malik, C. Feichtenhofer, and G. Gkioxari. Multiview compressive coding for 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9065–9075, 2023. 9

  47. [55]

    L. Wu, J. Y . Lee, A. Bhattad, Y .-X. Wang, and D. Forsyth. Diver: Real-time and accurate neural radiance fields with deterministic integration for volume rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16200–16209,

  48. [56]

    R. Wu, B. Mildenhall, P. Henzler, K. Park, R. Gao, D. Watson, P. P. Srinivasan, D. Verbin, J. T. Barron, B. Poole, and A. Holynski. Reconfusion: 3d reconstruction with diffusion priors. arXiv,

  49. [57]

    Xia, Z.-H

    H. Xia, Z.-H. Lin, W.-C. Ma, and S. Wang. Video2game: Real-time, interactive, realistic and browser-compatible environment from a single video, 2024. 1

  50. [58]

    Xiang, T

    Y . Xiang, T. Schmidt, V . Narayanan, and D. Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv, 2017. 1

  51. [59]

    D. Xu, Y . Jiang, P. Wang, Z. Fan, H. Shi, and Z. Wang. SinNeRF: Training neural radiance fields on complex scenes from a single image. In ECCV, 2022. 8

  52. [60]

    L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. arXiv preprint arXiv:2401.10891, 2024. 5 12

  53. [61]

    Z. Yang, Y . Chen, J. Wang, S. Manivasagam, W.-C. Ma, A. J. Yang, and R. Urtasun. Unisim: A neural closed-loop sensor simulator. CVPR, 2023. 1

  54. [62]

    Yen-Chen, P

    L. Yen-Chen, P. Florence, A. Zeng, J. T. Barron, Y . Du, W.-C. Ma, A. Simeonov, A. R. Garcia, and P. Isola. Mira: Mental imagery for robotic affordances, 2022. 1

  55. [63]

    A. Yu, R. Li, M. Tancik, H. Li, R. Ng, and A. Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5752–5761, 2021. 2

  56. [64]

    A. Yu, V . Ye, M. Tancik, and A. Kanazawa. pixelNeRF: Neural radiance fields from one or few images. In CVPR, 2021. 3, 8

  57. [65]

    X. Yu, M. Xu, Y . Zhang, H. Liu, C. Ye, Y . Wu, Z. Yan, C. Zhu, Z. Xiong, T. Liang, et al. Mvimgnet: A large-scale dataset of multi-view images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9150–9161, 2023. 2, 3, 5, 6

  58. [66]

    X. Zhao, A. Colburn, F. Ma, M. A. Bautista, J. M. Susskind, and A. G. Schwing. Is gener- alized dynamic novel view synthesis from monocular videos possible today? arXiv preprint arXiv:2310.08587, 2023. 1

  59. [67]

    B. Zhou, P. Krähenbühl, and V . Koltun. Does computer vision matter for action? Science Robotics, 2019. 1

  60. [68]

    T. Zhou, R. Tucker, J. Flynn, G. Fyffe, and N. Snavely. Stereo magnification: Learning view synthesis using multiplane images. ACM Trans. Graph. (Proc. SIGGRAPH), 37, 2018. 3, 6, 8 13 A Dataset Statistics 103 104 105 106 Count 0 2000 4000 6000 8000 10000 12000Duration (s) Figu...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.