REVIEW 3 major objections 7 minor 1 cited by
MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
T0 review · 3 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read MegaSynth, a 700K-scene procedurally generated dataset, improves feed-forward 3D scene reconstruction by 1.2–1.8 dB PSNR and matches real-data training when used alone.
desk verdict The MegaSynth dataset is a serious engineering contribution, but the headline PSNR gain is confounded with the geometry loss Lloc, and the 'comparable to real data' claim is overstated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a procedural generator that builds scenes from non-semantic primitives and provides precise per-pixel depth and camera metadata, which in turn enables a Gaussian-location loss $L_{\text{loc}}$ that stabilizes training and improves geometry. The generator's controllability over complexity, camera baselines, lighting, and textures allows it to loosely match real-world distributions while eliminating the need for semantically valid scenes.
What would settle it
Train GS-LRM on MegaSynth renderings without the geometry loss $L_{\text{loc}}$ and compare to training on real data alone; if the PSNR advantage vanishes or reverses, the claim that the synthetic data itself drives reconstruction quality is falsified.
Extended reading notes
Core claim
MegaSynth is a procedurally generated dataset of 700K scenes, over 50 times larger than the real DL3DV dataset, in which scenes are composed from non-semantic shape primitives (cubes, spheres, cylinders, cones) placed within simple floor plans and lit by randomized ambient, sunlight, and luminous light sources. By removing semantic priors such as object affordances and scene composition, the authors make generation highly scalable and controllable. Training large reconstruction models such as GS-LRM and Long-LRM jointly with MegaSynth and real data, or pre-training on MegaSynth before fine-tuning on real data, improves reconstruction quality by 1.2–1.8 dB PSNR relative to real-data-only training across DL3DV, Hypersim, MipNeRF360, and Tanks & Temples. When trained exclusively on MegaSynth, the model performs comparably to the real-data-trained model, indicating that scene semantics are largely unnecessary for learning multi-view reconstruction.
Load-bearing premise
The gains attributed to the synthetic data could actually come from the extra depth supervision that only synthetic data can provide; if that supervision is removed, the improvement may disappear.
Editorial extensions
If this is right
- Joint training and pre-training with MegaSynth improve accuracy for two different LRM architectures, at multiple resolutions, and on both in-domain and out-of-domain test sets.
- Models pre-trained on MegaSynth generalize better to out-of-domain scenes than models trained only on real data.
- MegaSynth-only training yields geometry and image quality comparable to real-data-only training, suggesting that semantic priors are not necessary for scene-level reconstruction.
- MegaSynth also benefits monocular depth estimation when fine-tuning a pretrained model.
Reading between the lines
- If low-level geometry and appearance are all that matter for multi-view reconstruction, synthetic data generation could expand to unbounded scale with simple procedural rules, making data collection for scene-level models nearly free.
- The geometry loss enabled by synthetic depth might be the real driver of gains; a controlled study on real data with pseudo-depth supervision would separate the effects of data distribution from supervision.
- Extending the generator to include outdoor-specific structures, such as open skies and distant terrain, could close the remaining indoor-bias gap the paper observes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MegaSynth, a procedurally generated dataset of 700K non-semantic 3D scenes for training feed-forward large reconstruction models (LRMs). The generation pipeline composes primitive shapes with random textures and lighting within a controlled floor plan, then samples camera poses and renders RGB and depth supervision. The authors train GS-LRM and Long-LRM on MegaSynth combined with the real DL3DV dataset, using either joint training or pre-training/fine-tuning, and report consistent PSNR improvements of 1.2–1.8 dB over real-data-only baselines across in-domain and out-of-domain benchmarks. They further claim that models trained solely on MegaSynth perform comparably to real-data-trained models, and demonstrate an application to monocular depth estimation.
Significance. If the claims are substantiated, the work would show that non-semantic synthetic scenes can provide scalable and effective training data for scene-level reconstruction, potentially reducing dependence on expensive real-world capture. The dataset itself (700K scenes), the detailed ablations of data controllability and scale, and the released code/data are valuable contributions. However, the central attribution of the performance gains to the synthetic scene distribution is currently confounded with an auxiliary geometry loss, so the significance of the data-distribution claim is not yet established.
major comments (3)
- [Sec. 5.3, Eq. (6), Table 2] The headline 1.2–1.8 dB gain is confounded with the geometry loss Lloc. In Table 2, adding Lloc (row 2 vs row 1) raises the real-data-tuned Hypersim PSNR from 21.87 to 25.12 dB, a 3.25 dB gain larger than the entire reported improvement. More importantly, row (1) (MegaSynth without Lloc) reaches only 21.87 dB, below the DL3DV-only baseline of 23.89 dB in Table 4; that run is also marked as unstable (Fail. Iter. 57k). Thus, MegaSynth data alone does not improve over real data, and the improvement in Table 1 should be attributed to the combination of synthetic data and the geometry supervision it enables, not to the synthetic scene distribution per se. Please provide an ablation that isolates the data distribution from the supervision signal, e.g., MegaSynth with and without Lloc alongside a real-data baseline with a comparable auxiliary loss, and adjust the wording of the central claim accordingly.
- [Sec. 6.4, Table 4] The abstract claim that models trained solely on MegaSynth 'perform comparably' to real-data-trained models is not supported by the paper's own numbers: MegaSynth-only achieves 21.50 dB PSNR on Hypersim versus 23.89 dB for DL3DV-only, a 2.39 dB gap larger than the headline gains. Since the paper elsewhere treats differences on the order of 1 dB as meaningful, this gap is substantial. Please either report additional benchmarks where the gap is smaller, provide statistical significance, or soften the claim to reflect that synthetic-only training substantially underperforms real-data training.
- [Sec. 6.5, Table 5] The comparison with other synthetic datasets (Kubric and Front3D) does not state whether the geometry loss Lloc was enabled in those runs. Because Table 2 shows Lloc to be the dominant factor in the MegaSynth gains, the conclusion that 'realistic 3D assets or scene composition is not the guarantee for improving reconstruction quality' is not established unless the auxiliary supervision is held constant across all synthetic data conditions. Please clarify the loss configurations for each row, and ideally re-run the comparison with the same geometry loss enabled for Kubric and Front3D.
minor comments (7)
- [Table 5 caption] The dataset name is misspelled as 'Kurbic'; it should be 'Kubric'.
- [Sec. 7] The phrase 'trained sorely with MegaSynth' should read 'trained solely with MegaSynth'.
- [Table 2] The column headers for the three boolean factors (Control, LSloc, Scale) are not clearly labeled in the table as printed; please make the factors explicit in the caption.
- [Sec. 4] The term 'non-semantic' is used in several senses (no semantic classes, no affordances, no object relationships); please provide an operational definition at first use.
- [Table 7] The user study reports only averaged rankings; please include the number of participants, the rating instructions, and inter-rater agreement.
- [Table 6] The fine-tuning protocol for Depth Anything V2 on MegaSynth is not described; please provide details in the appendix.
- [Figure 1 caption] The generation time of '700K scenes in 3 days' should be accompanied by the compute environment used for rendering.
Circularity Check
No circular derivation: the claimed gains rest on external held-out benchmarks, hand-specified generator ranges, and independently evaluated mixed training.
full rationale
None of the paper's load-bearing steps reduces to its own inputs. The MegaSynth distribution is generated from hand-specified procedural ranges (Tables 8-14), not inferred from LRM outputs or from evaluation metrics, so the dataset is not defined in terms of the reconstruction results it is supposed to improve. The headline PSNR gains are measured on held-out external benchmarks (the DL3DV test split, Hypersim, and MipNeRF360/Tanks-and-Temples) that were not used to set the generator ranges, so the improvements are not forced by construction. Equation 6 (Lloc) is an additional geometry supervision term enabled by synthetic ground-truth depth; Table 2 shows that it contributes a large part of the gain, and this is a genuine confound for attributing the gain to the data distribution rather than to the extra supervision. However, a confound between two training-input variables is a correctness or ablation concern, not a circularity: the model's output is not being compared with a fitted version of itself. The only overlapping-author citation, LRM-Zero [79], is used for context about primitive-based object-level data generation; the scene-level synthesis, camera sampling, mixed training, and evaluation are implemented and tested independently, so no load-bearing claim is imported by self-citation. There is no uniqueness theorem, no fitted parameter renamed as a prediction, and no equation that is identical to its input by definition. The abstract's wording that models trained solely on MegaSynth perform comparably to real-data training is in quantitative tension with Table 4 (21.50 dB versus 23.89 dB on Hypersim), but that is a factual/claim discrepancy rather than circular reasoning. Score 0.
Assumptions & free parameters
free parameters (6)
- Scene dimension ranges =
size [17,30], height [10,15]
- Object box category parameters =
7 categories, sizes 2-8, counts 1-16
- Primitive composition probabilities =
cube/sphere/cylinder/cone 0.25 each; wireframe and torus added
- Material and lighting probabilities =
specular 0.2, glass IOR 1.4-1.6, sunlight 0.6, luminous objects 0.7
- Camera sampling distribution =
FoV 45-70 deg, 36 outer/12 inner cameras, baseline ranges
- Loss weights and training hyperparameters =
Lloc weight 0.4, perceptual weight 0.2, LR schedules, camera scale 1.1-1.6
assumptions (4)
- domain assumption Multi-view 3D reconstruction is largely a low-level task; explicit semantics are not required for training.
- domain assumption The hand-set ranges for scene size, object boxes, primitives, materials, and lighting produce data 'loosely aligned' with the real-world distribution, sufficient for transfer.
- domain assumption DL3DV camera poses and the COLMAP initialization used for the 3DGS baseline are accurate enough for fair evaluation.
- domain assumption Results on GS-LRM and Long-LRM generalize to other feed-forward scene reconstruction models.
Cite this review
Pith. "Pith review of MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data." pith.science (2026). https://pith.science/paper/2YCRMGU3
@misc{pith2026241214166,
author = {Pith},
title = {Pith review of: MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/2YCRMGU3}},
note = {Machine review of arXiv:2412.14166}
}
read the original abstract
We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real dataset DL3DV - dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth's properties for enhancing model capability, training stability, and generalization, as well as application to other tasks.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Vision as Unified Multimodal Generation
A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
arXiv 2023
-
[2]
Nemotron- 4 340b technical report
Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al. Nemotron- 4 340b technical report. arXiv preprint arXiv:2406.11704, 2024
arXiv 2024
-
[3]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736, 2022
2022
-
[4]
Sequential modeling enables scalable learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros. Sequential modeling enables scalable learning for large vision models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22861– 22872, 2024
2024
-
[5]
Frozen in time: A joint video and image encoder for end-to- end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to- end retrieval. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1728–1738, 2021
2021
-
[6]
Mip-nerf 360: Unbounded anti-aliased neural radiance fields
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, pages 5470–5479, 2022
2022
-
[7]
Depth pro: Sharp monocular metric depth in less than a second
Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain, Mar- cel Santos, Yichao Zhou, Stephan R Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073, 2024
arXiv 2024
-
[8]
Deep local shapes: Learning local sdf priors for detailed 3d reconstruction
Rohan Chabra, Jan E Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, and Richard Newcombe. Deep local shapes: Learning local sdf priors for detailed 3d reconstruction. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX 16, pages 608–625. Springer, 2020
2020
Show all 93 references
-
[9]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19457–19467, 2024
2024
-
[10]
Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo
Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF international conference on computer vision, pages 14124–14133, 2021
2021
-
[11]
Explicit correspondence match- ing for generalizable neural radiance fields
Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat- Jen Cham, and Jianfei Cai. Explicit correspondence match- ing for generalizable neural radiance fields. arXiv preprint arXiv:2304.12294, 2023
2023 arXiv
-
[12]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. arXiv preprint arXiv:2403.14627, 2024
2024 arXiv
-
[13]
Mvsplat360: Feed-forward 360 scene synthesis from sparse views
Yuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang, Andrea Vedaldi, Tat-Jen Cham, and Jianfei Cai. Mvsplat360: Feed-forward 360 scene synthesis from sparse views. arXiv preprint arXiv:2411.04924, 2024
2024 arXiv
-
[14]
The manhattan world assumption: Regularities in scene statistics which enable bayesian inference
James Coughlan and Alan L Yuille. The manhattan world assumption: Regularities in scene statistics which enable bayesian inference. Advances in Neural Information Process- ing Systems, 13, 2000
2000
-
[15]
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017
2017
-
[16]
Procthor: Large-scale embodied ai using procedural generation
Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation. Ad- vances in Neural Information Processing Systems, 35:5982...
2022
-
[17]
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2023
-
[18]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2010 arXiv
-
[19]
Learning to render novel views from wide-baseline stereo pairs
Yilun Du, Cameron Smith, Ayush Tewari, and Vincent Sitz- mann. Learning to render novel views from wide-baseline stereo pairs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4970–4980, 2023
2023
-
[20]
3d-front: 3d furnished rooms with layouts and semantics
Huan Fu, Bowen Cai, Lin Gao, Ling-Xiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, et al. 3d-front: 3d furnished rooms with layouts and semantics. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10933– 10...
2021
-
[21]
Accurate, dense, and robust multiview stereopsis
Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multiview stereopsis. IEEE transactions on pattern analysis and machine intelligence, 32(8):1362–1376, 2009
2009
-
[22]
Multi-view stereo for com- munity photo collections
Michael Goesele, Noah Snavely, Brian Curless, Hugues Hoppe, and Steven M Seitz. Multi-view stereo for com- munity photo collections. In 2007 IEEE 11th International Conference on Computer Vision, pages 1–8. IEEE, 2007
2007
-
[23]
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapra- gasam, Florian Golemo, Charles Herrmann, et al. Kubric: A scalable dataset generator. In Proceedings of the IEEE/CVF conference on computer vision and pattern re...
2022
-
[24]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023
2023 arXiv
-
[25]
Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering
Antoine Guédon and Vincent Lepetit. Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5354–5363, 2024
2024
-
[26]
Vfusion3d: Learning scalable 3d generative models from video diffusion models
Junlin Han, Filippos Kokkinos, and Philip Torr. Vfusion3d: Learning scalable 3d generative models from video diffusion models. In European Conference on Computer Vision, pages 333–350. Springer, 2025
2025
-
[27]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020
2020
-
[28]
Text2room: Extracting textured 3d meshes from 2d text-to-image models
Lukas Höllein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner. Text2room: Extracting textured 3d meshes from 2d text-to-image models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7909–7920, 2023
2023
-
[29]
Lrm: Large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400, 2023
2023 arXiv
-
[30]
2d gaussian splatting for geometrically accu- rate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2d gaussian splatting for geometrically accu- rate radiance fields. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024
2024
-
[31]
Leap: Liberate sparse-view 3d modeling from camera poses
Hanwen Jiang, Zhenyu Jiang, Yue Zhao, and Qixing Huang. Leap: Liberate sparse-view 3d modeling from camera poses. arXiv preprint arXiv:2310.01410, 2023
2023 arXiv
-
[32]
Real3d: Scaling up large reconstruction models with real- world images
Hanwen Jiang, Qixing Huang, and Georgios Pavlakos. Real3d: Scaling up large reconstruction models with real- world images. arXiv preprint arXiv:2406.08479, 2024
2024 arXiv
-
[33]
Few-view object reconstruction with unknown cate- gories and camera poses
Hanwen Jiang, Zhenyu Jiang, Kristen Grauman, and Yuke Zhu. Few-view object reconstruction with unknown cate- gories and camera poses. In 2024 International Conference on 3D Vision (3DV), pages 31–41. IEEE, 2024
2024
-
[34]
Cofie: Learning compact neural surface representa- tions with coordinate fields
Hanwen Jiang, Haitao Yang, Georgios Pavlakos, and Qixing Huang. Cofie: Learning compact neural surface representa- tions with coordinate fields. arXiv preprint arXiv:2406.03417, 2024
2024 arXiv
-
[35]
Lvsm: A large view synthesis model with minimal 3d inductive bias
Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, and Zexiang Xu. Lvsm: A large view synthesis model with minimal 3d inductive bias. arXiv preprint arXiv:2410.17242, 2024
2024 arXiv
-
[36]
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceed- ings, Part II 14, pages 694–711. Springer, 2016
2016
-
[37]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1, 2023
2023
-
[38]
Tanks and temples: Benchmarking large-scale scene reconstruction
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (ToG), 36(4): 1–13, 2017
2017
-
[39]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and Jérôme Revaud. Ground- ing image matching in 3d with mast3r. arXiv preprint arXiv:2406.09756, 2024
2024 arXiv
-
[40]
Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model
Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi. Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model. arXiv preprint arXiv:2311.06214, 2023
2023 arXiv
-
[41]
Megadepth: Learning single- view depth prediction from internet photos
Zhengqi Li and Noah Snavely. Megadepth: Learning single- view depth prediction from internet photos. In Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[42]
Learning to recon- struct shape and spatially-varying reflectance from a single image
Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Learning to recon- struct shape and spatially-varying reflectance from a single image. ACM Transactions on Graphics (TOG), 37(6):1–11, 2018
2018
-
[43]
Through the looking glass: Neural 3d reconstruction of trans- parent shapes
Zhengqin Li, Yu-Ying Yeh, and Manmohan Chandraker. Through the looking glass: Neural 3d reconstruction of trans- parent shapes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1262– 1271, 2020
2020
-
[44]
Neuralangelo: High-fidelity neural surface reconstruction
Zhaoshuo Li, Thomas Müller, Alex Evans, Russell H Tay- lor, Mathias Unberath, Ming-Yu Liu, and Chen-Hsuan Lin. Neuralangelo: High-fidelity neural surface reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8456–8465, 2023
2023
-
[45]
Jamba: A hy- brid transformer-mamba language model
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. Jamba: A hy- brid transformer-mamba language model. arXiv preprint arXiv:2403.19887, 2024
2024 arXiv
-
[46]
Genusd: 3d scene generation made easy
Tsung-Yi Lin, Chen-Hsuan Lin, Yin Cui, Yunhao Ge, Se- ungjun Nah, Arun Mallya, Zekun Hao, Yifan Ding, Hanzi Mao, Zhaoshuo Li, et al. Genusd: 3d scene generation made easy. In ACM SIGGRAPH 2024 Real-Time Live!, pages 1–2, 2024
2024
-
[47]
Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p...
2024
-
[48]
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and ...
2016
-
[49]
Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar
Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: Practical view synthe- sis with prescriptive sampling guidelines. ACM Transactions on Graphics (TOG), 2019
2019
-
[50]
Nerf: Representing scenes as neural radiance fields for view synthe- sis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthe- sis. Communications of the ACM, 65(1):99–106, 2021
2021
-
[51]
Differentiable blocks world: Qualita- tive 3d decomposition by rendering primitives
Tom Monnier, Jake Austin, Angjoo Kanazawa, Alexei Efros, and Mathieu Aubry. Differentiable blocks world: Qualita- tive 3d decomposition by rendering primitives. Advances in Neural Information Processing Systems , 36:5791–5807, 2023
2023
-
[52]
Robocasa: Large-scale simulation of everyday tasks for generalist robots
Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots. arXiv preprint arXiv:2406.02523, 2024
2024 arXiv
-
[53]
NVIDIA, Maciej Bala, Yin Cui, Yifan Ding, Yunhao Ge, Zekun Hao, Jon Hasselgren, Jacob Huffman, Jingyi Jin, J.P. Lewis, Zhaoshuo Li, Chen-Hsuan Lin, Yen-Chen Lin, Tsung- Yi Lin, Ming-Yu Liu, Alice Luo, Qianli Ma, Jacob Munkberg, Stella Shi, Fangyin Wei, Donglai Xiang, Jiashu Xu...
2024 arXiv
-
[54]
Infinite photorealistic worlds using procedural generation
Alexander Raistrick, Lahav Lipson, Zeyu Ma, Lingjie Mei, Mingzhe Wang, Yiming Zuo, Karhan Kayan, Hongyu Wen, Beining Han, Yihan Wang, et al. Infinite photorealistic worlds using procedural generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern rec...
2023
-
[55]
Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation
Alexander Raistrick, Lingjie Mei, Karhan Kayan, David Yan, Yiming Zuo, Beining Han, Hongyu Wen, Meenal Parakh, Stamatis Alexandropoulos, Lahav Lipson, et al. Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation. In Proceedings of the IEEE/CVF Conferenc...
2024
-
[56]
Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding
Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Ku- mar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In Proceed- ings of the IEEE/CVF international conference...
2021
-
[57]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...
2015
-
[58]
Scene representation transformer: Geometry-free novel view synthe- sis through set-latent scene representations
Mehdi SM Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani V ora, Mario Luˇci´c, Daniel Duckworth, Alexey Dosovitskiy, et al. Scene representation transformer: Geometry-free novel view synthe- sis through set-latent scene representations. In P...
2022
-
[59]
Single-shot neural relighting and svbrdf estimation
Shen Sang and Manmohan Chandraker. Single-shot neural relighting and svbrdf estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part XIX 16, pages 85–101. Springer, 2020
2020
-
[60]
Structure- from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm. Structure- from-motion revisited. In Proceedings of the IEEE confer- ence on computer vision and pattern recognition, pages 4104– 4113, 2016
2016
-
[61]
Laion-5b: An open large-scale dataset for training next gen- eration image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next gen- eration image-text models. Advances in Neural Infor...
2022
-
[62]
A comparison and eval- uation of multi-view stereo reconstruction algorithms
Steven M Seitz, Brian Curless, James Diebel, Daniel Scharstein, and Richard Szeliski. A comparison and eval- uation of multi-view stereo reconstruction algorithms. In 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06), pages 519–528. IEEE, 2006
2006
-
[63]
Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Im- plicit neural representations with periodic activation functions. ArXiv, abs/2006.09661, 2020
2006 arXiv
-
[64]
Flowmap: High-quality camera poses, in- trinsics, and depth via gradient descent
Cameron Smith, David Charatan, Ayush Tewari, and Vin- cent Sitzmann. Flowmap: High-quality camera poses, in- trinsics, and depth via gradient descent. arXiv preprint arXiv:2404.15259, 2024
2024 arXiv
-
[65]
Model- ing the world from internet photo collections
Noah Snavely, Steven M Seitz, and Richard Szeliski. Model- ing the world from internet photo collections. International journal of computer vision, 80:189–210, 2008
2008
-
[66]
Behav- ior: Benchmark for everyday household activities in virtual, interactive, and ecological environments
Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al. Behav- ior: Benchmark for everyday household activities in virtual, interactive, and ecological environments. In Co...
2022
-
[67]
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. arXiv preprint arXiv:2402.05054, 2024
2024 arXiv
-
[68]
Solving olympiad geometry without human demon- strations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. Solving olympiad geometry without human demon- strations. Nature, 625(7995):476–482, 2024
2024
-
[69]
Megascenes: Scene-level view synthesis at scale
Joseph Tung, Gene Chou, Ruojin Cai, Guandao Yang, Kai Zhang, Gordon Wetzstein, Bharath Hariharan, and Noah Snavely. Megascenes: Scene-level view synthesis at scale. arXiv preprint arXiv:2406.11819, 2024
2024 arXiv
-
[70]
Attention is all you need
A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017
2017
-
[71]
Vggsfm: Visual geometry grounded deep structure from motion
Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. Vggsfm: Visual geometry grounded deep structure from motion. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 21686–21697, 2024
2024
-
[72]
Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction
Peng Wang, Hao Tan, Sai Bi, Yinghao Xu, Fujun Luan, Kalyan Sunkavalli, Wenping Wang, Zexiang Xu, and Kai Zhang. Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction. arXiv preprint arXiv:2311.12024, 2023
2023 arXiv
-
[73]
Ibrnet: Learning multi-view image-based rendering
Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P Srini- vasan, Howard Zhou, Jonathan T Barron, Ricardo Martin- Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In Proceedings of the IEEE/CVF conference on computer vision and p...
2021
-
[74]
Dust3r: Geometric 3d vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 20697–20709, 2024
2024
-
[75]
Tartanair: A dataset to push the limits of visual slam
Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer. Tartanair: A dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4909–4916...
2020
-
[76]
Meshlrm: Large reconstruction model for high- quality mesh
Xinyue Wei, Kai Zhang, Sai Bi, Hao Tan, Fujun Luan, Valentin Deschaintre, Kalyan Sunkavalli, Hao Su, and Zex- iang Xu. Meshlrm: Large reconstruction model for high- quality mesh. arXiv preprint arXiv:2404.12385, 2024
2024 arXiv
-
[77]
Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion
Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Ro- main Brégier, Yohann Cabon, Vaibhav Arora, Leonid Ants- feld, Boris Chidlovskii, Gabriela Csurka, and Jérôme Revaud. Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion. Advances in Neural Info...
2022
-
[78]
latentsplat: Autoencoding variational gaus- sians for fast generalizable 3d reconstruction
Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. latentsplat: Autoencoding variational gaus- sians for fast generalizable 3d reconstruction. arXiv preprint arXiv:2403.16292, 2024
2024 arXiv
-
[79]
Lrm- zero: Training large reconstruction models with synthesized data
Desai Xie, Sai Bi, Zhixin Shu, Kai Zhang, Zexiang Xu, Yi Zhou, Sören Pirk, Arie Kaufman, Xin Sun, and Hao Tan. Lrm- zero: Training large reconstruction models with synthesized data. arXiv preprint arXiv:2406.09371, 2024
2024 arXiv
-
[80]
Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models. arXiv preprint arXiv:2404.07191 , 2024
2024 arXiv
-
[81]
Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model, 2023
Yinghao Xu, Hao Tan, Fujun Luan, Sai Bi, Peng Wang, Jiahao Li, Zifan Shi, Kalyan Sunkavalli, Gordon Wetzstein, Zexiang Xu, and Kai Zhang. Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model, 2023
2023
-
[82]
Deep image-based relighting from optimal sparse samples
Zexiang Xu, Kalyan Sunkavalli, Sunil Hadap, and Ravi Ra- mamoorthi. Deep image-based relighting from optimal sparse samples. ACM Transactions on Graphics (ToG), 37(4):1–13, 2018
2018
-
[83]
Deep view synthesis from sparse photometric images
Zexiang Xu, Sai Bi, Kalyan Sunkavalli, Sunil Hadap, Hao Su, and Ravi Ramamoorthi. Deep view synthesis from sparse photometric images. ACM Transactions on Graphics (ToG), 38(4):1–13, 2019
2019
-
[84]
Scene synthesis via uncertainty-driven attribute syn- chronization
Haitao Yang, Zaiwei Zhang, Siming Yan, Haibin Huang, Chongyang Ma, Yi Zheng, Chandrajit Bajaj, and Qixing Huang. Scene synthesis via uncertainty-driven attribute syn- chronization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5630–5640, 2021
2021
-
[85]
Depth anything v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything v2. arXiv preprint arXiv:2406.09414, 2024
2024 arXiv
-
[86]
Holodeck: Language guided gen- eration of 3d embodied ai environments
Yue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt, Al- varo Herrasti, Winson Han, Jiajun Wu, Nick Haber, Ranjay Krishna, Lingjie Liu, et al. Holodeck: Language guided gen- eration of 3d embodied ai environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...
2024
-
[87]
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4578–4587, 2021
2021
-
[88]
Gs-lrm: Large recon- struction model for 3d gaussian splatting
Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. Gs-lrm: Large recon- struction model for 3d gaussian splatting. arXiv preprint arXiv:2404.19702, 2024
2024 arXiv
-
[89]
Freeman, Kai Zhang, and Fujun Luan
Tianyuan Zhang, Zhengfei Kuang, Haian Jin, Zexiang Xu, Sai Bi, Hao Tan, He Zhang, Yiwei Hu, Milos Hasan, William T. Freeman, Kai Zhang, and Fujun Luan. Relitlrm: Generative relightable radiance for large reconstruction models, 2024
2024
-
[90]
Vision-and-language navigation today and tomorrow: A survey in the era of foundation models
Yue Zhang, Ziqiao Ma, Jialu Li, Yanyuan Qiao, Zun Wang, Joyce Chai, Qi Wu, Mohit Bansal, and Parisa Kordjamshidi. Vision-and-language navigation today and tomorrow: A survey in the era of foundation models. arXiv preprint arXiv:2407.07035, 2024
2024 arXiv
-
[91]
Pointodyssey: A large-scale synthetic dataset for long-term point tracking
Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 19855–19865, 2023
2023
-
[92]
Stereo magnification: Learning view synthesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. arXiv preprint arXiv:1805.09817, 2018
2018 arXiv
-
[93]
Long-lrm: Long-sequence large reconstruction model for wide-coverage gaussian splats
Chen Ziwen, Hao Tan, Kai Zhang, Sai Bi, Fujun Luan, Yicong Hong, Li Fuxin, and Zexiang Xu. Long-lrm: Long-sequence large reconstruction model for wide-coverage gaussian splats. arXiv preprint 2410.12781, 2024. A. MegaSynth Details In this section, we include more details of ou...
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.