Pith. sign in

REVIEW 3 major objections 7 minor 1 cited by

MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

T0 review · 3 major / 7 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read MegaSynth, a 700K-scene procedurally generated dataset, improves feed-forward 3D scene reconstruction by 1.2–1.8 dB PSNR and matches real-data training when used alone.

desk verdict The MegaSynth dataset is a serious engineering contribution, but the headline PSNR gain is confounded with the geometry loss Lloc, and the 'comparable to real data' claim is overstated. read the letter →

arxiv 2412.14166 v2 pith:2YCRMGU3 submitted 2024-12-18 cs.CV

classification cs.CV
keywords MegaSynthproceduralgenerationlargereconstructionmodel3Dscenenon-semanticprimitivesGaussiansplattingnovelviewsynthesissynthetictrainingdata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Feed-forward 3D scene reconstruction models improve when trained on procedurally generated scenes that contain no semantic information. The paper introduces MegaSynth, a 700K-scene synthetic dataset built from shape primitives and controlled lighting, and shows that adding it to real data raises novel-view PSNR by 1.2–1.8 dB across in-domain and out-of-domain benchmarks. Models trained only on MegaSynth perform comparably to ones trained on real scenes, supporting the claim that multi-view reconstruction depends mainly on low-level geometry and appearance rather than on scene semantics. The synthetic pipeline also supplies exact depth and camera metadata, enabling a geometry loss that stabilizes training and strengthens reconstruction.

What carries the argument

The central mechanism is a procedural generator that builds scenes from non-semantic primitives and provides precise per-pixel depth and camera metadata, which in turn enables a Gaussian-location loss $L_{\text{loc}}$ that stabilizes training and improves geometry. The generator's controllability over complexity, camera baselines, lighting, and textures allows it to loosely match real-world distributions while eliminating the need for semantically valid scenes.

What would settle it

Train GS-LRM on MegaSynth renderings without the geometry loss $L_{\text{loc}}$ and compare to training on real data alone; if the PSNR advantage vanishes or reverses, the claim that the synthetic data itself drives reconstruction quality is falsified.

Watch

Extended reading notes

Core claim

MegaSynth is a procedurally generated dataset of 700K scenes, over 50 times larger than the real DL3DV dataset, in which scenes are composed from non-semantic shape primitives (cubes, spheres, cylinders, cones) placed within simple floor plans and lit by randomized ambient, sunlight, and luminous light sources. By removing semantic priors such as object affordances and scene composition, the authors make generation highly scalable and controllable. Training large reconstruction models such as GS-LRM and Long-LRM jointly with MegaSynth and real data, or pre-training on MegaSynth before fine-tuning on real data, improves reconstruction quality by 1.2–1.8 dB PSNR relative to real-data-only training across DL3DV, Hypersim, MipNeRF360, and Tanks & Temples. When trained exclusively on MegaSynth, the model performs comparably to the real-data-trained model, indicating that scene semantics are largely unnecessary for learning multi-view reconstruction.

Load-bearing premise

The gains attributed to the synthetic data could actually come from the extra depth supervision that only synthetic data can provide; if that supervision is removed, the improvement may disappear.

Editorial extensions

If this is right

  • Joint training and pre-training with MegaSynth improve accuracy for two different LRM architectures, at multiple resolutions, and on both in-domain and out-of-domain test sets.
  • Models pre-trained on MegaSynth generalize better to out-of-domain scenes than models trained only on real data.
  • MegaSynth-only training yields geometry and image quality comparable to real-data-only training, suggesting that semantic priors are not necessary for scene-level reconstruction.
  • MegaSynth also benefits monocular depth estimation when fine-tuning a pretrained model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If low-level geometry and appearance are all that matter for multi-view reconstruction, synthetic data generation could expand to unbounded scale with simple procedural rules, making data collection for scene-level models nearly free.
  • The geometry loss enabled by synthetic depth might be the real driver of gains; a controlled study on real data with pseudo-depth supervision would separate the effects of data distribution from supervision.
  • Extending the generator to include outdoor-specific structures, such as open skies and distant terrain, could close the remaining indoor-bias gap the paper observes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper proposes MegaSynth, a procedurally generated dataset of 700K non-semantic 3D scenes for training feed-forward large reconstruction models (LRMs). The generation pipeline composes primitive shapes with random textures and lighting within a controlled floor plan, then samples camera poses and renders RGB and depth supervision. The authors train GS-LRM and Long-LRM on MegaSynth combined with the real DL3DV dataset, using either joint training or pre-training/fine-tuning, and report consistent PSNR improvements of 1.2–1.8 dB over real-data-only baselines across in-domain and out-of-domain benchmarks. They further claim that models trained solely on MegaSynth perform comparably to real-data-trained models, and demonstrate an application to monocular depth estimation.

Significance. If the claims are substantiated, the work would show that non-semantic synthetic scenes can provide scalable and effective training data for scene-level reconstruction, potentially reducing dependence on expensive real-world capture. The dataset itself (700K scenes), the detailed ablations of data controllability and scale, and the released code/data are valuable contributions. However, the central attribution of the performance gains to the synthetic scene distribution is currently confounded with an auxiliary geometry loss, so the significance of the data-distribution claim is not yet established.

major comments (3)
  1. [Sec. 5.3, Eq. (6), Table 2] The headline 1.2–1.8 dB gain is confounded with the geometry loss Lloc. In Table 2, adding Lloc (row 2 vs row 1) raises the real-data-tuned Hypersim PSNR from 21.87 to 25.12 dB, a 3.25 dB gain larger than the entire reported improvement. More importantly, row (1) (MegaSynth without Lloc) reaches only 21.87 dB, below the DL3DV-only baseline of 23.89 dB in Table 4; that run is also marked as unstable (Fail. Iter. 57k). Thus, MegaSynth data alone does not improve over real data, and the improvement in Table 1 should be attributed to the combination of synthetic data and the geometry supervision it enables, not to the synthetic scene distribution per se. Please provide an ablation that isolates the data distribution from the supervision signal, e.g., MegaSynth with and without Lloc alongside a real-data baseline with a comparable auxiliary loss, and adjust the wording of the central claim accordingly.
  2. [Sec. 6.4, Table 4] The abstract claim that models trained solely on MegaSynth 'perform comparably' to real-data-trained models is not supported by the paper's own numbers: MegaSynth-only achieves 21.50 dB PSNR on Hypersim versus 23.89 dB for DL3DV-only, a 2.39 dB gap larger than the headline gains. Since the paper elsewhere treats differences on the order of 1 dB as meaningful, this gap is substantial. Please either report additional benchmarks where the gap is smaller, provide statistical significance, or soften the claim to reflect that synthetic-only training substantially underperforms real-data training.
  3. [Sec. 6.5, Table 5] The comparison with other synthetic datasets (Kubric and Front3D) does not state whether the geometry loss Lloc was enabled in those runs. Because Table 2 shows Lloc to be the dominant factor in the MegaSynth gains, the conclusion that 'realistic 3D assets or scene composition is not the guarantee for improving reconstruction quality' is not established unless the auxiliary supervision is held constant across all synthetic data conditions. Please clarify the loss configurations for each row, and ideally re-run the comparison with the same geometry loss enabled for Kubric and Front3D.
minor comments (7)
  1. [Table 5 caption] The dataset name is misspelled as 'Kurbic'; it should be 'Kubric'.
  2. [Sec. 7] The phrase 'trained sorely with MegaSynth' should read 'trained solely with MegaSynth'.
  3. [Table 2] The column headers for the three boolean factors (Control, LSloc, Scale) are not clearly labeled in the table as printed; please make the factors explicit in the caption.
  4. [Sec. 4] The term 'non-semantic' is used in several senses (no semantic classes, no affordances, no object relationships); please provide an operational definition at first use.
  5. [Table 7] The user study reports only averaged rankings; please include the number of participants, the rating instructions, and inter-rater agreement.
  6. [Table 6] The fine-tuning protocol for Depth Anything V2 on MegaSynth is not described; please provide details in the appendix.
  7. [Figure 1 caption] The generation time of '700K scenes in 3 days' should be accompanied by the compute environment used for rendering.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the claimed gains rest on external held-out benchmarks, hand-specified generator ranges, and independently evaluated mixed training.

full rationale

None of the paper's load-bearing steps reduces to its own inputs. The MegaSynth distribution is generated from hand-specified procedural ranges (Tables 8-14), not inferred from LRM outputs or from evaluation metrics, so the dataset is not defined in terms of the reconstruction results it is supposed to improve. The headline PSNR gains are measured on held-out external benchmarks (the DL3DV test split, Hypersim, and MipNeRF360/Tanks-and-Temples) that were not used to set the generator ranges, so the improvements are not forced by construction. Equation 6 (Lloc) is an additional geometry supervision term enabled by synthetic ground-truth depth; Table 2 shows that it contributes a large part of the gain, and this is a genuine confound for attributing the gain to the data distribution rather than to the extra supervision. However, a confound between two training-input variables is a correctness or ablation concern, not a circularity: the model's output is not being compared with a fitted version of itself. The only overlapping-author citation, LRM-Zero [79], is used for context about primitive-based object-level data generation; the scene-level synthesis, camera sampling, mixed training, and evaluation are implemented and tested independently, so no load-bearing claim is imported by self-citation. There is no uniqueness theorem, no fitted parameter renamed as a prediction, and no equation that is identical to its input by definition. The abstract's wording that models trained solely on MegaSynth perform comparably to real-data training is in quantitative tension with Table 4 (21.50 dB versus 23.89 dB on Hypersim), but that is a factual/claim discrepancy rather than circular reasoning. Score 0.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The central claim rests on hand-set procedural generation parameters and on the assumption that the resulting non-semantic distribution is aligned enough with real scenes to transfer. Most parameters are design choices described in Appendix A; they are not fit to the evaluation metrics, but several directly control the difficulty and realism of the training data.

free parameters (6)
  • Scene dimension ranges = size [17,30], height [10,15]
    Hand-chosen spatial scale of generated rooms; not fit to any target metric, but determines camera and depth distributions.
  • Object box category parameters = 7 categories, sizes 2-8, counts 1-16
    Hand-set distributions over object placement and size to mimic real indoor clutter.
  • Primitive composition probabilities = cube/sphere/cylinder/cone 0.25 each; wireframe and torus added
    Choice of shape primitive frequencies; affects geometric difficulty and topology.
  • Material and lighting probabilities = specular 0.2, glass IOR 1.4-1.6, sunlight 0.6, luminous objects 0.7
    Hand-tuned to approximate real-world appearance and to increase complexity.
  • Camera sampling distribution = FoV 45-70 deg, 36 outer/12 inner cameras, baseline ranges
    Empirically chosen to stabilize training and match real capture geometry; directly affects view coverage.
  • Loss weights and training hyperparameters = Lloc weight 0.4, perceptual weight 0.2, LR schedules, camera scale 1.1-1.6
    Hand-set; the geometry loss weight directly controls the main supervision signal causing the reported gain.
assumptions (4)
  • domain assumption Multi-view 3D reconstruction is largely a low-level task; explicit semantics are not required for training.
    Motivates the non-semantic generator; the paper tests it but does not prove it independent of the confounded comparison.
  • domain assumption The hand-set ranges for scene size, object boxes, primitives, materials, and lighting produce data 'loosely aligned' with the real-world distribution, sufficient for transfer.
    No quantitative distributional matching is provided; transfer gains are the only evidence.
  • domain assumption DL3DV camera poses and the COLMAP initialization used for the 3DGS baseline are accurate enough for fair evaluation.
    Evaluation assumes annotations are clean; this is standard but unverified here.
  • domain assumption Results on GS-LRM and Long-LRM generalize to other feed-forward scene reconstruction models.
    Only two architectures are tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data." pith.science (2026). https://pith.science/paper/2YCRMGU3

@misc{pith2026241214166,
  author       = {Pith},
  title        = {Pith review of: MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2YCRMGU3}},
  note         = {Machine review of arXiv:2412.14166}
}
read the original abstract

We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real dataset DL3DV - dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth's properties for enhancing model capability, training stability, and generalization, as well as application to other tasks.

Figures

Figures reproduced from arXiv: 2412.14166 by the authors.

Figure 1
Figure 1. We introduce MegaSynth, a non-semantic synthesized dataset for training LRMs. MegaSynth benefits from its scalability and controllability, enabling us to generate 700K scenes in 3 days. We train LRMs with both the large-scale MegaSynth data and small-scale real data, improving LRMs for reconstructing wide-coverage scenes from dense-view images. Abstract We propose scaling up 3D scene reconstruction by train￾ing with… view at source ↗
Figure 2
Figure 2. MegaSynth generation pipeline. We first generate the scene floor plan, where each 3D box represents a shape and different colors represent different object types. We compose shape primitives into objects with geometry augmentations, where these objects further compose the scene. We randomize the texture and lighting, and generate random cameras for rendering. types of geometry. First, we add thin structures, such as… view at source ↗
Figure 3
Figure 3. Reconstruction visualization on the in-domain DL3DV data. The results are from Long-LRM at resolution 256. We present both indoor and outdoor results in the first and second rows, respectively. With our MegaSynth (denoted as ‘w. MegaSynth’), the model performs better on thin structures (e.g., bottom left), complicated lighting (e.g., top middle), and cluttered scenes (e.g., top right). Inf. Time In-Domain Out-of-Dom… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Reconstruction visualization on the out-of-domain data. The results are from Long-LRM at resolution 256. We include results for both Hypersim and MipNeRF360 are presented in the first and second rows, respectively. MegaSynth-only Training Real Data Tuning Data LS loc S…
Figure 5
Figure 5. Figure 5: Visual comparison between training with only MegaSynth and only real data. We include two failure cases of MegaSynth-only with failures highlighted [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Visualizaton of input views (first row of each example), render target view and ground-truth target views (last two rows of each [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Visualizaton of input views (first row of each example), render target view and ground-truth target views (last two rows of each [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision as Unified Multimodal Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.

Reference graph

Works this paper leans on

93 extracted references · 44 canonical work pages · cited by 1 Pith paper

  1. [1]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  2. [2]

    Nemotron- 4 340b technical report

    Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al. Nemotron- 4 340b technical report. arXiv preprint arXiv:2406.11704, 2024

  3. [3]

    Flamingo: a visual language model for few-shot learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736, 2022

  4. [4]

    Sequential modeling enables scalable learning for large vision models

    Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros. Sequential modeling enables scalable learning for large vision models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22861– 22872, 2024

  5. [5]

    Frozen in time: A joint video and image encoder for end-to- end retrieval

    Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to- end retrieval. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1728–1738, 2021

  6. [6]

    Mip-nerf 360: Unbounded anti-aliased neural radiance fields

    Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, pages 5470–5479, 2022

  7. [7]

    Depth pro: Sharp monocular metric depth in less than a second

    Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain, Mar- cel Santos, Yichao Zhou, Stephan R Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073, 2024

  8. [8]

    Deep local shapes: Learning local sdf priors for detailed 3d reconstruction

    Rohan Chabra, Jan E Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, and Richard Newcombe. Deep local shapes: Learning local sdf priors for detailed 3d reconstruction. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX 16, pages 608–625. Springer, 2020

Show all 93 references
  1. [9]

    pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction

    David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19457–19467, 2024

  2. [10]

    Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo

    Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF international conference on computer vision, pages 14124–14133, 2021

  3. [11]

    Explicit correspondence match- ing for generalizable neural radiance fields

    Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat- Jen Cham, and Jianfei Cai. Explicit correspondence match- ing for generalizable neural radiance fields. arXiv preprint arXiv:2304.12294, 2023

  4. [12]

    Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images

    Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. arXiv preprint arXiv:2403.14627, 2024

  5. [13]

    Mvsplat360: Feed-forward 360 scene synthesis from sparse views

    Yuedong Chen, Chuanxia Zheng, Haofei Xu, Bohan Zhuang, Andrea Vedaldi, Tat-Jen Cham, and Jianfei Cai. Mvsplat360: Feed-forward 360 scene synthesis from sparse views. arXiv preprint arXiv:2411.04924, 2024

  6. [14]

    The manhattan world assumption: Regularities in scene statistics which enable bayesian inference

    James Coughlan and Alan L Yuille. The manhattan world assumption: Regularities in scene statistics which enable bayesian inference. Advances in Neural Information Process- ing Systems, 13, 2000

  7. [15]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017

  8. [16]

    Procthor: Large-scale embodied ai using procedural generation

    Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation. Ad- vances in Neural Information Processing Systems, 35:5982...

  9. [17]

    Objaverse: A universe of annotated 3d objects

    Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...

  10. [18]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...

  11. [19]

    Learning to render novel views from wide-baseline stereo pairs

    Yilun Du, Cameron Smith, Ayush Tewari, and Vincent Sitz- mann. Learning to render novel views from wide-baseline stereo pairs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4970–4980, 2023

  12. [20]

    3d-front: 3d furnished rooms with layouts and semantics

    Huan Fu, Bowen Cai, Lin Gao, Ling-Xiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, et al. 3d-front: 3d furnished rooms with layouts and semantics. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10933– 10...

  13. [21]

    Accurate, dense, and robust multiview stereopsis

    Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multiview stereopsis. IEEE transactions on pattern analysis and machine intelligence, 32(8):1362–1376, 2009

  14. [22]

    Multi-view stereo for com- munity photo collections

    Michael Goesele, Noah Snavely, Brian Curless, Hugues Hoppe, and Steven M Seitz. Multi-view stereo for com- munity photo collections. In 2007 IEEE 11th International Conference on Computer Vision, pages 1–8. IEEE, 2007

  15. [23]

    Kubric: A scalable dataset generator

    Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapra- gasam, Florian Golemo, Charles Herrmann, et al. Kubric: A scalable dataset generator. In Proceedings of the IEEE/CVF conference on computer vision and pattern re...

  16. [24]

    Mamba: Linear-time sequence modeling with selective state spaces

    Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023

  17. [25]

    Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering

    Antoine Guédon and Vincent Lepetit. Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5354–5363, 2024

  18. [26]

    Vfusion3d: Learning scalable 3d generative models from video diffusion models

    Junlin Han, Filippos Kokkinos, and Philip Torr. Vfusion3d: Learning scalable 3d generative models from video diffusion models. In European Conference on Computer Vision, pages 333–350. Springer, 2025

  19. [27]

    Denoising diffu- sion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020

  20. [28]

    Text2room: Extracting textured 3d meshes from 2d text-to-image models

    Lukas Höllein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner. Text2room: Extracting textured 3d meshes from 2d text-to-image models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7909–7920, 2023

  21. [29]

    Lrm: Large reconstruction model for single image to 3d

    Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400, 2023

  22. [30]

    2d gaussian splatting for geometrically accu- rate radiance fields

    Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2d gaussian splatting for geometrically accu- rate radiance fields. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024

  23. [31]

    Leap: Liberate sparse-view 3d modeling from camera poses

    Hanwen Jiang, Zhenyu Jiang, Yue Zhao, and Qixing Huang. Leap: Liberate sparse-view 3d modeling from camera poses. arXiv preprint arXiv:2310.01410, 2023

  24. [32]

    Real3d: Scaling up large reconstruction models with real- world images

    Hanwen Jiang, Qixing Huang, and Georgios Pavlakos. Real3d: Scaling up large reconstruction models with real- world images. arXiv preprint arXiv:2406.08479, 2024

  25. [33]

    Few-view object reconstruction with unknown cate- gories and camera poses

    Hanwen Jiang, Zhenyu Jiang, Kristen Grauman, and Yuke Zhu. Few-view object reconstruction with unknown cate- gories and camera poses. In 2024 International Conference on 3D Vision (3DV), pages 31–41. IEEE, 2024

  26. [34]

    Cofie: Learning compact neural surface representa- tions with coordinate fields

    Hanwen Jiang, Haitao Yang, Georgios Pavlakos, and Qixing Huang. Cofie: Learning compact neural surface representa- tions with coordinate fields. arXiv preprint arXiv:2406.03417, 2024

  27. [35]

    Lvsm: A large view synthesis model with minimal 3d inductive bias

    Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, and Zexiang Xu. Lvsm: A large view synthesis model with minimal 3d inductive bias. arXiv preprint arXiv:2410.17242, 2024

  28. [36]

    Perceptual losses for real-time style transfer and super-resolution

    Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceed- ings, Part II 14, pages 694–711. Springer, 2016

  29. [37]

    3d gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1, 2023

  30. [38]

    Tanks and temples: Benchmarking large-scale scene reconstruction

    Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (ToG), 36(4): 1–13, 2017

  31. [39]

    Ground- ing image matching in 3d with mast3r

    Vincent Leroy, Yohann Cabon, and Jérôme Revaud. Ground- ing image matching in 3d with mast3r. arXiv preprint arXiv:2406.09756, 2024

  32. [40]

    Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model

    Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi. Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model. arXiv preprint arXiv:2311.06214, 2023

  33. [41]

    Megadepth: Learning single- view depth prediction from internet photos

    Zhengqi Li and Noah Snavely. Megadepth: Learning single- view depth prediction from internet photos. In Computer Vision and Pattern Recognition (CVPR), 2018

  34. [42]

    Learning to recon- struct shape and spatially-varying reflectance from a single image

    Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Learning to recon- struct shape and spatially-varying reflectance from a single image. ACM Transactions on Graphics (TOG), 37(6):1–11, 2018

  35. [43]

    Through the looking glass: Neural 3d reconstruction of trans- parent shapes

    Zhengqin Li, Yu-Ying Yeh, and Manmohan Chandraker. Through the looking glass: Neural 3d reconstruction of trans- parent shapes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1262– 1271, 2020

  36. [44]

    Neuralangelo: High-fidelity neural surface reconstruction

    Zhaoshuo Li, Thomas Müller, Alex Evans, Russell H Tay- lor, Mathias Unberath, Ming-Yu Liu, and Chen-Hsuan Lin. Neuralangelo: High-fidelity neural surface reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8456–8465, 2023

  37. [45]

    Jamba: A hy- brid transformer-mamba language model

    Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. Jamba: A hy- brid transformer-mamba language model. arXiv preprint arXiv:2403.19887, 2024

  38. [46]

    Genusd: 3d scene generation made easy

    Tsung-Yi Lin, Chen-Hsuan Lin, Yin Cui, Yunhao Ge, Se- ungjun Nah, Arun Mallya, Zekun Hao, Yifan Ding, Hanzi Mao, Zhaoshuo Li, et al. Genusd: 3d scene generation made easy. In ACM SIGGRAPH 2024 Real-Time Live!, pages 1–2, 2024

  39. [47]

    Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision

    Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p...

  40. [48]

    A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation

    Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and ...

  41. [49]

    Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar

    Ben Mildenhall, Pratul P. Srinivasan, Rodrigo Ortiz-Cayon, Nima Khademi Kalantari, Ravi Ramamoorthi, Ren Ng, and Abhishek Kar. Local light field fusion: Practical view synthe- sis with prescriptive sampling guidelines. ACM Transactions on Graphics (TOG), 2019

  42. [50]

    Nerf: Representing scenes as neural radiance fields for view synthe- sis

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthe- sis. Communications of the ACM, 65(1):99–106, 2021

  43. [51]

    Differentiable blocks world: Qualita- tive 3d decomposition by rendering primitives

    Tom Monnier, Jake Austin, Angjoo Kanazawa, Alexei Efros, and Mathieu Aubry. Differentiable blocks world: Qualita- tive 3d decomposition by rendering primitives. Advances in Neural Information Processing Systems , 36:5791–5807, 2023

  44. [52]

    Robocasa: Large-scale simulation of everyday tasks for generalist robots

    Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots. arXiv preprint arXiv:2406.02523, 2024

  45. [53]

    NVIDIA, Maciej Bala, Yin Cui, Yifan Ding, Yunhao Ge, Zekun Hao, Jon Hasselgren, Jacob Huffman, Jingyi Jin, J.P. Lewis, Zhaoshuo Li, Chen-Hsuan Lin, Yen-Chen Lin, Tsung- Yi Lin, Ming-Yu Liu, Alice Luo, Qianli Ma, Jacob Munkberg, Stella Shi, Fangyin Wei, Donglai Xiang, Jiashu Xu...

  46. [54]

    Infinite photorealistic worlds using procedural generation

    Alexander Raistrick, Lahav Lipson, Zeyu Ma, Lingjie Mei, Mingzhe Wang, Yiming Zuo, Karhan Kayan, Hongyu Wen, Beining Han, Yihan Wang, et al. Infinite photorealistic worlds using procedural generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern rec...

  47. [55]

    Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation

    Alexander Raistrick, Lingjie Mei, Karhan Kayan, David Yan, Yiming Zuo, Beining Han, Hongyu Wen, Meenal Parakh, Stamatis Alexandropoulos, Lahav Lipson, et al. Infinigen indoors: Photorealistic indoor scenes using procedural gener- ation. In Proceedings of the IEEE/CVF Conferenc...

  48. [56]

    Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding

    Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Ku- mar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In Proceed- ings of the IEEE/CVF international conference...

  49. [57]

    U- net: Convolutional networks for biomedical image segmen- tation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...

  50. [58]

    Scene representation transformer: Geometry-free novel view synthe- sis through set-latent scene representations

    Mehdi SM Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani V ora, Mario Luˇci´c, Daniel Duckworth, Alexey Dosovitskiy, et al. Scene representation transformer: Geometry-free novel view synthe- sis through set-latent scene representations. In P...

  51. [59]

    Single-shot neural relighting and svbrdf estimation

    Shen Sang and Manmohan Chandraker. Single-shot neural relighting and svbrdf estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part XIX 16, pages 85–101. Springer, 2020

  52. [60]

    Structure- from-motion revisited

    Johannes L Schonberger and Jan-Michael Frahm. Structure- from-motion revisited. In Proceedings of the IEEE confer- ence on computer vision and pattern recognition, pages 4104– 4113, 2016

  53. [61]

    Laion-5b: An open large-scale dataset for training next gen- eration image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next gen- eration image-text models. Advances in Neural Infor...

  54. [62]

    A comparison and eval- uation of multi-view stereo reconstruction algorithms

    Steven M Seitz, Brian Curless, James Diebel, Daniel Scharstein, and Richard Szeliski. A comparison and eval- uation of multi-view stereo reconstruction algorithms. In 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06), pages 519–528. IEEE, 2006

  55. [63]

    Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Im- plicit neural representations with periodic activation functions. ArXiv, abs/2006.09661, 2020

  56. [64]

    Flowmap: High-quality camera poses, in- trinsics, and depth via gradient descent

    Cameron Smith, David Charatan, Ayush Tewari, and Vin- cent Sitzmann. Flowmap: High-quality camera poses, in- trinsics, and depth via gradient descent. arXiv preprint arXiv:2404.15259, 2024

  57. [65]

    Model- ing the world from internet photo collections

    Noah Snavely, Steven M Seitz, and Richard Szeliski. Model- ing the world from internet photo collections. International journal of computer vision, 80:189–210, 2008

  58. [66]

    Behav- ior: Benchmark for everyday household activities in virtual, interactive, and ecological environments

    Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al. Behav- ior: Benchmark for everyday household activities in virtual, interactive, and ecological environments. In Co...

  59. [67]

    Lgm: Large multi-view gaussian model for high-resolution 3d content creation

    Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. arXiv preprint arXiv:2402.05054, 2024

  60. [68]

    Solving olympiad geometry without human demon- strations

    Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. Solving olympiad geometry without human demon- strations. Nature, 625(7995):476–482, 2024

  61. [69]

    Megascenes: Scene-level view synthesis at scale

    Joseph Tung, Gene Chou, Ruojin Cai, Guandao Yang, Kai Zhang, Gordon Wetzstein, Bharath Hariharan, and Noah Snavely. Megascenes: Scene-level view synthesis at scale. arXiv preprint arXiv:2406.11819, 2024

  62. [70]

    Attention is all you need

    A Vaswani. Attention is all you need. Advances in Neural Information Processing Systems, 2017

  63. [71]

    Vggsfm: Visual geometry grounded deep structure from motion

    Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. Vggsfm: Visual geometry grounded deep structure from motion. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 21686–21697, 2024

  64. [72]

    Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction

    Peng Wang, Hao Tan, Sai Bi, Yinghao Xu, Fujun Luan, Kalyan Sunkavalli, Wenping Wang, Zexiang Xu, and Kai Zhang. Pf-lrm: Pose-free large reconstruction model for joint pose and shape prediction. arXiv preprint arXiv:2311.12024, 2023

  65. [73]

    Ibrnet: Learning multi-view image-based rendering

    Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P Srini- vasan, Howard Zhou, Jonathan T Barron, Ricardo Martin- Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In Proceedings of the IEEE/CVF conference on computer vision and p...

  66. [74]

    Dust3r: Geometric 3d vision made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 20697–20709, 2024

  67. [75]

    Tartanair: A dataset to push the limits of visual slam

    Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor, and Sebastian Scherer. Tartanair: A dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 4909–4916...

  68. [76]

    Meshlrm: Large reconstruction model for high- quality mesh

    Xinyue Wei, Kai Zhang, Sai Bi, Hao Tan, Fujun Luan, Valentin Deschaintre, Kalyan Sunkavalli, Hao Su, and Zex- iang Xu. Meshlrm: Large reconstruction model for high- quality mesh. arXiv preprint arXiv:2404.12385, 2024

  69. [77]

    Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion

    Philippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Ro- main Brégier, Yohann Cabon, Vaibhav Arora, Leonid Ants- feld, Boris Chidlovskii, Gabriela Csurka, and Jérôme Revaud. Croco: Self-supervised pre-training for 3d vision tasks by cross-view completion. Advances in Neural Info...

  70. [78]

    latentsplat: Autoencoding variational gaus- sians for fast generalizable 3d reconstruction

    Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. latentsplat: Autoencoding variational gaus- sians for fast generalizable 3d reconstruction. arXiv preprint arXiv:2403.16292, 2024

  71. [79]

    Lrm- zero: Training large reconstruction models with synthesized data

    Desai Xie, Sai Bi, Zhixin Shu, Kai Zhang, Zexiang Xu, Yi Zhou, Sören Pirk, Arie Kaufman, Xin Sun, and Hao Tan. Lrm- zero: Training large reconstruction models with synthesized data. arXiv preprint arXiv:2406.09371, 2024

  72. [80]

    Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models

    Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models. arXiv preprint arXiv:2404.07191 , 2024

  73. [81]

    Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model, 2023

    Yinghao Xu, Hao Tan, Fujun Luan, Sai Bi, Peng Wang, Jiahao Li, Zifan Shi, Kalyan Sunkavalli, Gordon Wetzstein, Zexiang Xu, and Kai Zhang. Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model, 2023

  74. [82]

    Deep image-based relighting from optimal sparse samples

    Zexiang Xu, Kalyan Sunkavalli, Sunil Hadap, and Ravi Ra- mamoorthi. Deep image-based relighting from optimal sparse samples. ACM Transactions on Graphics (ToG), 37(4):1–13, 2018

  75. [83]

    Deep view synthesis from sparse photometric images

    Zexiang Xu, Sai Bi, Kalyan Sunkavalli, Sunil Hadap, Hao Su, and Ravi Ramamoorthi. Deep view synthesis from sparse photometric images. ACM Transactions on Graphics (ToG), 38(4):1–13, 2019

  76. [84]

    Scene synthesis via uncertainty-driven attribute syn- chronization

    Haitao Yang, Zaiwei Zhang, Siming Yan, Haibin Huang, Chongyang Ma, Yi Zheng, Chandrajit Bajaj, and Qixing Huang. Scene synthesis via uncertainty-driven attribute syn- chronization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5630–5640, 2021

  77. [85]

    Depth anything v2

    Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything v2. arXiv preprint arXiv:2406.09414, 2024

  78. [86]

    Holodeck: Language guided gen- eration of 3d embodied ai environments

    Yue Yang, Fan-Yun Sun, Luca Weihs, Eli VanderBilt, Al- varo Herrasti, Winson Han, Jiajun Wu, Nick Haber, Ranjay Krishna, Lingjie Liu, et al. Holodeck: Language guided gen- eration of 3d embodied ai environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...

  79. [87]

    pixelnerf: Neural radiance fields from one or few images

    Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4578–4587, 2021

  80. [88]

    Gs-lrm: Large recon- struction model for 3d gaussian splatting

    Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. Gs-lrm: Large recon- struction model for 3d gaussian splatting. arXiv preprint arXiv:2404.19702, 2024

  81. [89]

    Freeman, Kai Zhang, and Fujun Luan

    Tianyuan Zhang, Zhengfei Kuang, Haian Jin, Zexiang Xu, Sai Bi, Hao Tan, He Zhang, Yiwei Hu, Milos Hasan, William T. Freeman, Kai Zhang, and Fujun Luan. Relitlrm: Generative relightable radiance for large reconstruction models, 2024

  82. [90]

    Vision-and-language navigation today and tomorrow: A survey in the era of foundation models

    Yue Zhang, Ziqiao Ma, Jialu Li, Yanyuan Qiao, Zun Wang, Joyce Chai, Qi Wu, Mohit Bansal, and Parisa Kordjamshidi. Vision-and-language navigation today and tomorrow: A survey in the era of foundation models. arXiv preprint arXiv:2407.07035, 2024

  83. [91]

    Pointodyssey: A large-scale synthetic dataset for long-term point tracking

    Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wetzstein, and Leonidas J Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 19855–19865, 2023

  84. [92]

    Stereo magnification: Learning view synthesis using multiplane images

    Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. arXiv preprint arXiv:1805.09817, 2018

  85. [93]

    Long-lrm: Long-sequence large reconstruction model for wide-coverage gaussian splats

    Chen Ziwen, Hao Tan, Kai Zhang, Sai Bi, Fujun Luan, Yicong Hong, Li Fuxin, and Zexiang Xu. Long-lrm: Long-sequence large reconstruction model for wide-coverage gaussian splats. arXiv preprint 2410.12781, 2024. A. MegaSynth Details In this section, we include more details of ou...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.