Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

Sparse-View 3D Reconstruction: Recent Advances and Open Challenges

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read One map unifies three method families for 3D recovery from sparse views.

desk verdict A broad, current survey of sparse-view 3D reconstruction with a useful taxonomy, but its central comparative/convergence claim is not backed by the data it presents; worth refereeing after substantial revisions. read the letter →

arxiv 2507.16406 v1 pith:PG7FVUWX submitted 2025-07-22 cs.CV

classification cs.CV
keywords sparse-view3DreconstructionneuralradiancefieldsGaussiansplattingdiffusionmodelsvisionfoundationpose-freecomputersurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish that the scattered landscape of sparse-view 3D reconstruction can be understood as three converging method families: geometry-based pipelines, neural implicit fields such as NeRF, and explicit Gaussian splatting, with diffusion and vision foundation models acting as a fourth hybrid layer. It argues that, although earlier reviews covered individual threads, none systematically analyzed how these paradigms converge in the low-image regime. The payoff of the unified view is practical: it turns dozens of seemingly unrelated fixes for floaters, background collapse, and pose ambiguity into a small set of design choices, and it exposes the trade-offs in accuracy, efficiency, and generalization that any real application must navigate. The survey also contends that the next bottleneck is not representation but priors, pointing to 3D-native generative models as the way toward real-time, pose-free reconstruction.

What carries the argument

The load-bearing device is the survey's four-part taxonomy of methods—geometry-based, neural implicit (NeRF), 3D Gaussian Splatting, and diffusion/vision-foundation-model hybrids—combined with the screening protocol shown in Figure 3 and a comparative evaluation using rendering metrics (PSNR, SSIM, LPIPS), geometric error metrics (Chamfer distance, F-score, pose errors), and runtime and input-scaling measures. The taxonomy does the argumentative work: it converts hundreds of papers into a small number of design moves, such as depth regularization, pseudo-view densification, co-training, and pose-free initialization, and it is what lets the survey claim a unified perspective that earlier reviews lacked.

What would settle it

Compile the set of peer-reviewed sparse-view 3D reconstruction papers from 2020–2025 indexed in a mainstream database and test whether every one falls into one of the four survey categories; finding a substantial method family that fits none of them, or locating a prior review that already systematically compared geometry-based, NeRF, and diffusion approaches, would overturn the survey's central claims of completeness and novelty.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that present-day sparse-view 3D reconstruction is no longer a bag of tricks but a convergent field: methods that start from classical geometry, from neural radiance fields, and from explicit 3D Gaussian primitives increasingly lean on the same crutches—monocular depth priors, diffusion-generated pseudo-views, and joint pose-geometry optimization—to fight the same failure modes. The survey's central claim is that organizing this convergence into a taxonomy, with comparative tables across datasets and metrics, is both possible and useful, and that doing so reveals a shift from per-scene optimization toward generalizable, pose-free, real-time systems. It further claims that the persistent "chicken-and-egg" problem of limited correspondences and unknown poses is being broken by generative and foundation-model priors rather than by better feature matching alone.

Load-bearing premise

The load-bearing premise is that the literature search behind Figure 3 was complete and representative; the survey does not report how many records were screened, excluded, or included, so if major 2020–2025 work is missing, the taxonomy and comparative conclusions could be incomplete.

Editorial extensions

If this is right

  • If the unified perspective is right, a practitioner can read off the trade-off from the method family: geometry-based methods win on pose accuracy in planar and indoor settings, NeRF variants win on consistency with enough regularization, 3DGS wins on speed, and diffusion hybrids buy generalization at the cost of compute.
  • The survey's reading implies that depth priors and pseudo-view generation are interchangeable substitutes: almost any sparse-view method in any family can be improved by injecting monocular depth or diffusion-generated observations, which is why the same regularization ideas recur across NeRF and 3DGS.
  • Pose-free reconstruction is moving from exception to default for sparse inputs: methods such as InstantSplat and CF-3DGS show that jointly optimizing pose and geometry can replace SfM initialization, which makes uncalibrated video capture a viable input source.
  • The paper's future directions predict that 2D-based diffusion guidance will hit a consistency wall, and that models trained directly on 3D data will be required to combine semantic, geometric, and material outputs.
  • Standardized benchmarks across DTU, LLFF, Tanks and Temples, CO3D, ScanNet, and RealEstate10K/DL3DV are sufficient for comparing current methods, but the survey notes that real-world robustness to specularity, transparency, dynamic content, and noise is not captured by those controlled settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • What is left implicit in the survey's own tables is that definitions of "sparse" vary widely, from 2–3 views in some papers to up to 12 in others; this suggests that cross-family quantitative comparisons are partly comparing different problems, and a shared view-count protocol would make the trade-off claims testable.
  • A testable extension follows from the paper's logic: if diffusion priors are truly filling in missing geometry, their benefit should shrink as input views grow, and the PSNR gap between diffusion hybrids and pure-geometry methods should close around 9–12 views; that curve can be measured directly from the type of data shown in Figure 6.
  • The unified taxonomy implies that next-generation "3D-native generative priors" could be evaluated by a simple criterion: whether they reduce floaters and Janus artifacts without additional per-scene optimization, a metric the survey discusses qualitatively but does not formalize.
  • The same pose-free, diffusion-guided machinery surveyed here connects directly to simultaneous localization and mapping and to dynamic scene reconstruction from handheld video, which the paper lists as future work but does not develop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript is a survey of sparse-view 3D reconstruction, organized into four methodological families: geometry-based methods, neural implicit representations (NeRF and variants), 3D Gaussian Splatting, and diffusion/vision-foundation-model-based hybrid methods. It claims to be the first study to systematically analyze the convergence of geometry-based, neural implicit, and generative approaches in this area, and to provide a unified cross-paradigm comparison of their accuracy, efficiency, and generalization. The paper covers roughly 150 papers from 2020–2025, proposes a taxonomy, lists benchmark datasets and metrics, and discusses current challenges and future directions.

Significance. If the advertised comparative analysis were delivered, the survey would be a useful map of a rapidly evolving field. Its strengths are the breadth of coverage, the explicit review protocol in Figure 3, the clear taxonomy, and the identification of live open problems such as pose-free reconstruction, domain generalization, and 3D-native generative priors. However, the central claim of a systematic convergence analysis and cross-paradigm comparison is not currently supported by the body of the paper: the comparison figures lack numerical provenance, no quantitative comparison under a common protocol is provided, and the term 'convergence' is never defined. The citation errors and the incomplete coverage table further weaken the survey's reliability as a reference. The manuscript is therefore a promising draft whose main claim must either be substantiated with real comparative data or appropriately weakened.

major comments (4)
  1. [Section VI.C; Figures 1 and 6] The abstract and Introduction claim a 'systematic analysis' and 'comparative evaluation' across geometry-based, neural implicit, and generative methods, but the body does not deliver this. Section VI lists datasets and metric definitions only; Tables II–IV report settings and runtimes as stated in original papers, and Table V is a check-mark coverage table. Figure 1 shows six 'normalized metrics' but does not name the methods compared, the normalization scheme, or the source of the scores. Figure 6 plots PSNR versus number of views without axis values, data points, or a data source. The term 'convergence' (of the three method families) is never defined, so the central claim is not testable. The authors should add a cross-paradigm comparison table with numbers drawn from a common evaluation protocol, or explicitly state that no such comparison is currently possible, and they must provide sources for all figures. If the claim is meant qualitatively, it should be reworded to match the narrative actually presented.
  2. [Introduction, second paragraph] The statement 'no previous study has systematically analyzed the convergence of geometry-based, neural implicit, and generative (diffusion-based) approaches' is asserted rather than demonstrated. The related-work discussion of [3]–[7] is a single sentence, with no comparison of the scopes, method families, or comparison methodologies of those surveys. This matters because [7] is already a review of 3D Gaussian Splatting for sparse-view reconstruction, and [6] is a 3DGS survey. To support the novelty claim, the authors should add a short table or paragraph contrasting the coverage of each prior survey and stating precisely what the present survey adds. Otherwise the claim should be softened to 'to our knowledge' and justified by that comparison.
  3. [Section III.C and Section II.B] There are citation errors that affect the survey's reliability as a reference. SparseNeuS is cited as [97], but reference [97] is K. Zhou's 'Neural surface reconstruction from sparse views using epipolar geometry,' which is also listed as [36]; the actual SparseNeuS paper (Long et al., CVPR 2022) is not in the reference list. In Section II.B, the sentence introducing EpiS cites [41], which is the Hartley–Zisserman computer vision textbook, not the EpiS paper. These references should be corrected against the primary sources.
  4. [Table V] Table V is presented as a comprehensive dataset-coverage summary, but many methods that are discussed as having benchmark evaluations have no checkmarks at all, including 3DFIRES [25], 'A Semantically Aware Multi-View 3D Reconstruction' [26], Dust to Tower [34], SparseAGS [86], SparseNeuS [97], and DNGaussian [13]. As a result, the table cannot support the survey's claim of comprehensive coverage. The authors should either complete the table with specific citations to evaluation sections in the primary papers, or replace it with a statement of coverage as reported where available and clearly mark unknown entries.
minor comments (5)
  1. [Figures 5 and 6] Figure 5 ('Distribution of sparse-view 3D reconstruction papers') and Figure 6 ('PSNR versus number of input views') have no numerical axis labels, data points, or source information; even as schematic illustrations they should be labeled or described as qualitative.
  2. [Section IV.A.1] There is a typo, 'oversized Gausssians,' in the description of LoopSparseGS; it should read 'Gaussians.'
  3. [Section VIII.A] There are missing spaces in the phrases 'asSp2360[161] andGenFusion[159]'; these should be formatted as 'as Sp2360 [161] and GenFusion [159].'
  4. [Table II] The caption defines 'sparse' as 3–10 views and 'few-shot' as 2–5 views, but some entries use 'Few-shot' for two views (e.g., Preface [93]) and others use 'Sparse' for ranges that overlap; the usage should be made consistent.
  5. [General] Several table entries contain inconsistent spacing and capitalization, such as 'V oxel' in Table II, 'V olumetric' in Table II, and 'Surfel Splatting' in Table III; a careful proofreading pass over all tables is recommended.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a literature survey without a derivation chain; its claims are external assertions about other work.

full rationale

This manuscript is a survey, not a derivation. It fits no parameters, introduces no equations, and makes no predictions that reduce to its inputs by construction. Its central novelty claim—that no previous study has systematically analyzed the convergence of geometry-based, neural implicit, and generative approaches in sparse-view 3D reconstruction—is an external literature claim, not a derived result. Every method description cites independent primary papers with their own benchmarks and reported settings. No self-citation is load-bearing: the authors do not rely on their own prior work, and no uniqueness theorem or ansatz is imported from the authors' previous papers. The skeptical concern that Figures 1 and 6 lack provenance and that no quantitative cross-paradigm comparison is provided is a rigor or completeness issue, not a circularity issue: missing evidence is not the same as equating outputs with inputs. Likewise, the unquantified screening protocol in Figure 3 weakens the coverage claim but does not make the survey circular. Under the hard rules, unsupported assertions and 'not standard consensus' concerns are correctness risks, not circular steps. Honest non-finding is therefore appropriate: the survey's organization and narrative do not derive from their own conclusions.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey introduces no free parameters or invented entities. It depends on two domain assumptions about completeness of the literature search and fidelity to the cited works.

assumptions (2)
  • domain assumption The literature search and screening protocol described in Figure 3 yields a complete and representative set of relevant papers.
    The survey's taxonomy and landscape conclusions depend on this assumption, but the paper does not report the number of records screened or the flow of studies, so completeness is unverifiable.
  • domain assumption The descriptions and reported performance numbers for the cited methods are accurate as summarized.
    The survey relies on the original papers without independent verification; any mischaracterization would propagate into the tables and discussion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sparse-View 3D Reconstruction: Recent Advances and Open Challenges." pith.science (2026). https://pith.science/paper/PG7FVUWX

@misc{pith2026250716406,
  author       = {Pith},
  title        = {Pith review of: Sparse-View 3D Reconstruction: Recent Advances and Open Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PG7FVUWX}},
  note         = {Machine review of arXiv:2507.16406}
}
read the original abstract

Sparse-view 3D reconstruction is essential for applications in which dense image acquisition is impractical, such as robotics, augmented/virtual reality (AR/VR), and autonomous systems. In these settings, minimal image overlap prevents reliable correspondence matching, causing traditional methods, such as structure-from-motion (SfM) and multiview stereo (MVS), to fail. This survey reviews the latest advances in neural implicit models (e.g., NeRF and its regularized versions), explicit point-cloud-based approaches (e.g., 3D Gaussian Splatting), and hybrid frameworks that leverage priors from diffusion and vision foundation models (VFMs).We analyze how geometric regularization, explicit shape modeling, and generative inference are used to mitigate artifacts such as floaters and pose ambiguities in sparse-view settings. Comparative results on standard benchmarks reveal key trade-offs between the reconstruction accuracy, efficiency, and generalization. Unlike previous reviews, our survey provides a unified perspective on geometry-based, neural implicit, and generative (diffusion-based) methods. We highlight the persistent challenges in domain generalization and pose-free reconstruction and outline future directions for developing 3D-native generative priors and achieving real-time, unconstrained sparse-view reconstruction.

Figures

Figures reproduced from arXiv: 2507.16406 by the authors.

Figure 1
Figure 1. Comparative performance of leading sparse-view 3D recon [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Structure of this survey: major topics and subtopics covered in sparse-view 3D reconstruction. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Methodological Review Protocol outlining the systematic process of literature identification, screening, data extraction, categorization, [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Taxonomy of sparse view 3D reconstruction methods by core categories. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Distribution of sparse-view 3D reconstruction papers [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Rendering quality (PSNR) versus number of input views for leading sparse-view 3D reconstruction methods. The plot compares NeRF [PITH_FULL_IMAGE:figures/full_fig_p021_6.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    A relightable Gaussian Splatting method for virtual production decomposes scenes into fixed appearance and variable lighting by parameterizing primitives to directly sample high-resolution background textures, enablin...

  2. MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Semantically enriched MASt3R correspondences plus a multi-attribute 3D consistency loss raise sparse-view ScanNet++ PSNR by >4.5 dB over Splatt3R and preserve quality under wide baselines.

Reference graph

Works this paper leans on

204 extracted references · 40 canonical work pages · cited by 2 Pith papers

  1. [97]

    Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry

    K. Zhou, “Neural surface reconstruction from sparse views using epipolar geometry,”arXiv preprint arXiv:2406.04301, 2024

  2. [3]

    BeyondPixels: A Comprehensive Review of the Evolution of Neural Radiance Fields

    A. Rabby and C. Zhang, “Beyondpixels: A comprehensive review of the evolution of neural radiance fields,”arXiv preprint arXiv:2306.03000, 2023

  3. [7]

    A review on 3d gaussian splatting for sparse view reconstruction,

    H. Liu, B. Liu, Q. Hu, P. Du, J. Li, Y . Bao, and F. Wang, “A review on 3d gaussian splatting for sparse view reconstruction,” Artif. Intell. Rev., vol. 58, p. 215, 2025

  4. [6]

    Recent advances in 3d gaussian splatting,

    T. Wu, Y .-J. Yuan, L.-X. Zhang, J. Yang, Y .-P. Cao, L.-Q. Yan, and L. Gao, “Recent advances in 3d gaussian splatting,” Computational Visual Media, vol. 10, no. 4, pp. 613–642, 2024

  5. [41]

    Hartley and A

    R. Hartley and A. Zisserman,Multiple View Geometry in Computer Vision, 2nd ed. Cambridge University Press, 2004

  6. [25]

    3dfires: Few image 3d reconstruction for scenes with hidden surfaces,

    L. Jin, N. Kulkarni, and D. F. Fouhey, “3dfires: Few image 3d reconstruction for scenes with hidden surfaces,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 9742–9751

  7. [26]

    A Semantically Aware Multi-View 3D Reconstruction Method for Urban Applications,

    R. Wei, H. Pei, D. Wu, C. Zeng, X. Ai, and H. Duan, “A Semantically Aware Multi-View 3D Reconstruction Method for Urban Applications,”Applied Sciences, vol. 14, 2024

  8. [34]

    Dust to tower: Coarse-to-fine photo-realistic scene reconstruction from sparse uncalibrated images,

    X. Cai, Y . Wang, Z. Fan, D. Haoran, S. Wang, W. Li, D. Li, L. Luo, M. Wang, and J. Xu, “Dust to tower: Coarse-to-fine photo-realistic scene reconstruction from sparse uncalibrated images,”arXiv preprint arXiv:2412.19518, 2024

  9. [86]

    Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis

    Q. Zhao and S. Tulsiani, “Sparse-view pose estimation and reconstruction via analysis by generative synthesis,”arXiv preprint arXiv:2412.03570, 2024

  10. [13]

    DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization,

    J. Li, J. Zhang, X. Bai, J. Zheng, X. Ning, J. Zhou, and L. Gu, “DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, W A, USA: IEEE, 2024

Show all 204 references
  1. [1]

    Structure-from-motion revisited,

    J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion revisited,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104–4113

  2. [2]

    Pixelwise view selection for unstructured multi-view stereo,

    J. L. Sch ¨onberger, E. Zheng, M. Pollefeys, and J.-M. Frahm, “Pixelwise view selection for unstructured multi-view stereo,” inEuropean Conference on Computer Vision (ECCV), 2016, pp. 501–518

  3. [4]

    Deep-learning-based 3-d surface reconstruction—a survey,

    A. Farshian, M. G ¨otz, G. Cavallaro, C. Debus, M. Nießner, J. A. Benediktsson, and A. Streit, “Deep-learning-based 3-d surface reconstruction—a survey,”Proceedings of the IEEE, vol. 111, pp. 1464–1501, 2023

  4. [5]

    Multi- view 3d reconstruction based on deep learning: A survey and comparison of methods,

    J. Wu, O. Wyman, Y . Tang, D. Pasini, and W. Wang, “Multi- view 3d reconstruction based on deep learning: A survey and comparison of methods,”Neurocomputing, vol. 582, p. 127553, 2024

  5. [8]

    Point Cloud Densification for 3D Gaussian Splatting from Sparse Input Views,

    K.-C. Chan, J. Xiao, H. L. Goshu, and K.-M. Lam, “Point Cloud Densification for 3D Gaussian Splatting from Sparse Input Views,” inProceedings of the 32nd ACM International Conference on Multimedia. Melbourne VIC Australia: ACM, 2024

  6. [9]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” inEuropean Conference on Computer Vision (ECCV), 2020, pp. 405–421

  7. [10]

    3d gaussian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” in ACM SIGGRAPH 2023 Conference Proceedings. ACM, 2023

  8. [11]

    SparseNeRF: Dis- tilling Depth Ranking for Few-shot Novel View Synthesis,

    G. Wang, Z. Chen, C. C. Loy, and Z. Liu, “SparseNeRF: Dis- tilling Depth Ranking for Few-shot Novel View Synthesis,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023

  9. [12]

    Zerorf: Fast sparse view 360° reconstruction with zero pretraining,

    R. Shi, X. Wei, C. Wang, and H. Su, “Zerorf: Fast sparse view 360° reconstruction with zero pretraining,”arXiv preprint arXiv:2312.09249, 2023

  10. [14]

    Sparsecraft: Few- shot neural reconstruction through stereopsis guided geometric linearization,

    M. Younes, A. Ouasfi, and A. Boukhayma, “Sparsecraft: Few- shot neural reconstruction through stereopsis guided geometric linearization,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 37–56

  11. [15]

    Pointgs: Point attention-aware sparse view synthesis with gaussian splat- ting,

    L. Xiang, H. Zheng, Y . Huang, Q. Yang, and H. Yin, “Pointgs: Point attention-aware sparse view synthesis with gaussian splat- ting,”arXiv preprint arXiv:2506.10335, 2025

  12. [16]

    Instantsplat: Unbounded sparse-view pose-free gaussian splatting in 40 sec- onds,

    Z. Fan, W. Cong, K. Wen, K. Wang, J. Zhang, X. Ding, D. Xu, B. Ivanovic, M. Pavone, G. Pavlakoset al., “Instantsplat: Unbounded sparse-view pose-free gaussian splatting in 40 sec- onds,”arXiv preprint arXiv:2403.20309, vol. 2, no. 3, p. 4, 2024

  13. [17]

    COLMAP-Free 3D Gaussian Splatting,

    Y . Fu, S. Liu, A. Kulkarni, J. Kautz, A. A. Efros, and X. Wang, “COLMAP-Free 3D Gaussian Splatting,”2024 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pp. 20 796–20 805, 2023

  14. [18]

    Cor-gs: sparse-view 3d gaussian splatting via co- regularization,

    J. Zhang, J. Li, X. Yu, L. Huang, L. Gu, J. Zheng, and X. Bai, “Cor-gs: sparse-view 3d gaussian splatting via co- regularization,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 335–352

  15. [19]

    Fsgs: Real-time few-shot view synthesis using gaussian splatting,

    Z. Zhu, Z. Fan, Y . Jiang, and Z. Wang, “Fsgs: Real-time few-shot view synthesis using gaussian splatting,” inEuropean conference on computer vision. Springer, 2024, pp. 145–163

  16. [20]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” inAdvances in Neural Information Processing Systems (NeurIPS), 2020

  17. [21]

    Cat3d: Create anything in 3d with multi-view diffusion models,

    R. Gao, A. Holynski, P. Henzler, A. Brussee, R. Martin Brualla, P. Srinivasan, J. Barron, and B. Poole, “Cat3d: Create anything in 3d with multi-view diffusion models,”Advances in Neural Information Processing Systems, vol. 37, pp. 75 468–75 494, 2024

  18. [22]

    Gaussian splatting decoder for 3d-aware generative adversarial networks,

    F. Barthel, A. Beckmann, W. Morgenstern, A. Hilsmann, and P. Eisert, “Gaussian splatting decoder for 3d-aware generative adversarial networks,” inProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2024, pp. 7963–7972

  19. [23]

    Mv-dust3r+: Single-stage scene reconstruction from sparse views in 2 seconds,

    Z. Tang, Y . Fan, D. Wang, H. Xu, R. Ranjan, A. Schwing, and Z. Yan, “Mv-dust3r+: Single-stage scene reconstruction from sparse views in 2 seconds,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 5283– 5293

  20. [24]

    NOPE-SAC: Neural One-Plane RANSAC for Sparse-View Planar 3D Reconstruc- tion,

    B. Tan, N. Xue, T. Wu, and G.-S. Xia, “NOPE-SAC: Neural One-Plane RANSAC for Sparse-View Planar 3D Reconstruc- tion,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, 2023

  21. [27]

    3-D Reconstruction from Sparse Views using Monocular Vision,

    A. Saxena, M. Sun, and A. Y . Ng, “3-D Reconstruction from Sparse Views using Monocular Vision,” in2007 IEEE 11th International Conference on Computer Vision. Rio de Janeiro, Brazil: IEEE, 2007

  22. [28]

    Single and sparse view 3D recon- struction by learning shape priors,

    Y . Chen and R. Cipolla, “Single and sparse view 3D recon- struction by learning shape priors,”Computer Vision and Image Understanding, vol. 115, 2011

  23. [29]

    Planar surface re- construction from sparse views,

    L. Jin, S. Qian, A. Owens, and D. F. Fouhey, “Planar surface re- construction from sparse views,”2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 12 971–12 980, 2021

  24. [30]

    Matterport3d: Learning from rgb-d data in indoor environments,

    A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y . Zhang, “Matterport3d: Learning from rgb-d data in indoor environments,” in2017 International Conference on 3D Vision (3DV). IEEE Computer Society, 2017, pp. 667–676

  25. [31]

    Neural 3D reconstruction from sparse views using geometric priors,

    T.-J. Mu, H.-X. Chen, J.-X. Cai, and N. Guo, “Neural 3D reconstruction from sparse views using geometric priors,”Com- putational Visual Media, vol. 9, 2023

  26. [32]

    pixelnerf: Neural radiance fields from one or few images,

    A. Yu, V . Ye, M. Tancik, and A. Kanazawa, “pixelnerf: Neural radiance fields from one or few images,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 4576–4585

  27. [33]

    Stereo Radiance Fields (SRF): Learning View Synthesis for Sparse Views of Novel Scenes,

    J. Chibane, A. Bansal, V . Lazova, and G. Pons-Moll, “Stereo Radiance Fields (SRF): Learning View Synthesis for Sparse Views of Novel Scenes,” in2021 IEEE/CVF Conference on 26 Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2021

  28. [35]

    Gs4: Gen- eralizable sparse splatting semantic slam,

    M. Jiang, C. Kim, C. Ziwen, and L. Fuxin, “Gs4: Gen- eralizable sparse splatting semantic slam,”arXiv preprint arXiv:2506.06517, 2025

  29. [37]

    Accurate and efficient stereo processing by semi-global matching and mutual information,

    H. Hirschm ¨uller, “Accurate and efficient stereo processing by semi-global matching and mutual information,” in2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), vol. 2, 2005, pp. 807–814

  30. [38]

    Directed ray distance functions for 3d scene reconstruction,

    N. Kulkarni, J. Johnson, and D. F. Fouhey, “Directed ray distance functions for 3d scene reconstruction,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 201–219

  31. [39]

    Vision transformers for dense prediction,

    R. Ranftl, A. Bochkovskiy, and V . Koltun, “Vision transformers for dense prediction,” inProceedings of the IEEE/CVF interna- tional conference on computer vision, 2021, pp. 12 179–12 188

  32. [40]

    Geometry-free view synthesis: Transformers and no 3d priors,

    R. Rombach, P. Esser, and B. Ommer, “Geometry-free view synthesis: Transformers and no 3d priors,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 356–14 366

  33. [42]

    3d clothed human reconstruction from sparse multi- view images,

    J. G. Hong, S. Y . Noh, H.-K. Lee, W.-S. Cheong, and J. Y . Chang, “3d clothed human reconstruction from sparse multi- view images,”2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 677–687, 2024

  34. [43]

    Pvp-recon: Progressive view planning via warping consistency for sparse-view surface reconstruction,

    S. Ye, Y . He, M. Lin, J. Sheng, R. Fan, Y . Han, Y . Hu, R. Yi, Y .-H. Wen, Y .-J. Liu, and W. Wang, “Pvp-recon: Progressive view planning via warping consistency for sparse-view surface reconstruction,”ACM Transactions on Graphics (TOG), vol. 43, pp. 1 – 13, 2024

  35. [44]

    RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs,

    M. Niemeyer, J. T. Barron, B. Mildenhall, M. S. M. Sajjadi, A. Geiger, and N. Radwan, “RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, LA, USA: IEEE, 2022

  36. [45]

    Density estimation using real nvp,

    L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using real nvp,” inInternational Conference on Learning Rep- resentations (ICLR), 2017

  37. [46]

    FlipNeRF: Flipped Reflection Rays for Few-shot Novel View Synthesis,

    S. Seo, Y . Chang, and N. Kwak, “FlipNeRF: Flipped Reflection Rays for Few-shot Novel View Synthesis,” in2023 IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023

  38. [47]

    Fast sparse view guided nerf update for object reconfigurations,

    Z. Lu, J. Ye, X. Fei, X. Li, J. Mo, A. Swaminathan, and S. Soatto, “Fast sparse view guided nerf update for object reconfigurations,”arXiv preprint arXiv:2403.11024, 2024

  39. [48]

    As-nerf: Learning auxiliary sampling for generalizable novel view syn- thesis from sparse views,

    J. Tang, L. Li, X. Qi, Y . Chen, C. Fan, and X. Yu, “As-nerf: Learning auxiliary sampling for generalizable novel view syn- thesis from sparse views,”2024 IEEE International Conference on Multimedia and Expo (ICME), pp. 1–6, 2024

  40. [49]

    Sc-nerf: Self-correcting neural radiance field with sparse views,

    L. Song, G. Wang, J. Liu, Z. Fu, Y . Miaoet al., “Sc-nerf: Self-correcting neural radiance field with sparse views,”arXiv preprint arXiv:2309.05028, 2023

  41. [50]

    Sparse- derf: Deblurred neural radiance fields from sparse view,

    D. Lee, D. Kim, J. Lee, M. Lee, S. Lee, and S. Lee, “Sparse- derf: Deblurred neural radiance fields from sparse view,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, pp. 6842–6858, 2024

  42. [51]

    Dense Depth Priors for Neural Radiance Fields from Sparse Input Views,

    B. Roessle, J. T. Barron, B. Mildenhall, P. P. Srinivasan, and M. Niebner, “Dense Depth Priors for Neural Radiance Fields from Sparse Input Views,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, LA, USA: IEEE, 2022

  43. [52]

    Depth-supervised NeRF: Fewer Views and Faster Training for Free,

    K. Deng, A. Liu, J.-Y . Zhu, and D. Ramanan, “Depth-supervised NeRF: Fewer Views and Faster Training for Free,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR). New Orleans, LA, USA: IEEE, 2022

  44. [53]

    Sparsesat-nerf: Dense depth super- vised neural radiance fields for sparse satellite images,

    L. Zhang and E. Rupnik, “Sparsesat-nerf: Dense depth super- vised neural radiance fields for sparse satellite images,” inISPRS Annals 2023, 2023

  45. [54]

    Ibrnet: Learning multi-view image-based rendering,

    Q. Wang, Z. Wang, K. Genova, P. P. Srinivasan, H. Zhou, J. T. Barron, R. Martin-Brualla, N. Snavely, and T. Funkhouser, “Ibrnet: Learning multi-view image-based rendering,” inPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 4690–4699

  46. [55]

    X-nerf: Explicit neural radiance field for multi- scene 360deg insufficient rgb-d views,

    H. Zhu, “X-nerf: Explicit neural radiance field for multi- scene 360deg insufficient rgb-d views,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 5766–5775

  47. [56]

    6img-to-3d: Few-image large-scale outdoor driving scene re- construction,

    T. Gieruc, M. K ¨astingsch¨afer, S. Bernhard, and M. Salzmann, “6img-to-3d: Few-image large-scale outdoor driving scene re- construction,”arXiv preprint arXiv:2404.12378, 2024

  48. [57]

    Ners: Neural reflectance surfaces for sparse-view 3d reconstruction in the wild,

    J. Zhang, G. Yang, S. Tulsiani, and D. Ramanan, “Ners: Neural reflectance surfaces for sparse-view 3d reconstruction in the wild,”Advances in Neural Information Processing Systems, vol. 34, pp. 29 835–29 847, 2021

  49. [58]

    Is vanilla mlp in neural radiance field enough for few-shot view synthesis?

    H. Zhu, T. He, X. Li, B. Li, and Z. Chen, “Is vanilla mlp in neural radiance field enough for few-shot view synthesis?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 288–20 298

  50. [59]

    Cmc: few-shot novel view synthesis via cross-view multiplane consistency,

    H. Zhu and Z. Chen, “Cmc: few-shot novel view synthesis via cross-view multiplane consistency,” in2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR). IEEE, 2024, pp. 960–968

  51. [60]

    Cvt-xrf: Contrastive in-voxel transformer for 3d consistent radiance fields from sparse inputs,

    Y . Zhong, L. Hong, Z. Li, and D. Xu, “Cvt-xrf: Contrastive in-voxel transformer for 3d consistent radiance fields from sparse inputs,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 466– 21 475

  52. [61]

    Where and how: Mitigating confusion in neural radiance fields from sparse inputs,

    Y . Bao, Y . Li, J. Huo, T. Ding, X. Liang, W. Li, and Y . Gao, “Where and how: Mitigating confusion in neural radiance fields from sparse inputs,”arXiv preprint arXiv:2308.02908, 2023

  53. [62]

    Nerf-or: neural radiance fields for operating room scene recon- struction from sparse-view rgb-d videos,

    B. G. A. Gerats, J. M. Wolterink, and I. A. M. J. Broeders, “Nerf-or: neural radiance fields for operating room scene recon- struction from sparse-view rgb-d videos,”International Journal of Computer Assisted Radiology and Surgery, vol. 20, pp. 147 – 156, 2024

  54. [63]

    Repurposing diffusion-based image gener- ators for monocular depth estimation,

    B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler, “Repurposing diffusion-based image gener- ators for monocular depth estimation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2024, pp. 9492–9502

  55. [64]

    D ¨arf: Boosting radiance fields from sparse input views with monocular depth adaptation,

    J. Song, S. Park, H. An, S. Cho, M.-S. Kwak, S. Cho, and S. Kim, “D ¨arf: Boosting radiance fields from sparse input views with monocular depth adaptation,”Advances in Neural Information Processing Systems, vol. 36, pp. 68 458–68 470, 2023

  56. [65]

    Divinet: 3d reconstruction from disparate views via neural template regularization,

    A. V ora, A. G. Patil, and H. Zhang, “Divinet: 3d reconstruction from disparate views via neural template regularization,”arXiv preprint arXiv:2306.04699, 2023

  57. [66]

    Regularizing Neural Radiance Fields from Sparse Rgb-D Inputs,

    Q. Li, F. Multon, and A. Boukhayma, “Regularizing Neural Radiance Fields from Sparse Rgb-D Inputs,” in2023 IEEE International Conference on Image Processing (ICIP). Kuala Lumpur, Malaysia: IEEE, 2023

  58. [67]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” inInternational Conference on Learning Representations (ICLR), 2015

  59. [68]

    Empowering sparse-input neural radiance fields with dual-level semantic guidance from dense novel views,

    Y . Zhong, K. Zhou, Z. Li, L. Hong, Z. Li, and D. Xu, “Empowering sparse-input neural radiance fields with dual-level semantic guidance from dense novel views,”arXiv preprint arXiv:2503.02230, 2025

  60. [69]

    Hg3-nerf: Hierarchical geomet- ric, semantic, and photometric guided neural radiance fields for sparse view inputs,

    Z. Gao, W. Dai, and Y . Zhang, “Hg3-nerf: Hierarchical geomet- ric, semantic, and photometric guided neural radiance fields for sparse view inputs,”arXiv preprint arXiv:2401.11711, 2024

  61. [70]

    ViP-NeRF: Visibility Prior for Sparse Input Neural Radiance Fields,

    N. Somraj and R. Soundararajan, “ViP-NeRF: Visibility Prior for Sparse Input Neural Radiance Fields,” inSpecial Interest 27 Group on Computer Graphics and Interactive Techniques Con- ference Conference Proceedings, 2023

  62. [71]

    SimpleNeRF: Regularizing Sparse Input Neural Radiance Fields with Simpler Solutions,

    N. Somraj, A. Karanayil, and R. Soundararajan, “SimpleNeRF: Regularizing Sparse Input Neural Radiance Fields with Simpler Solutions,” inSIGGRAPH Asia 2023 Conference Papers, 2023

  63. [72]

    Consistentnerf: Enhancing neural radiance fields with 3d consistency for sparse view synthesis,

    S. Hu, K. Zhou, K. Li, L. Yu, L. Hong, T. Hu, Z. Li, G. H. Lee, and Z. Liu, “Consistentnerf: Enhancing neural radiance fields with 3d consistency for sparse view synthesis,”arXiv preprint arXiv:2305.11031, 2023

  64. [73]

    Vgos: V oxel grid optimization for view synthesis from sparse inputs,

    J. Sun, Z. Zhang, J. Chen, G. Li, B. Ji, L. Zhao, W. Xing, and H. Lin, “Vgos: V oxel grid optimization for view synthesis from sparse inputs,”arXiv preprint arXiv:2304.13386, 2023

  65. [74]

    Spatial annealing for efficient few-shot neural rendering,

    Y . Xiao, D. Zhai, W. Zhao, K. Jiang, J. Jiang, and X. Liu, “Spatial annealing for efficient few-shot neural rendering,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 8, 2025, pp. 8691–8699

  66. [75]

    Freenerf: Improving few- shot neural rendering with free frequency regularization,

    J. Yang, M. Pavone, and Y . Wang, “Freenerf: Improving few- shot neural rendering with free frequency regularization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 8254–8263

  67. [76]

    Framenerf: A simple and efficient framework for few-shot novel view synthesis,

    Y . Xing, P. Wang, L. Liu, D. Li, and L. Zhang, “Framenerf: A simple and efficient framework for few-shot novel view synthesis,”arXiv preprint arXiv:2402.14586, 2024

  68. [77]

    Tensorf: Tensorial radiance fields,

    A. Chen, Z. Xu, A. Zhao, V . J. Prabhakaran, A. Ghosh, H. Su, and J. Yu, “Tensorf: Tensorial radiance fields,” inEuropean Conference on Computer Vision (ECCV). Springer, 2022, pp. 333–350

  69. [78]

    Arc- nerf: Area ray casting for broader unseen view coverage in few- shot object rendering,

    S. Seo, Y . Chang, J. Yoo, S. Lee, H. Lee, and N. Kwak, “Arc- nerf: Area ray casting for broader unseen view coverage in few- shot object rendering,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 273–283

  70. [79]

    Few-shot neural radiance fields under unconstrained illumination,

    S. Lee, J. Choi, S. Kim, I.-J. Kim, and J. Cho, “Few-shot neural radiance fields under unconstrained illumination,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 4, 2024, pp. 2938–2946

  71. [80]

    Caesarnerf: Calibrated semantic representation for few-shot generalizable neural rendering,

    H. Zhu, T. Ding, T. Chen, I. Zharkov, R. Nevatia, and L. Liang, “Caesarnerf: Calibrated semantic representation for few-shot generalizable neural rendering,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 71–89

  72. [81]

    Manifoldnerf: View-dependent image feature supervi- sion for few-shot neural radiance fields,

    D. Kanaoka, M. Sonogashira, H. Tamukoh, and Y . Kawan- ishi, “Manifoldnerf: View-dependent image feature supervi- sion for few-shot neural radiance fields,”arXiv preprint arXiv:2310.13670, 2023

  73. [82]

    Frugalnerf: Fast convergence for extreme few-shot novel view synthesis without learned priors,

    C.-Y . Lin, C.-H. Wu, C.-H. Yeh, S.-H. Yen, C. Sun, and Y .-L. Liu, “Frugalnerf: Fast convergence for extreme few-shot novel view synthesis without learned priors,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 11 227–11 238

  74. [83]

    Neo 360: Neural fields for sparse view synthesis of outdoor scenes,

    M. Z. Irshad, S. Zakharov, K. Liu, V . Guizilini, T. Kollar, A. Gaidon, Z. Kira, and R. Ambrus, “Neo 360: Neural fields for sparse view synthesis of outdoor scenes,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 9187–9198

  75. [84]

    Leap: Liberate sparse-view 3d modeling from camera poses,

    H. Jiang, Z. Jiang, Y . Zhao, and Q. Huang, “Leap: Liberate sparse-view 3d modeling from camera poses,”arXiv preprint arXiv:2310.01410, 2023

  76. [85]

    Sparsepose: Sparse-view camera pose regression and refinement,

    S. Sinha, J. Y . Zhang, A. Tagliasacchi, I. Gilitschenski, and D. B. Lindell, “Sparsepose: Sparse-view camera pose regression and refinement,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 21 349– 21 359

  77. [87]

    Barf: Bundle- adjusting neural radiance fields,

    C.-H. Lin, W.-C. Ma, A. Torralba, and S. Lucey, “Barf: Bundle- adjusting neural radiance fields,”2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5721–5731, 2021

  78. [88]

    Mixnerf: Modeling a ray with mixture density for novel view synthesis from sparse inputs,

    S. Seo, D. Han, Y . Chang, and N. Kwak, “Mixnerf: Modeling a ray with mixture density for novel view synthesis from sparse inputs,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 20 659– 20 668

  79. [89]

    Panerf: Pseudo-view augmentation for improved neural radiance fields based on few-shot inputs,

    Y . C. Ahn, S. Jang, S. Park, J.-Y . Kim, and N. Kang, “Panerf: Pseudo-view augmentation for improved neural radiance fields based on few-shot inputs,”arXiv preprint arXiv:2211.12758, 2022

  80. [90]

    InfoNeRF: Ray Entropy Min- imization for Few-Shot Neural V olume Rendering,

    M. Kim, S. Seo, and B. Han, “InfoNeRF: Ray Entropy Min- imization for Few-Shot Neural V olume Rendering,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR). New Orleans, LA, USA: IEEE, 2022

  81. [91]

    Geconerf: Few-shot neural radiance fields via geometric consistency,

    M.-S. Kwak, J. Song, and S. Kim, “Geconerf: Few-shot neural radiance fields via geometric consistency,” inProceedings of the 40th International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol. 202. PMLR, 2023, pp. 7437–7444

  82. [92]

    Putting nerf on a diet: Se- mantically consistent few-shot view synthesis,

    A. Jain, M. Tancik, and P. Abbeel, “Putting nerf on a diet: Se- mantically consistent few-shot view synthesis,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5885–5894

  83. [93]

    Preface: A data-driven volumetric prior for few-shot ultra high- resolution face synthesis,

    M. C. B ¨uhler, K. Sarkar, T. Shah, G. Li, D. Wang, L. Helminger, S. Orts-Escolano, D. Lagun, O. Hilliges, T. Beeler, and A. Meka, “Preface: A data-driven volumetric prior for few-shot ultra high- resolution face synthesis,”arXiv preprint arXiv:2309.16859, 2023

  84. [94]

    Neurallift-360: Lifting an in-the-wild 2d photo to a 3d object with 360deg views,

    D. Xu, Y . Jiang, P. Wang, Z. Fan, Y . Wang, and Z. Wang, “Neurallift-360: Lifting an in-the-wild 2d photo to a 3d object with 360deg views,” inProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, 2023, pp. 4479–4489

  85. [95]

    Sc-neus: Consistent neural surface reconstruction from sparse and noisy views,

    S.-S. Huang, Z. Zou, Y . Zhang, Y .-P. Cao, and Y . Shan, “Sc-neus: Consistent neural surface reconstruction from sparse and noisy views,” inProceedings of the AAAI conference on artificial intelligence, vol. 38, no. 3, 2024, pp. 2357–2365

  86. [96]

    Generic objects as pose probes for few-shot view synthesis,

    Z. Gao, R. Yi, C. Zhu, K. Zhuang, W. Chen, and K. Xu, “Generic objects as pose probes for few-shot view synthesis,” IEEE Transactions on Circuits and Systems for Video Technol- ogy, 2025

  87. [98]

    SPARF: Large- Scale Learning of 3D Sparse Radiance Fields from Few Input Images,

    A. Hamdi, B. Ghanem, and M. Nießner, “SPARF: Large- Scale Learning of 3D Sparse Radiance Fields from Few Input Images,” in2023 IEEE/CVF International Conference on Com- puter Vision Workshops (ICCVW), 2023

  88. [99]

    Eg-humannerf: Efficient generalizable human nerf utilizing human prior for sparse view,

    Z. Wang, Y . Kanamori, and Y . Endo, “Eg-humannerf: Efficient generalizable human nerf utilizing human prior for sparse view,” arXiv preprint arXiv:2410.12242, 2024

  89. [100]

    Flexnerf: Photorealistic free-viewpoint rendering of moving humans from sparse views,

    V . Jayasundara, A. Agrawal, N. Heron, A. Shrivastava, and L. S. Davis, “Flexnerf: Photorealistic free-viewpoint rendering of moving humans from sparse views,”2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pp. 21 118–21 127, 2023

  90. [101]

    Sparsegs: Real-time 360° sparse view synthesis using gaussian splatting,

    H. Xiong, S. Muttukuru, R. Upadhyay, P. Chari, and A. Kadambi, “Sparsegs: Real-time 360° sparse view synthesis using gaussian splatting,”arXiv preprint arXiv:2312.00206, 2025

  91. [102]

    Dreamfusion: Text-to-3d using 2d diffusion,

    B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “Dreamfusion: Text-to-3d using 2d diffusion,”arXiv preprint arXiv:2209.14988, 2022

  92. [103]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10 684– 10 695

  93. [104]

    Loopsparsegs: Loop based sparse-view friendly gaussian splat- ting,

    Z. Bao, G. Liao, K. Zhou, K. Liu, Q. Li, and G. Qiu, “Loopsparsegs: Loop based sparse-view friendly gaussian splat- ting,”arXiv preprint arXiv:2408.00254, 2024

  94. [105]

    Gaussianobject: High-quality 3d object re- construction from four views with gaussian splatting,

    C. Yang, S. Li, J. Fang, R. Liang, L. Xie, X. Zhang, W. Shen, and Q. Tian, “Gaussianobject: High-quality 3d object re- construction from four views with gaussian splatting,”arXiv preprint arXiv:2402.10259, 2024. 28

  95. [106]

    Adding conditional control to text-to-image diffusion models,

    B. Zhang, S. Gu, J. Zhao, T. Wu, Y . Luo, J. Zhu, and P. Luo, “Adding conditional control to text-to-image diffusion models,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 18 338–18 348

  96. [107]

    Dropgaussian: Structural regu- larization for sparse-view gaussian splatting,

    H. Park, G. Ryu, and W. Kim, “Dropgaussian: Structural regu- larization for sparse-view gaussian splatting,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 600–21 609

  97. [108]

    S2gaussian: Sparse- view super-resolution 3d gaussian splatting,

    Y . Wan, M. Shao, Y . Cheng, and W. Zuo, “S2gaussian: Sparse- view super-resolution 3d gaussian splatting,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 711–721

  98. [109]

    Structure consistent gaussian splatting with matching prior for few-shot novel view synthesis,

    R. Peng, W. Xu, L. Tang, J. Jiao, R. Wanget al., “Structure consistent gaussian splatting with matching prior for few-shot novel view synthesis,”Advances in Neural Information Process- ing Systems, vol. 37, pp. 97 328–97 352, 2024

  99. [110]

    Uncertainty-guided optimal transport in depth supervised sparse-view 3d gaussian,

    W. Sun, Q. Zhang, Y . Zhou, Q. Ye, J. Jiao, and Y . Li, “Uncertainty-guided optimal transport in depth supervised sparse-view 3d gaussian,”arXiv preprint arXiv:2405.19657, 2024

  100. [111]

    Speedy-splat: Fast 3d gaussian splatting with sparse pixels and sparse primitives,

    A. Hanson, A. Tu, G. Lin, V . Singla, M. Zwicker, and T. Gold- stein, “Speedy-splat: Fast 3d gaussian splatting with sparse pixels and sparse primitives,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 537– 21 546

  101. [112]

    Transplat: Gen- eralizable 3d gaussian splatting from sparse multi-view images with transformers,

    C. Zhang, Y . Zou, Z. Li, M. Yi, and H. Wang, “Transplat: Gen- eralizable 3d gaussian splatting from sparse multi-view images with transformers,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 9, 2025, pp. 9869–9877

  102. [113]

    Vgnc: Reducing the overfitting of sparse-view 3dgs via validation-guided gaussian number control,

    L. Lin, R. Lu, Q. Chen, H. Ren, M. Lu, Y . Sun, C. Yan, and A. Xue, “Vgnc: Reducing the overfitting of sparse-view 3dgs via validation-guided gaussian number control,”arXiv preprint arXiv:2504.14548, 2025

  103. [114]

    Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis,

    W. Yu, J. Xing, L. Yuan, W. Hu, X. Li, Z. Huang, X. Gao, T.-T. Wong, Y . Shan, and Y . Tian, “Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis,”arXiv preprint arXiv:2409.02048, 2024

  104. [115]

    Uniforward: Unified 3d scene and semantic field reconstruction via feed- forward gaussian splatting from only sparse-view images,

    Q. Tian, X. Tan, J. Gong, Y . Xie, and L. Ma, “Uniforward: Unified 3d scene and semantic field reconstruction via feed- forward gaussian splatting from only sparse-view images,”arXiv preprint arXiv:2506.09378, 2025

  105. [116]

    Sparsplat: Fast multi-view reconstruction with generalizable 2d gaussian splat- ting,

    S. Jena, S. R. Vutukur, and A. Boukhayma, “Sparsplat: Fast multi-view reconstruction with generalizable 2d gaussian splat- ting,”arXiv preprint arXiv:2505.02175, 2025

  106. [117]

    General- izable human gaussians for sparse view synthesis,

    Y . Kwon, B. Fang, Y . Lu, H. Dong, C. Zhang, F. V . Carrasco, A. Mosella-Montoro, J. Xu, S. Takagi, D. Kimet al., “General- izable human gaussians for sparse view synthesis,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 451–468

  107. [118]

    Sparse2dgs: Geometry-prioritized gaussian splatting for surface reconstruc- tion from sparse views,

    J. Wu, R. Li, Y . Zhu, R. Guo, J. Sun, and Y . Zhang, “Sparse2dgs: Geometry-prioritized gaussian splatting for surface reconstruc- tion from sparse views,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 11 307–11 316

  108. [119]

    Hisplat: Hierarchical 3d gaussian splatting for generalizable sparse-view reconstruction,

    S. Tang, W. Ye, P. Ye, W. Lin, Y . Zhou, T. Chen, and W. Ouyang, “Hisplat: Hierarchical 3d gaussian splatting for generalizable sparse-view reconstruction,”arXiv preprint arXiv:2410.06245, 2024

  109. [120]

    Vggt: Visual geometry grounded transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 5294–5306

  110. [121]

    Improving geometry in sparse-view 3dgs via reprojection-based dof separation,

    Y . Kim, M. Park, J. Choi, and S. Yoon, “Improving geometry in sparse-view 3dgs via reprojection-based dof separation,”arXiv preprint arXiv:2412.14568, 2024

  111. [122]

    Spars3r: Semantic prior alignment and regularization for sparse 3d reconstruction,

    Y . Tang, Y . Guo, D. Li, and C. Peng, “Spars3r: Semantic prior alignment and regularization for sparse 3d reconstruction,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 26 810–26 821

  112. [123]

    Dust3r: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,”2024 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pp. 20 697–20 709, 2023

  113. [124]

    Mast3r: Multi-scale attention for stereo matching with transformers,

    X. Li, Y . Yuan, E. Xie, Y . Sun, Z. Chen, W. Wang, P. Luo, and L. Shao, “Mast3r: Multi-scale attention for stereo matching with transformers,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 18 103–18 113

  114. [125]

    Comapgs: Covisibility map- based gaussian splatting for sparse novel view synthesis,

    Y . Jang and E. P ´erez-Pellitero, “Comapgs: Covisibility map- based gaussian splatting for sparse novel view synthesis,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 26 779–26 788

  115. [126]

    Fewviewgs: Gaussian splatting with few view matching and multi-stage training,

    R. Yin, V . Yugay, Y . Li, S. Karaoglu, and T. Gevers, “Fewviewgs: Gaussian splatting with few view matching and multi-stage training,”Advances in Neural Information Process- ing Systems, vol. 37, pp. 127 204–127 225, 2024

  116. [127]

    Solidgs: Consolidating gaussian surfel splatting for sparse-view surface reconstruction,

    Z. Shen, Y . Liu, Z. Chen, Z. Li, J. Wang, Y . Liang, Z. Yu, J. Zhang, Y . Xu, S. Schaeferet al., “Solidgs: Consolidating gaussian surfel splatting for sparse-view surface reconstruction,” arXiv preprint arXiv:2412.15400, 2024

  117. [128]

    Mvpgs: Excavating multi-view priors for gaussian splatting from sparse input views,

    W. Xu, H. Gao, S. Shen, R. Peng, J. Jiao, and R. Wang, “Mvpgs: Excavating multi-view priors for gaussian splatting from sparse input views,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 203–220

  118. [129]

    GeoRGS: Geometric Regularization for Real-Time Novel View Synthesis From Sparse Inputs,

    Z. Liu, J. Su, G. Cai, Y . Chen, B. Zeng, and Z. Wang, “GeoRGS: Geometric Regularization for Real-Time Novel View Synthesis From Sparse Inputs,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, 2024

  119. [130]

    Depth-regularized optimization for 3d gaussian splatting in few-shot images,

    J. Chung, J. Oh, and K. M. Lee, “Depth-regularized optimization for 3d gaussian splatting in few-shot images,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 811–820

  120. [131]

    Mvg-splatting: Multi-view guided gaussian splatting with adaptive quantile-based geometric consistency densification,

    Z. Li, S. Yao, Y . Chu, A. F. Garcia-Fernandez, Y . Yue, E. G. Lim, and X. Zhu, “Mvg-splatting: Multi-view guided gaussian splatting with adaptive quantile-based geometric consistency densification,”arXiv preprint arXiv:2407.11840, 2024

  121. [132]

    Infonorm: Mutual information shaping of normals for sparse-view reconstruction,

    X. Wang, S. Dong, Y . Zheng, and Y . Yang, “Infonorm: Mutual information shaping of normals for sparse-view reconstruction,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 242–258

  122. [133]

    Unig: Modelling unitary 3d gaussians for view-consistent 3d reconstruction,

    J. Wu, K. Liu, Y . Shi, X. Jiang, Y . Yao, and L. Zhang, “Unig: Modelling unitary 3d gaussians for view-consistent 3d reconstruction,”arXiv preprint arXiv:2410.13195, 2024

  123. [134]

    End-to-end object detection with trans- formers,

    N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with trans- formers,” inEuropean Conference on Computer Vision (ECCV). Springer, 2020, pp. 213–229

  124. [135]

    Intern-gs: Vision model guided sparse-view 3d gaussian splatting,

    X. Sun, R. Chen, M. Gong, D. Xu, and T. Liu, “Intern-gs: Vision model guided sparse-view 3d gaussian splatting,”arXiv preprint arXiv:2505.20729, 2025

  125. [136]

    pix- elsplat: 3d gaussian splats from image pairs for scalable gen- eralizable 3d reconstruction,

    D. Charatan, S. L. Li, A. Tagliasacchi, and V . Sitzmann, “pix- elsplat: 3d gaussian splats from image pairs for scalable gen- eralizable 3d reconstruction,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 19 457–19 467

  126. [137]

    Deceptive-nerf/3dgs: Diffusion-generated pseudo-observations for high-quality sparse-view reconstruction,

    X. Liu, J. Chen, S.-H. Kao, Y .-W. Tai, and C.-K. Tang, “Deceptive-nerf/3dgs: Diffusion-generated pseudo-observations for high-quality sparse-view reconstruction,” inEuropean Con- ference on Computer Vision. Springer, 2024, pp. 337–355

  127. [138]

    How to use diffusion priors under sparse views?

    Q. Wang, Y . Zhao, J. Ma, and J. Li, “How to use diffusion priors under sparse views?”arXiv preprint arXiv:2412.02225, 2024

  128. [139]

    Lm-gaussian: Boost sparse-view 3d gaussian splatting with large model priors,

    H. Yu, X. Long, and P. Tan, “Lm-gaussian: Boost sparse-view 3d gaussian splatting with large model priors,”arXiv preprint arXiv:2409.03456, 2024

  129. [140]

    Mvsplat360: Feed-forward 360 scene synthesis from sparse views,

    Y . Chen, C. Zheng, H. Xu, B. Zhuang, A. Vedaldi, T.-J. Cham, and J. Cai, “Mvsplat360: Feed-forward 360 scene synthesis from sparse views,”arXiv preprint arXiv:2411.04924, 2024

  130. [141]

    Stable video diffusion: Scaling latent video diffusion models to large datasets,

    A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kil- ian, D. Lorenz, Y . Levi, Z. English, V . V oleti, A. Lettset al., “Stable video diffusion: Scaling latent video diffusion models to large datasets,”arXiv preprint arXiv:2311.15127, 2023

  131. [142]

    Prosplat: Improved feed-forward 3d gaussian splatting for wide-baseline sparse views,

    X. Lu, J. Fu, J. Zhang, Z. Song, C. Jia, and S. Ma, “Prosplat: Improved feed-forward 3d gaussian splatting for wide-baseline sparse views,”arXiv preprint arXiv:2506.07670, 2025. 29

  132. [143]

    Flowr: Flowing from sparse to dense 3d reconstructions,

    T. Fischer, S. R. Bul `o, Y .-H. Yang, N. V . Keetha, L. Porzi, N. M ¨uller, K. Schwarz, J. Luiten, M. Pollefeys, and P. Kontschieder, “Flowr: Flowing from sparse to dense 3d reconstructions,”arXiv preprint arXiv:2504.01647, 2025

  133. [144]

    Auggs: Self-augmented gaussians with structural masks for sparse-view 3d reconstruction,

    B. Du, L. Meng, and W. Hu, “Auggs: Self-augmented gaussians with structural masks for sparse-view 3d reconstruction,”arXiv preprint arXiv:2408.04831, 2024

  134. [145]

    V3d: Video diffusion models are effective 3d generators,

    Z. Chen, Y . Wang, F. Wang, Z. Wang, and H. Liu, “V3d: Video diffusion models are effective 3d generators,”arXiv preprint arXiv:2403.06738, 2024

  135. [146]

    Video diffusion models,

    J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,”Advances in neural information processing systems, vol. 35, pp. 8633–8646, 2022

  136. [147]

    Cat4d: Create anything in 4d with multi-view video diffusion models,

    R. Wu, R. Gao, B. Poole, A. Trevithick, C. Zheng, J. T. Barron, and A. Holynski, “Cat4d: Create anything in 4d with multi-view video diffusion models,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 26 057–26 068

  137. [148]

    Ri3d: Few-shot gaussian splatting with repair and inpainting diffusion priors,

    A. Paliwal, X. Zhou, W. Ye, J. Xiong, R. Ranjan, and N. K. Kalantari, “Ri3d: Few-shot gaussian splatting with repair and inpainting diffusion priors,”arXiv preprint arXiv:2503.10860, 2025

  138. [149]

    la- tentsplat: Autoencoding variational gaussians for fast general- izable 3d reconstruction,

    C. Wewer, K. Raj, E. Ilg, B. Schiele, and J. E. Lenssen, “la- tentsplat: Autoencoding variational gaussians for fast general- izable 3d reconstruction,” inEuropean conference on computer vision. Springer, 2024, pp. 456–473

  139. [150]

    Improving novel view synthesis of 360° scenes in extremely sparse views by jointly training hemisphere-sampled synthetic images,

    G. Chen, A. M. Truong, H. Lin, M. Vlaminck, W. Philips, and H. Luong, “Improving novel view synthesis of 360° scenes in extremely sparse views by jointly training hemisphere-sampled synthetic images,”arXiv preprint arXiv:2505.19264, 2025

  140. [151]

    Freesplatter: Pose-free gaussian splatting for sparse-view 3d reconstruction,

    J. Xu, S. Gao, and Y . Shan, “Freesplatter: Pose-free gaussian splatting for sparse-view 3d reconstruction,”arXiv preprint arXiv:2412.09573, 2024

  141. [152]

    Gaussian scenes: Pose- free sparse-view scene reconstruction using depth-enhanced diffusion priors,

    S. Paul, P. Kaushik, and A. Yuille, “Gaussian scenes: Pose- free sparse-view scene reconstruction using depth-enhanced diffusion priors,”arXiv preprint arXiv:2411.15966, 2024

  142. [153]

    ifusion: Inverting diffusion for pose-free reconstruction from sparse views,

    C.-H. Wu, Y .-C. Chen, B. Solarte, L. Yuan, and M. Sun, “ifusion: Inverting diffusion for pose-free reconstruction from sparse views,”arXiv preprint arXiv:2312.17250, 2023

  143. [154]

    Seeing a 3d world in a grain of sand,

    Y . Zhang, Y . Ji, Y . Guo, and J. Ye, “Seeing a 3d world in a grain of sand,”arXiv preprint arXiv:2503.00260, 2025

  144. [155]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, L. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inProceedings of the International Conference on Machine Learning (ICML), 2021

  145. [156]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo et al., “Segment anything,” inProceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015– 4026

  146. [157]

    Emerging properties in self-supervised vision transformers,

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bo- janowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” inProceedings of the International Con- ference on Computer Vision (ICCV), 2021

  147. [158]

    Diffusion models in 3d vision: A survey,

    Z. Wang, D. Li, Y . Wu, T. He, J. Bian, and R. Jiang, “Diffusion models in 3d vision: A survey,”arXiv preprint arXiv:2410.04738, 2025

  148. [159]

    Genfusion: Closing the loop between reconstruction and generation via videos,

    S. Wu, C. Xu, B. Huang, A. Geiger, and A. Chen, “Genfusion: Closing the loop between reconstruction and generation via videos,”arXiv preprint arXiv:2503.21219, 2025

  149. [160]

    Sir-diff: Sparse image sets restoration with multi-view diffusion model,

    Y . Mao, B. Wang, N. Kulkarni, and J. J. Park, “Sir-diff: Sparse image sets restoration with multi-view diffusion model,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 620–21 630

  150. [161]

    Sp2360: Sparse-view 360 scene reconstruction using cascaded 2d diffu- sion priors,

    S. Paul, C. Wewer, B. Schiele, and J. E. Lenssen, “Sp2360: Sparse-view 360 scene reconstruction using cascaded 2d diffu- sion priors,”arXiv preprint arXiv:2405.16517, 2024

  151. [162]

    Vi3drm: Towards meticulous 3d reconstruction from sparse views via photo-realistic novel view synthesis,

    H. Chen, J. Wu, Y . Jin, J. Peng, X. Mao, M. Chi, M. Yao, B. Peng, J. Li, and Y . Cao, “Vi3drm: Towards meticulous 3d reconstruction from sparse views via photo-realistic novel view synthesis,”arXiv preprint arXiv:2409.08207, 2024

  152. [163]

    Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse Views,

    Z. Zou, W. Cheng, Y .-P. Cao, S.-S. Huang, Y . Shan, and S.- H. Zhang, “Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse Views,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, 2024

  153. [164]

    Generating Material-Aware 3D Models from Sparse Views,

    S. Mao, C. Wu, R. Yi, Z. Shen, L. Zhang, and W. Heidrich, “Generating Material-Aware 3D Models from Sparse Views,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1400–1409, 2024

  154. [165]

    Recon- fusion: 3d reconstruction with diffusion priors,

    R. Wu, B. Mildenhall, P. Henzler, K. Park, R. Gao, D. Watson, P. P. Srinivasan, D. Verbin, J. T. Barron, B. Pooleet al., “Recon- fusion: 3d reconstruction with diffusion priors,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 21 ...

  155. [166]

    Reconx: Reconstruct any scene from sparse views with video diffusion model,

    F. Liu, W. Sun, H. Wang, Y . Wang, H. Sun, J. Ye, J. Zhang, and Y . Duan, “Reconx: Reconstruct any scene from sparse views with video diffusion model,”arXiv preprint arXiv:2408.16767, 2024

  156. [167]

    Taming video diffusion prior with scene-grounding guidance for 3d gaussian splatting from sparse inputs,

    Y . Zhong, Z. Li, D. Z. Chen, L. Hong, and D. Xu, “Taming video diffusion prior with scene-grounding guidance for 3d gaussian splatting from sparse inputs,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 6133– 6143

  157. [168]

    Mvdiffusion++: A dense high-resolution multi-view diffusion model for single or sparse-view 3d object reconstruction,

    S. Tang, J. Chen, D. Wang, C. Tang, F. Zhang, Y . Fan, V . Chandra, Y . Furukawa, and R. Ranjan, “Mvdiffusion++: A dense high-resolution multi-view diffusion model for single or sparse-view 3d object reconstruction,” inEuropean Conference on Computer Vision. Springer, 2024, pp...

  158. [169]

    Id-pose: Sparse-view camera pose estimation by inverting diffusion models,

    W. Cheng, Y .-P. Cao, and Y . Shan, “Id-pose: Sparse-view camera pose estimation by inverting diffusion models,”arXiv preprint arXiv:2306.17140, 2023

  159. [170]

    Fine-tuning the diffusion model and distilling informative priors for sparse-view 3d reconstruction,

    J. Tang, Y . Gao, T. Jiang, Y . Yang, and M. Fu, “Fine-tuning the diffusion model and distilling informative priors for sparse-view 3d reconstruction,”2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 7437–7444, 2024

  160. [171]

    Instant neural graphics primitives with a multiresolution hash encoding,

    T. M ¨uller, A. Evans, N. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” in ACM SIGGRAPH 2022 Conference Proceedings. Association for Computing Machinery, 2022, pp. 1–13

  161. [172]

    Fine-tuning the Diffusion Model and Distilling Informative Priors for Sparse- view 3D Reconstruction,

    J. Tang, Y . Gao, T. Jiang, Y . Yang, and M. Fu, “Fine-tuning the Diffusion Model and Distilling Informative Priors for Sparse- view 3D Reconstruction,” in2024 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS). Abu Dhabi, United Arab Emirates: IEEE, 2024

  162. [173]

    How to use diffusion priors under sparse views?

    Q. Wang, Y . Zhao, J. Ma, and J. Li, “How to use diffusion priors under sparse views?”Advances in Neural Information Processing Systems, vol. 37, pp. 30 394–30 424, 2024

  163. [174]

    SparseFusion: Distilling View- Conditioned Diffusion for 3D Reconstruction,

    Z. Zhou and S. Tulsiani, “SparseFusion: Distilling View- Conditioned Diffusion for 3D Reconstruction,” in2023 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR). Vancouver, BC, Canada: IEEE, 2023

  164. [175]

    The less you depend, the more you learn: Synthesizing novel views from sparse, unposed images without any 3d knowledge,

    H. Wang, K. Ye, Y . Li, W. Chen, and B. Chen, “The less you depend, the more you learn: Synthesizing novel views from sparse, unposed images without any 3d knowledge,”arXiv preprint arXiv:2506.09885, 2025

  165. [176]

    Sparp: Fast 3d object reconstruction and pose estimation from sparse views,

    C. Xu, A. Li, L. Chen, Y . Liu, R. Shi, H. Su, and M. Liu, “Sparp: Fast 3d object reconstruction and pose estimation from sparse views,”arXiv preprint arXiv:2408.10195, 2024

  166. [177]

    sshelf: Single-shot hierarchical extrapolation of latent features for 3d reconstruction from sparse-views,

    E. Najafli, M. K ¨astingsch¨afer, S. Bernhard, T. Brox, and A. Geiger, “sshelf: Single-shot hierarchical extrapolation of latent features for 3d reconstruction from sparse-views,”arXiv preprint arXiv:2502.04318, 2025

  167. [178]

    3d vessel reconstruction from sparse-view dynamic dsa images via vessel probability guided attenuation learning,

    Z. Liu, H. Zhao, W. Qin, Z. Zhou, X. Wang, W. Wang, X. Lai, C. Zheng, D. Shen, and Z. Cui, “3d vessel reconstruction from sparse-view dynamic dsa images via vessel probability guided attenuation learning,”arXiv preprint arXiv:2405.10705, 2024

  168. [179]

    Optimizing 3d gaussian splat- ting for sparse viewpoint scene reconstruction,

    S. Chen, J. Zhou, and L. Li, “Optimizing 3d gaussian splat- ting for sparse viewpoint scene reconstruction,”arXiv preprint arXiv:2409.03213, 2024

  169. [180]

    Jointsplat: Probabilistic joint flow-depth optimization for sparse-view gaussian splat- ting,

    Y . Xiao, G. Xu, Q. Wu, and W. Jia, “Jointsplat: Probabilistic joint flow-depth optimization for sparse-view gaussian splat- ting,”arXiv preprint arXiv:2506.03872, 2025. 30

  170. [181]

    Free360: Layered gaussian splatting for unbounded 360-degree view synthesis from extremely sparse and unposed views,

    C. Bao, X. Zhang, Z. Yu, J. Shi, G. Zhang, S. Peng, and Z. Cui, “Free360: Layered gaussian splatting for unbounded 360-degree view synthesis from extremely sparse and unposed views,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 16 377–16 387

  171. [182]

    Spatialsplat: Efficient semantic 3d from sparse unposed images,

    Y . Sheng, J. Deng, X. Zhang, Y . Zhang, B. Hua, Y . Zhang, and J. Ji, “Spatialsplat: Efficient semantic 3d from sparse unposed images,”arXiv preprint arXiv:2505.23044, 2025

  172. [183]

    Large scale multi-view stereopsis evaluation,

    R. Jensen, A. L. Dahl, G. V ogiatzis, E. Tola, and H. Aanæs, “Large scale multi-view stereopsis evaluation,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR), 2014, pp. 406–413

  173. [184]

    Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,”ACM Transactions on Graphics (TOG), vol. 38, no. 4, pp. 29:1–29:14, 2019

  174. [185]

    Tanks and temples: Benchmarking large-scale scene reconstruction,

    A. Knapitsch, J. Park, Q.-Y . Zhou, and V . Koltun, “Tanks and temples: Benchmarking large-scale scene reconstruction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4410–4419

  175. [186]

    Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction,

    J. Reizenstein, G. Riegler, T. Sattler, F. Tombari, M. Polle- feys, and D. Novotny, “Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. ...

  176. [187]

    Stereo magnification: Learning view synthesis using multiplane images,

    T. Zhou, R. Tucker, J. Flynn, G. Fyffe, and N. Snavely, “Stereo magnification: Learning view synthesis using multiplane images,”ACM Transactions on Graphics (TOG), vol. 37, no. 4, pp. 65:1–65:12, 2018

  177. [188]

    Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision,

    L. Ling, Y . Sheng, Z. Tu, W. Zhao, C. Xin, K. Wan, L. Yu, Q. Guo, Z. Yu, Y . Lu, X. Li, X. Sun, R. Ashok, A. Mukherjee, H. Kang, X. Kong, G. Hua, T. Zhang, B. Benes, and A. Bera, “Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision,” inProceedings of the ...

  178. [189]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes,

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2432–2443

  179. [190]

    Scannet++: Large-scale indoor 3d scene under- standing benchmark with densely annotated rgb-d sequences,

    J. Tang, A. Dai, H. Yan, M. Halber, M. Nießner, and T. Funkhouser, “Scannet++: Large-scale indoor 3d scene under- standing benchmark with densely annotated rgb-d sequences,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. ...

  180. [191]

    Mvimgnet: A large-scale dataset of multi-view images,

    X. Yu, M. Xu, Y . Zhang, H. Liu, C. Ye, Y . Wu, Z. Yan, C. Zhu, Z. Xiong, T. Liang, G. Chen, S. Cui, and X. Han, “Mvimgnet: A large-scale dataset of multi-view images,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 9150–9161

  181. [192]

    Shapenet: An information-rich 3d model repository,

    A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Suet al., “Shapenet: An information-rich 3d model repository,”arXiv preprint arXiv:1512.03012, 2015

  182. [193]

    Omniobject3d: Large- vocabulary 3d object dataset for realistic perception, recon- struction and generation,

    T. Wu, J. Zhang, X. Fu, Y . Wang, J. Ren, L. Pan, W. Wu, L. Yang, J. Wang, C. Qianet al., “Omniobject3d: Large- vocabulary 3d object dataset for realistic perception, recon- struction and generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  183. [194]

    Openillumination: A multi-illumination dataset for inverse rendering evaluation on real objects,

    I. Liu, L. Chen, Z. Fu, L. Wu, H. Jin, Z. Li, C. M. R. Wong, Y . Xu, R. Ramamoorthi, Z. Xuet al., “Openillumination: A multi-illumination dataset for inverse rendering evaluation on real objects,”Advances in Neural Information Processing Systems, vol. 36, pp. 36 951–36 962, 2023

  184. [195]

    Development of an image data set of construction machines for deep learning object detection,

    B. Xiao and S. C. Kang, “Development of an image data set of construction machines for deep learning object detection,” Journal of Computing in Civil Engineering, vol. 35, no. 2, p. 05020005, 2021

  185. [196]

    Metrics for objective quality assessment of video from lossy compression,

    Q. Huynh-Thu and M. Ghanbari, “Metrics for objective quality assessment of video from lossy compression,”Signal Process- ing: Image Communication, vol. 25, no. 7, pp. 477–486, 2008

  186. [197]

    Image quality assessment: From error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,”IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004

  187. [198]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 586–595

  188. [199]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium,

    M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,”Advances in neural information processing systems, vol. 30, 2017

  189. [200]

    Image quality assessment: Unifying structure and texture similarity,

    K. Ding, K. Ma, S. Wang, and E. P. Simoncelli, “Image quality assessment: Unifying structure and texture similarity,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 5, pp. 2567–2581, 2022

  190. [201]

    A point set generation network for 3d object reconstruction from a single image,

    H. Fan, H. Su, and L. J. Guibas, “A point set generation network for 3d object reconstruction from a single image,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 605–613

  191. [202]

    Occupancy networks: Learning 3d reconstruction in function space,

    L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger, “Occupancy networks: Learning 3d reconstruction in function space,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4460–4470

  192. [203]

    A benchmark for the evaluation of rgb-d slam systems,

    J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers, “A benchmark for the evaluation of rgb-d slam systems,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2012, pp. 573–580

  193. [204]

    Lsd-slam: Large-scale direct monocular slam,

    J. Engel, T. Sch ¨ops, and D. Cremers, “Lsd-slam: Large-scale direct monocular slam,”European Conference on Computer Vision (ECCV), pp. 834–849, 2014

  194. [205]

    Neural semantic scene synthesis with uncertainty modeling,

    Z. Yuan, J. Fu, L. Yu, T. Zhou, H. Lu, and W. Liu, “Neural semantic scene synthesis with uncertainty modeling,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 1920–1929

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.