Pith. sign in

REVIEW 4 major objections 6 minor 68 references

RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read RDG-GS claims sparse-view 3D rendering can reach state-of-the-art quality by replacing absolute monocular depth supervision with relative depth guidance.

desk verdict Promising sparse-view 3DGS method with a genuinely new relative-depth guidance loss, but the evaluation is so internally inconsistent that the SOTA claim cannot be verified from the paper. read the letter →

arxiv 2501.11102 v1 pith:XEPF4QPE submitted 2025-01-19 cs.CV

classification cs.CV
keywords sparse-view3DrenderingGaussianSplattingrelativedepthguidancemonocularpriorview-consistentgeometryadaptivesamplingnovelviewsynthesisreal-time
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes RDG-GS, a 3D Gaussian Splatting pipeline for rendering novel views from very few input images, typically 3 to 24 views. Its central claim is that supervising geometry with relative depth relationships, rather than absolute monocular depth values, yields view-consistent geometry and state-of-the-art rendering quality while keeping rendering real-time. The paper reports large gains over prior NeRF-based and Gaussian-based sparse-view methods on four benchmarks, including PSNR 26.03 on Mip-NeRF360's 1/8-resolution setting with 24 training views and 112 FPS rendering. A sympathetic reader would care because this targets the practical regime where dense coverage is unavailable and fast, accurate reconstruction matters.

What carries the argument

The central object is the relative depth guidance loss, $L_{rdg} = \sum_{hw,ij} \log(1 + \exp^{-(D^{hw,ij}-b)\max(F^{hw,ij},0)})$, where $D^{hw,ij}$ is the cosine similarity between two patches of rendered depth and $F^{hw,ij}$ is the cosine similarity between the corresponding image-feature patches. This loss steers the Gaussian field toward view-consistent geometry by increasing image-feature similarity when depth similarity exceeds an adaptive threshold $b$ and decreasing it when depth similarity falls below the threshold. Supporting machinery includes an energy-function refinement of coarse monocular depth that injects global and local RGB structure into the depth prior, and an adaptive sampling strategy that densifies regions with high depth training error to accelerate convergence.

What would settle it

Train RDG-GS on a sparse-view scene with large untextured or repeating-pattern regions, such as a white wall with a poster, and measure rendered depth error against held-out views. If the relative depth guidance loss fails to reduce depth error compared with using only refined depth regularization, or actively degrades geometry because DINO features are not depth-ordered in such regions, the central claim that the loss produces view-consistent geometry would be refuted.

Watch

Extended reading notes

Core claim

RDG-GS's central claim is that the spatial relationships between patches, rather than absolute depth values, are the reliable geometric signal for sparse-view Gaussian optimization. The paper argues that monocular depth estimates are coarse and view-inconsistent, so supervising Gaussians with absolute depth pushes them toward wrong shapes. Instead, RDG-GS computes cosine similarities between patches of rendered depth and between patches of image features extracted by DINOv2, then aligns these two similarity tensors through a loss with an adaptive bias, encouraging patches that are nearby in depth to be nearby in feature space and distant patches to be pushed apart. Combined with refined depth priors and adaptive densification, the paper reports superior PSNR, SSIM, and LPIPS across Mip-NeRF360, LLFF, DTU, and Blender, while retaining real-time rendering.

Load-bearing premise

The load-bearing premise is that DINOv2 patch-feature similarity computed from rendered depth and images is a faithful proxy for 3D spatial proximity, so aligning image-feature similarity with depth similarity actually pushes Gaussians toward correct shapes; if DINO features do not preserve relative depth ordering, the loss can push geometry in the wrong direction.

Editorial extensions

If this is right

  • Sparse-view 3D Gaussian Splatting can achieve strong rendering quality without dense view coverage, with gains persisting from 3 to 24 training views and across multiple resolutions.
  • The method keeps rendering real-time at roughly 112 FPS, so the quality improvement does not sacrifice interactivity for applications such as virtual reality and autonomous driving.
  • The relative depth guidance is not tied to a specific monocular depth estimator; the paper shows consistent gains with different DPT variants, suggesting robustness to the choice of coarse-depth backbone.
  • Training remains fast at minutes per scene, roughly 20 to 40 times faster than NeRF-based sparse-view methods, making the approach practical for real-world use.
  • The paper's ablations attribute the improvements specifically to refined depth priors, relative depth guidance, and adaptive sampling, implying each module contributes meaningfully to the final result.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to apply the same relative-depth loss to NeRF-style volume rendering or to stereo-derived depth maps; because the loss only compares patch similarities, it may transfer without Gaussian-specific machinery.
  • The adaptive bias $b$ acts like a contrastive margin, and one could study its scheduling as a trade-off between geometry fidelity and texture-copy collapse, possibly making it per-patch rather than global.
  • The paper's own limitations, including artifacts in planar regions, mirror reflections, and added training time, point toward hierarchical or asymmetric Gaussian representations and reflection-aware features as the natural next steps.
  • A stronger validation would be to compare depth accuracy directly against a metric-depth ground truth, since the current evaluation is primarily through rendered image quality.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes RDG-GS, a sparse-view novel-view-synthesis method built on 3D Gaussian Splatting. The main contributions are (i) an energy-based refinement of monocular depth priors that injects global and fine-grained RGB structure into the depth used for regularization, (ii) a relative-depth-guidance loss that aligns patch-wise cosine similarities of rendered depth and image features, and (iii) an adaptive sampling strategy that densifies Gaussians in high-error regions. The authors report state-of-the-art PSNR/SSIM/LPIPS results on Mip-NeRF360, LLFF, DTU, and Blender, with real-time rendering speeds around 112 FPS. The paper also includes a detailed ablation study and additional robustness experiments over DPT variants, spherical-harmonic degree, initialization, and energy-function weights.

Significance. If the reported numbers are reliable, the method would be a meaningful advance for sparse-view 3DGS rendering, offering substantial quality gains over FSGS, CoR-GS, and geometry-aware 3DGS baselines while preserving real-time rendering. The paper deserves credit for its extensive ablation suite (Tables 9-13), for testing robustness across different DPT models and initialization strategies, and for explicitly listing limitations regarding training time, anisotropic Gaussians, and mirror reflections. I do not find the circularity concern raised in the stress-test note to be substantiated: the depth supervision comes from an external estimator and an energy refinement, and the relative-depth loss operates on rendered images and rendered depths rather than on the test target, so no step is fitted to the evaluation metric. The decisive weakness is that the central SOTA claim rests on an internally inconsistent evaluation protocol: conflicting baseline numbers and dataset descriptions prevent a reader from verifying the main contribution. The manuscript also provides no code, no per-scene breakdown, and no variance estimates, which at present makes the headline numbers uncheckable.

major comments (4)
  1. [Sec. 4.1.1 and Supplement, 'Introduction of Datasets'] The Mip-NeRF360 evaluation protocol is described inconsistently. The main text says 24 viewpoints from 7 scenes, while the supplement says 8 scenes and explicitly lists 'bear' as one of them; the official Mip-NeRF360 dataset contains 9 scenes and does not include a scene named 'bear'. The authors must state exactly which scenes are used, justify the subset, and ensure the main text and supplement agree, because every comparison in Tables 1, 2, 6, 7, and 9 depends on this protocol.
  2. [Tables 1, 2, and 7] The baseline numbers are mutually inconsistent, which makes the SOTA claim undecidable. For example, at 1/8 resolution Table 1 reports 3D-GS as PSNR 20.89 / SSIM 0.633 / LPIPS 0.317 and FSGS as 23.70 / 0.745 / 0.230, whereas Table 2 under the 24-view setting gives 3DGS 22.80 / 0.708 / 0.276 and FSGS 23.28 / 0.715 / 0.274 without specifying the resolution. In Table 7, the '3D-GS [26] None' row at 1/2 resolution reports exactly the same numbers as the FreeNeRF row (18.35 / 0.476 / 0.514), while Table 1 at 1/2 resolution reports 3D-GS as 17.12 / 0.476 / 0.514. These discrepancies directly affect the claimed advantage of 26.03 PSNR and must be resolved with one consistent protocol before the paper can be evaluated.
  3. [Sec. 4.1.1 and Tables 1-6] No variance or per-scene statistics are reported, although Sec. 4.1.1 states that the LLFF results are averaged over 10 experiments. With only 8, 15, or 9 scenes per dataset and 3-24 training views, a difference of 1-2 dB can fall within run-to-run noise, and the paper provides no way to assess this. The authors should report standard deviations or per-scene tables, and should clarify the resolution and train/test split for each table, including Table 5, where the 3-view LLFF 3D-GS number (19.22) conflicts with the value in Table 4 (17.83 at 503x381).
  4. [Sec. 3.3, Eq. (7)] The central relative-depth-guidance mechanism is motivated by an assumption that is not independently verified: that DINOv2 patch-feature similarity computed on depth maps is a faithful proxy for 3D spatial proximity, and that matching RGB patch similarity to depth patch similarity yields view-consistent geometry. The paper should provide a direct validation of this premise, for example an analysis on rendered depth maps showing that the depth patch-similarity tensor preserves relative depth ordering, or an ablation that replaces the DINO features with a simpler spatial-distance baseline. As written, the mechanism's success is only measured indirectly through the final rendering metrics.
minor comments (6)
  1. [Sec. 3.2, Eq. (3)] Eq. (3) writes Dr = arg max Dr E(Dr | I, Dc), while the surrounding text says the refined depth is obtained by 'minimizing the energy function E'; the sign convention should be made consistent.
  2. [Sec. 3.2.1 and Table 13] The notation for the energy-function weights is inconsistent: Eq. (4) uses wu, wp, wh, while Table 13 and Sec. 4.4.4 refer to wg and wh. Please align the notation.
  3. [Figure captions] Several figure captions contain typos or placeholder text: Fig. 3 uses 'RADG-GS' instead of RDG-GS, Figs. 5-8 use 'RDG-DS' instead of RDG-GS, and Fig. 4 contains the garbled string 'FPS 310 FPS221FPS0.03FPS'.
  4. [Table 2 and Table 8 captions] Table 2 cites the Mip-NeRF360 dataset as reference [45] but the correct reference is [3], and Table 8 cites the Blender dataset as [68] instead of [34]; the reference list also contains duplicate entries for DietNeRF ([23] and [24]).
  5. [Sec. 4.2.1 and Sec. 4.2.6] The speed-up claim is stated inconsistently: Sec. 4.2.1 says 'over 4000x faster' while Sec. 4.2.6 says 'over 3500x acceleration'; Table 7 implies about 3733x. Please use one consistent number.
  6. [Table 10] The variant name 'dpt large-384' is missing a hyphen and should read 'dpt-large-384' for consistency with the other entries.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the depth priors and relative-depth guidance are auxiliary regularizers driven by an external depth estimator (DPT) and model-internal consistency, not by the evaluation metrics.

full rationale

The derivation chain is self-contained rather than circular. Coarse depth D_c comes from the external monocular estimator F (DPT, Sec. 3.2); refined depth D_r is obtained by minimizing an energy function E(D_r | I, D_c) (Eqs. 3-5), and the rendered depth D_o is produced by the alpha-blend rasterizer (Eq. 2). The relative depth guidance loss L_rdg (Eq. 7) aligns image-feature patch similarity with depth-patch similarity on rendered outputs, and the final objective (Eq. 13) combines L_color against ground-truth training views, L_depth against the D_r pseudo-label, and L_rdg. No quantity in this chain is defined in terms of the evaluation metrics PSNR/SSIM/LPIPS, and no parameter is fitted to the test views; the hyperparameters are fixed constants reported in Sec. 4.1.2, with ablations in Tables 9-13. The closest loop is L_rdg, which matches two rendered quantities to each other, but it is an auxiliary regularizer and the paper validates it through held-out-view comparisons and ablations, so it is not a fitted prediction or a self-definition of the reported numbers. The serious protocol inconsistencies (7 vs 8 vs 9 Mip-NeRF360 scenes; conflicting baseline PSNR/SSIM in Tables 1 and 2; the duplicated 3D-GS row in Table 7) are correctness and reproducibility concerns, not circularity, and therefore do not raise the circularity score. The citation [55] for the CRF-style energy function is methodological borrowing, not a load-bearing self-citation chain.

Assumptions & free parameters 7 free parameters · 4 assumptions · 0 invented entities

The method does not introduce new physical entities. The free parameters are hyperparameters of the depth refinement energy, the relative depth guidance loss, and the training objective. Most are hand-set or tuned on the benchmark datasets, which weakens the claim of a parameter-free derivation.

free parameters (7)
  • Energy weight w_g (global structural consistency) = 10
    Chosen via Table 13 hyperparameter search; governs the global depth-image alignment term in Eq. 4.
  • Energy weight w_h (texture detail constraint) = 5
    Chosen via Table 13 hyperparameter search; balances high-frequency edge alignment in Eq. 4.
  • Energy weight w_u (local similarity baseline) = 1
    Set as baseline per Sec 4.4.4; not explicitly tuned.
  • Gaussian kernel stds theta_alpha, theta_mu, theta_beta = 35,10,10 and 10,2,2
    Hand-set in Sec 4.1.2 for coarse and fine-grained modules of Eq. 5.
  • High-frequency sensitivity tau and gamma = tau=5, gamma=10
    Hand-set in Sec 4.1.2; control the high-frequency weights g_u and g_p.
  • Relative depth guidance bias b (initial) = 0.4
    Hand-set in Sec 4.1.2; adapts via b(t)=b(t-1)|t/m| in Eq. 7.
  • Loss weights beta, lambda, omega = 0.4, 0.1, 0.05
    Set in Sec 4.1.2; omega is adaptive with an initial value of 0.05.
assumptions (4)
  • domain assumption Coarse depth from a pre-trained monocular depth estimator (DPT) is a useful starting point for sparse-view geometry.
    Sec 3.2 and Sec 4.4.1: the entire pipeline relies on this external prior; no analysis of failure cases for the estimator is given.
  • domain assumption DINOv2 features computed on a depth map encode meaningful spatial similarity.
    Eq. 6 and Sec 3.3 use the same feature extractor S for both image and depth patches; this is not validated independently.
  • domain assumption The RGB-guided energy function of [55] transfers correct geometry without texture-copy artifacts.
    Sec 3.2.1 and Eq. 4-5: the paper claims it suppresses redundant textures, but this depends on hyperparameters that are tuned on the evaluation datasets.
  • ad hoc to paper Uniform ray sampling within an adaptive depth range fills under-covered regions beneficially.
    Sec 3.4.1: no theoretical guarantee that patch-wise mean-depth-loss thresholding identifies the correct regions to densify.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering." pith.science (2026). https://pith.science/paper/XEPF4QPE

@misc{pith2026250111102,
  author       = {Pith},
  title        = {Pith review of: RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XEPF4QPE}},
  note         = {Machine review of arXiv:2501.11102}
}
read the original abstract

Efficiently synthesizing novel views from sparse inputs while maintaining accuracy remains a critical challenge in 3D reconstruction. While advanced techniques like radiance fields and 3D Gaussian Splatting achieve rendering quality and impressive efficiency with dense view inputs, they suffer from significant geometric reconstruction errors when applied to sparse input views. Moreover, although recent methods leverage monocular depth estimation to enhance geometric learning, their dependence on single-view estimated depth often leads to view inconsistency issues across different viewpoints. Consequently, this reliance on absolute depth can introduce inaccuracies in geometric information, ultimately compromising the quality of scene reconstruction with Gaussian splats. In this paper, we present RDG-GS, a novel sparse-view 3D rendering framework with Relative Depth Guidance based on 3D Gaussian Splatting. The core innovation lies in utilizing relative depth guidance to refine the Gaussian field, steering it towards view-consistent spatial geometric representations, thereby enabling the reconstruction of accurate geometric structures and capturing intricate textures. First, we devise refined depth priors to rectify the coarse estimated depth and insert global and fine-grained scene information to regular Gaussians. Building on this, to address spatial geometric inaccuracies from absolute depth, we propose relative depth guidance by optimizing the similarity between spatially correlated patches of depth and images. Additionally, we also directly deal with the sparse areas challenging to converge by the adaptive sampling for quick densification. Across extensive experiments on Mip-NeRF360, LLFF, DTU, and Blender, RDG-GS demonstrates state-of-the-art rendering quality and efficiency, making a significant advancement for real-world application.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 53 canonical work pages

  1. [26]

    Kerbl, G

    B. Kerbl, G. Kopanas, T. Leimk¨ uhler, and G. Drettakis. 3d gaussian splatting for real- time radiance field rendering. ACM Transac- tions on Graphics , 42(4), 2023

  2. [1]

    Avidan and A

    S. Avidan and A. Shashua. Novel view syn- thesis in tensor space. In Proceedings of IEEE Computer Society Conference on Com- puter Vision and Pattern Recognition , pages 1034–1040. IEEE, 1997

  3. [2]

    J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan. Mip-nerf: A multiscale represen- tation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 5855–5864, 2021. Springer Nature 2021 LATEX template 20 Article Title

  4. [3]

    J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5470–5479, 2022

  5. [4]

    J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman. Zip-nerf: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 19697–19705, 2023

  6. [5]

    S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. M¨ uller. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023

  7. [6]

    Birkl, D

    R. Birkl, D. Wofk, and M. M¨ uller. Midas v3. 1–a model zoo for robust monocular relative depth estimation. arXiv preprint arXiv:2307.14460, 2023

  8. [7]

    A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su. Tensorf: Tensorial radiance fields, 2022

Show all 68 references
  1. [8]

    A. Chen, Z. Xu, F. Zhao, X. Zhang, F. Xiang, J. Yu, and H. Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF international conference on computer vision , pages 14124–14133, 2021

  2. [9]

    D. Chen, H. Li, W. Ye, Y. Wang, W. Xie, S. Zhai, N. Wang, H. Liu, H. Bao, and G. Zhang. Pgsr: Planar-based gaus- sian splatting for efficient and high-fidelity surface reconstruction. arXiv preprint arXiv:2406.06521, 2024

  3. [10]

    T. Chen, P. Wang, Z. Fan, and Z. Wang. Aug- nerf: Training stronger neural radiance fields with triple-level physically-grounded aug- mentations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15191–15202, 2022

  4. [11]

    Cheng, Y.-P

    W. Cheng, Y.-P. Cao, and Y. Shan. Sparsegnv: Generating novel views of indoor scenes with sparse rgb-d images. In Pro- ceedings of the AAAI Conference on Artifi- cial Intelligence, volume 38, pages 1308–1316, 2024

  5. [12]

    Chung, J

    J. Chung, J. Oh, and K. M. Lee. Depth- regularized optimization for 3d gaussian splatting in few-shot images. arXiv preprint arXiv:2311.13398, 2023

  6. [13]

    W. Cong, H. Liang, P. Wang, Z. Fan, T. Chen, M. Varma, Y. Wang, and Z. Wang. Enhancing nerf akin to enhanc- ing llms: Generalizable nerf transformer with mixture-of-view-experts. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 3193–3204, 2023

  7. [14]

    K. Deng, A. Liu, J.-Y. Zhu, and D. Ramanan. Depth-supervised nerf: Fewer views and faster training for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12882–12891, 2022

  8. [15]

    Fridovich-Keil, G

    S. Fridovich-Keil, G. Meanti, F. R. War- burg, B. Recht, and A. Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12479–12488, 2023

  9. [16]

    Fridovich-Keil, A

    S. Fridovich-Keil, A. Yu, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5501–5510, 2022

  10. [17]

    K. Gao, Y. Gao, H. He, D. Lu, L. Xu, and J. Li. Nerf: Neural radiance field in 3d vision, a comprehensive review. arXiv preprint arXiv:2210.00379, 2022

  11. [18]

    S. J. Garbin, M. Kowalski, M. Johnson, J. Shotton, and J. Valentin. Fastnerf: High- fidelity neural rendering at 200fps, 2021

  12. [19]

    Gu´ edon and V

    A. Gu´ edon and V. Lepetit. Sugar: Surface- aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering. In Proceedings of the IEEE/CVF Springer Nature 2021 LATEX template Article Title 21 Conference on Computer Vision and Pattern Recognitio...

  13. [20]

    S. Guo, Q. Wang, Y. Gao, R. Xie, and L. Song. Depth-guided robust and fast point cloud fusion nerf for sparse input views. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 1976– 1984, 2024

  14. [21]

    Y.-C. Guo, D. Kang, L. Bao, Y. He, and S.-H. Zhang. Nerfren: Neural radiance fields with reflections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18409–18418, 2022

  15. [22]

    Huang, Z

    B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao. 2d gaussian splatting for geomet- rically accurate radiance fields. In ACM SIGGRAPH 2024 Conference Papers , pages 1–11, 2024

  16. [24]

    A. Jain, M. Tancik, and P. Abbeel. Putting nerf on a diet: Semantically consistent few- shot view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 5885–5894, 2021

  17. [25]

    Jensen, A

    R. Jensen, A. Dahl, G. Vogiatzis, E. Tola, and H. Aanæs. Large scale multi-view stereop- sis evaluation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition , pages 406–413, 2014

  18. [27]

    Khosla, P

    P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan. Supervised contrastive learning. Advances in neural information processing systems, 33:18661–18673, 2020

  19. [28]

    M. Kim, S. Seo, and B. Han. Infonerf: Ray entropy minimization for few-shot neu- ral volume rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12912–12921, 2022

  20. [29]

    C. Li, B. Y. Feng, Y. Liu, H. Liu, C. Wang, W. Yu, and Y. Yuan. Endosparse: Real-time sparse view synthesis of endoscopic scenes using gaussian splatting, 2024

  21. [30]

    J. Li, J. Zhang, X. Bai, J. Zheng, X. Ning, J. Zhou, and L. Gu. Dngaussian: Optimiz- ing sparse-view 3d gaussian radiance fields with global-local depth normalization. arXiv preprint arXiv:2403.06912, 2024

  22. [31]

    Z. Li, Z. Chen, Z. Li, and Y. Xu. Space- time gaussian feature splatting for real-time dynamic view synthesis. arXiv preprint arXiv:2312.16812, 2023

  23. [32]

    L. Liu, J. Gu, K. Zaw Lin, T.-S. Chua, and C. Theobalt. Neural sparse voxel fields. Advances in Neural Information Processing Systems, 33:15651–15663, 2020

  24. [33]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, R. Ortiz- Cayon, N. K. Kalantari, R. Ramamoorthi, R. Ng, and A. Kar. Local light field fusion: Practical view synthesis with prescriptive sampling guidelines. ACM Transactions on Graphics (TOG), 38(4):1–14, 2019

  25. [34]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021

  26. [35]

    M¨ uller, A

    T. M¨ uller, A. Evans, C. Schied, and A. Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM trans- actions on graphics (TOG), 41(4):1–15, 2022

  27. [36]

    Niedermayr, J

    S. Niedermayr, J. Stumpfegger, and R. West- ermann. Compressed 3d gaussian splatting for accelerated novel view synthesis. arXiv preprint arXiv:2401.02436, 2023. Springer Nature 2021 LATEX template 22 Article Title

  28. [37]

    Niemeyer, J

    M. Niemeyer, J. T. Barron, B. Mildenhall, M. S. Sajjadi, A. Geiger, and N. Radwan. Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5480–5490, 2022

  29. [38]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernan- dez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual fea- tures without supervision. arXiv preprint arXiv:2304.07193, 2023

  30. [39]

    Rabby and C

    A. Rabby and C. Zhang. Beyondpixels: A comprehensive review of the evolution of neural radiance fields. arXiv preprint arXiv:2306.03000, 2023

  31. [40]

    Ranftl, A

    R. Ranftl, A. Bochkovskiy, and V. Koltun. Vision transformers for dense prediction, 2021

  32. [41]

    Roessle, J

    B. Roessle, J. T. Barron, B. Mildenhall, P. P. Srinivasan, and M. Nießner. Dense depth priors for neural radiance fields from sparse input views. CoRR, abs/2112.03288, 2021

  33. [42]

    J. L. Schonberger and J.-M. Frahm. Structure-from-motion revisited. In Proceed- ings of the IEEE conference on computer vision and pattern recognition , pages 4104–4113, 2016

  34. [43]

    Schwarz, A

    K. Schwarz, A. Sauer, M. Niemeyer, Y. Liao, and A. Geiger. Voxgraf: Fast 3d-aware image synthesis with sparse voxel grids. Advances in Neural Information Processing Systems , 35:33999–34011, 2022

  35. [44]

    Z. Shao, Z. Wang, Z. Li, D. Wang, X. Lin, Y. Zhang, M. Fan, and Z. Wang. Splattinga- vatar: Realistic real-time human avatars with mesh-embedded gaussian splatting. arXiv preprint arXiv:2403.05087, 2024

  36. [45]

    Somraj, A

    N. Somraj, A. Karanayil, and R. Soundarara- jan. Simplenerf: Regularizing sparse input neural radiance fields with simpler solu- tions. In SIGGRAPH Asia 2023 Conference Papers, pages 1–11, 2023

  37. [46]

    Somraj and R

    N. Somraj and R. Soundararajan. Vip- nerf: Visibility prior for sparse input neural radiance fields. In ACM SIGGRAPH 2023 Conference Proceedings, pages 1–11, 2023

  38. [47]

    J. Song, S. Park, H. An, S. Cho, M.-S. Kwak, S. Cho, and S. Kim. D \” arf: Boost- ing radiance fields from sparse inputs with monocular depth adaptation. arXiv preprint arXiv:2305.19201, 2023

  39. [48]

    Suhail, C

    M. Suhail, C. Esteves, L. Sigal, and A. Maka- dia. Light field neural rendering. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8269–8279, 2022

  40. [49]

    C. Sun, M. Sun, and H. Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In CVPR, 2022

  41. [50]

    C. Sun, M. Sun, and H.-T. Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction, 2022

  42. [51]

    Tewari, J

    A. Tewari, J. Thies, B. Mildenhall, P. Srini- vasan, E. Tretschk, W. Yifan, C. Lassner, V. Sitzmann, R. Martin-Brualla, S. Lom- bardi, et al. Advances in neural rendering. In Computer Graphics Forum, volume 41, pages 703–735. Wiley Online Library, 2022

  43. [52]

    M. A. Uy, R. Martin-Brualla, L. Guibas, and K. Li. Scade: Nerfs from space carving with ambiguity-aware depth estimates. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16518–16527, 2023

  44. [53]

    C. Wang, M. Chai, M. He, D. Chen, and J. Liao. Clip-nerf: Text-and-image driven manipulation of neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 3835–3844, 2022

  45. [54]

    G. Wang, Z. Chen, C. C. Loy, and Z. Liu. Sparsenerf: Distilling depth ranking for few- shot novel view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 9065–9076, 2023. Springer Nature 2021 LATEX template Article Title 23

  46. [55]

    H. Wang, M. Yang, C. Zhu, and N. Zheng. Rgb-guided depth map recovery by two-stage coarse-to-fine dense crf models. IEEE Trans- actions on Image Processing , 32:1315–1328, 2023

  47. [56]

    L. Wang, J. Zhang, X. Liu, F. Zhao, Y. Zhang, Y. Zhang, M. Wu, J. Yu, and L. Xu. Fourier plenoctrees for dynamic radiance field rendering in real-time. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 13524–13534, 2022

  48. [57]

    P. Wang, Y. Liu, Z. Chen, L. Liu, Z. Liu, T. Komura, C. Theobalt, and W. Wang. F2- nerf: Fast neural radiance field training with free camera trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4150–4159, 2023

  49. [58]

    R. Wu, B. Mildenhall, P. Henzler, K. Park, R. Gao, D. Watson, P. P. Srinivasan, D. Verbin, J. T. Barron, B. Poole, et al. Reconfusion: 3d reconstruction with diffusion priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21551–21561, 2024

  50. [59]

    Wynn and D

    J. Wynn and D. Turmukhambetov. Diffu- sionerf: Regularizing neural radiance fields with denoising diffusion models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 4180–4189, 2023

  51. [60]

    Xiong, S

    H. Xiong, S. Muttukuru, R. Upadhyay, P. Chari, and A. Kadambi. Sparsegs: Real- time 360 sparse view synthesis using gaussian splatting. 2023

  52. [61]

    C. Yang, S. Li, J. Fang, R. Liang, L. Xie, X. Zhang, W. Shen, and Q. Tian. Gaus- sianobject: Just taking four images to get a high-quality 3d object with gaussian splat- ting, 2024

  53. [62]

    J. Yang, M. Pavone, and Y. Wang. Freen- erf: Improving few-shot neural rendering with free frequency regularization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8254– 8263, 2023

  54. [63]

    Z. Yang, X. Gao, W. Zhou, S. Jiao, Y. Zhang, and X. Jin. Deformable 3d gaussians for high-fidelity monocular dynamic scene recon- struction. arXiv preprint arXiv:2309.13101 , 2023

  55. [64]

    R. Yin, V. Yugay, Y. Li, S. Karaoglu, and T. Gevers. Fewviewgs: Gaussian splatting with few view matching and multi-stage training. arXiv preprint arXiv:2411.02229 , 2024

  56. [65]

    A. Yu, V. Ye, M. Tancik, and A. Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021

  57. [66]

    Z. Yu, T. Sattler, and A. Geiger. Gaus- sian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes. ACM Transactions on Graphics (TOG) , 43(6):1– 13, 2024

  58. [67]

    Zhang, J

    J. Zhang, J. Li, X. Yu, L. Huang, L. Gu, J. Zheng, and X. Bai. Cor-gs: sparse-view 3d gaussian splatting via co-regularization. In European Conference on Computer Vision , pages 335–352. Springer, 2025

  59. [68]

    Zhang, P

    R. Zhang, P. Isola, A. A. Efros, E. Shecht- man, and O. Wang. The unreasonable effec- tiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 586–595, 2018

  60. [69]

    counter”, “room

    Z. Zhu, Z. Fan, Y. Jiang, and Z. Wang. Fsgs: Real-time few-shot view synthesis using gaussian splatting. arXiv preprint arXiv:2312.00451, 2023. Springer Nature 2021 LATEX template 24 Article Title Supplemental Detail Theoretical Analysis Our refined depth restoration hinges on...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.