REVIEW 4 major objections 7 minor 2 cited by
SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction
T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read SPARS3R fuses a dense depth-prior point cloud with SfM-accurate poses through global Procrustes alignment and semantic local fixes, producing a Gaussian-splatting initialization that outperforms prior sparse-view NVS methods by about 2.7…
desk verdict Solid sparse-view NVS system: the semantic outlier alignment is genuinely new, but most of the reported gain comes from global fusion with SfM poses, and 'consistent' needs per-scene support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is global-plus-semantic alignment. Global Fusion Alignment uses RANSAC-robust Procrustes analysis on forward-backward projected correspondences to map the dense prior onto the sparse SfM cloud, producing a single scale, rotation, and translation. Semantic Outlier Alignment then iteratively seeds masks from the Segment-Anything Model at outlier projections, requiring each mask to contain at least a threshold number $T$ of SfM outliers before estimating a per-mask rigid transform. The output is a fused point cloud $\chi^*$ concatenated with the SfM point cloud, which serves as the initialization for a Gaussian splatting optimizer. The intended effect is that accurate SfM poses and dense depth coverage coexist, so optimization avoids floaters and background blur.
What would settle it
On a scene with a large smooth curved surface, measure the residual alignment error inside a single semantic mask after Semantic Outlier Alignment; if the error grows with depth bias or the mask lacks SfM support, the piecewise-rigid assumption fails.
Extended reading notes
Core claim
The discovery is a two-stage alignment recipe. Given a dense point cloud from MASt3R/DUSt3R and a sparse point cloud from COLMAP SfM, correspondences are obtained by projecting SfM points into the dense cloud's coordinate frame, and a RANSAC-filtered Procrustes fit provides a global transform. Points that remain outliers are grouped by prompting the Segment-Anything Model with their projected 2D locations; each semantic mask is then given its own local rigid transform estimated from the SfM points inside it. The transformed dense cloud is concatenated with the SfM cloud and used to initialize Gaussian splatting. The paper reports that this fused initialization consistently outperforms existing sparse-view NVS methods across MipNeRF360, Tanks & Temples, and MVimgNet, with an average PSNR gain of about 2.7 dB.
Load-bearing premise
The load-bearing premise is that the depth errors left after a global rigid fit are rigid per semantically coherent region, and every such region has enough SfM points to estimate its own local transform.
Editorial extensions
If this is right
- Sparse-view Gaussian splatting can be initialized with both dense coverage and SfM-accurate poses instead of trading one against the other.
- Residual depth bias in learned dense clouds can be corrected locally by grouping outlier regions with semantic masks, as long as each group has enough SfM support.
- Sparse-NVS evaluation becomes fairer when test camera poses are aligned with rotation-aware RANSAC and quality is measured with a pose-shift-robust metric such as DSIM.
- On the three tested benchmarks, the method reports the best sparse-view NVS numbers, with the largest gain over the strongest prior baseline on MipNeRF360.
Reading between the lines
- Editorial inference: The same global-plus-semantic alignment recipe should transfer to other dense depth estimators beyond MASt3R/DUSt3R, provided their depth errors concentrate at object boundaries; swapping the prior and rerunning the fusion would test this.
- Editorial inference: The evaluation changes imply that earlier sparse-NVS numbers relying on PSNR/SSIM under imperfect test-pose alignment may undervalue methods with accurate geometry; DSIM plus rotation-aware camera alignment could become a fairer default.
- Editorial inference: A natural extension, named by the paper as future work, is replacing the rigid-per-mask transform with a smooth non-rigid deformation inside each semantic region; an intermediate test is to fit locally affine rather than rigid transforms per mask.
- Editorial inference: Because the fused cloud inherits SfM's scale and pose accuracy, it could serve as a metric-scale geometric prior for robotics or navigation, not only as a rendering initialization.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SPARS3R is a sparse-view 3D reconstruction and novel view synthesis method. It combines a sparse point cloud and camera poses from COLMAP SfM (run with MASt3R feature matches) with a dense point cloud from MASt3R/DUSt3R. In the first stage, Global Fusion Alignment estimates a global similarity transform (Procrustes with RANSAC) that maps the dense prior point cloud into the SfM coordinate frame. In the second stage, Semantic Outlier Alignment groups the remaining alignment outliers by prompting SAM with outlier points, and estimates a local rigid transform per semantic mask; the locally transformed points are merged with the SfM cloud to initialize a 3DGS/Splatfacto optimization. The paper also contributes evaluation modifications: a rotation-point-aware Procrustes alignment with RANSAC for camera pose evaluation, and the use of DreamSim (DSIM) as an additional perceptual metric. Experiments on MipNeRF360, Tanks & Temples, and MVimgNet report improved PSNR, SSIM, LPIPS, and DSIM over existing sparse-view NVS methods, with the largest gains on MipNeRF360 (e.g., 18.85 vs. 16.23 PSNR over InstantSplat).
Significance. If the results hold, SPARS3R is a practical and conceptually simple way to combine the dense geometric prior of DUSt3R/MASt3R with the accurate camera poses of SfM, and the reported margins over InstantSplat and FSGS on three benchmarks are substantial. The paper ships code, builds on publicly available components, and includes a useful evaluation improvement (rotation-point camera alignment with RANSAC). However, the novel Semantic Outlier Alignment step is currently supported by only a 0.3 dB average PSNR gain in a one-dataset ablation, and the absence of per-scene statistics and error bars weakens the claim that SPARS3R 'consistently' outperforms prior methods. The test-pose-optimization protocol and the non-standard DSIM metric also require clearer justification before the quantitative claims can be fully assessed.
major comments (4)
- [4.2 / Table 2] The only ablation of Semantic Outlier Alignment is the MipNeRF360 average, where the full SPARS3R pipeline reaches 18.9 dB PSNR versus 18.6 dB for Global Fusion Alignment alone. The text mentions a 1.4 dB gain on the Bonsai scene, but no per-scene breakdown or variance estimate is given for the nine MipNeRF360 scenes or for the other two datasets. Because Section 4.3 claims that SPARS3R 'consistently improves' over prior methods, and Table 4 reports only dataset averages over 24 scenes, the evidence does not yet rule out the possibility that SOA helps one or two scenes and slightly hurts others. Please provide per-scene results, standard deviations or error bars, and ablations on Tanks & Temples and MVimgNet.
- [3.2.2 / Algorithm 1] The RANSAC parameters (sample size n, number of iterations R, error threshold epsilon) and the mask support threshold T in Eq. (5) are never specified. These parameters directly determine the inlier/outlier partition and which semantic masks are accepted for local alignment, so they have a direct effect on the fused point cloud and on the reported render metrics. Without these values (or a sensitivity analysis), the experiments are not reproducible, and the 0.3 dB SOA gain cannot be assessed for threshold dependence.
- [3.2.2 / Section 4.4] The piecewise-rigid assumption per SAM mask is the core of Semantic Outlier Alignment, but the manuscript's own limitations state that masks can split connected surfaces (leaving insufficient SfM support) or group areas with disparate depths (making local alignment ineffective). The paper gives no indication of how often these failure modes occur across the 24 scenes, no sensitivity to SAM mask granularity, and no diagnostic of the fraction of outliers that are discarded rather than aligned. Please quantify these cases and their effect on the final render metrics.
- [4.1 / Table 4] All methods are evaluated after 500 steps of test-pose optimization using the test images. This post-processing fits renders to the ground-truth test views and can differentially inflate PSNR/SSIM across methods, so the headline numbers in Table 4 may reflect pose-fitting ability rather than pure NVS quality. Please report metrics with and without test-pose optimization, and specify exactly what is optimized and whether ground-truth test images are used in that optimization.
minor comments (7)
- [Section 1, Contribution list] The word 'Gloabl' should be 'Global'.
- [Section 3.2.1] The word 'btween' should be 'between'.
- [Figure 1 caption] The method name is misspelled as 'SPAS3R'; it should be 'SPARS3R'.
- [Section 3.2.1] The notation for the visibility indicator, 'V in R^{N*3} -> R^{N*3}', is unclear; please define V as a function or set-valued map.
- [Section 4.1 / Figure 3] The claim that DSIM is the most pose-shift invariant metric is based only on the qualitative curves in Figure 3; please provide numeric values or a table.
- [Table 3] SPARS3R itself is not listed in Table 3; since SPARS3R uses COLMAP poses, please state explicitly that its pose accuracy is that of the COLMAP+MASt3R row.
- [Table 4 header] The dataset name is misspelled as 'MipsNeRF360'; it should be 'MipNeRF360'.
Circularity Check
No significant circularity: SPARS3R's components are independently fitted to reconstruction inputs and evaluated on held-out views.
full rationale
SPARS3R's derivation chain is self-contained rather than circular. The dense point cloud comes from the external pretrained MASt3R/DUSt3R models, the sparse reference point cloud and poses come from COLMAP SfM, and the semantic grouping comes from the external SAM model. The global and local rigid alignments in Eqs. (4), (6), and (7) are fitted to SfM correspondences, not to the reported novel-view-synthesis metrics. The rendering quality is then evaluated on held-out test views after Gaussian optimization, so no fitted parameter is renamed as a prediction. The paper's introduction of DSIM as an evaluation metric is not circular: DreamSim is an external pretrained similarity model, the method's parameters are not optimized against DSIM, and PSNR, SSIM, and LPIPS also improve in Table 4. Section 4.4 honestly states limitations of the semantic-mask and piecewise-rigid assumptions; those are robustness caveats, not reductions of the method's output to its inputs. No load-bearing self-citations, imported uniqueness theorems, or ansatz-by-citation steps are present. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- RANSAC inlier threshold epsilon
- Outlier mask threshold T
- RANSAC sample size n and iterations R
- Loss weights lambda1 and lambda2 =
0.8 and 0.2
assumptions (4)
- domain assumption The dense point cloud from DUSt3R/MASt3R is locally coherent but globally biased toward smoothness between objects.
- domain assumption The SfM point cloud from COLMAP with MASt3R features has accurate poses and reliable geometry under sparse views.
- ad hoc to paper Scene regions corresponding to SAM masks each follow a single rigid transformation.
- domain assumption SAM segmentation groups outliers into semantically coherent regions with consistent depth bias.
Cite this review
Pith. "Pith review of SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction." pith.science (2026). https://pith.science/paper/AFAJX4IS
@misc{pith2026241112592,
author = {Pith},
title = {Pith review of: SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction},
year = {2026},
howpublished = {\url{https://pith.science/paper/AFAJX4IS}},
note = {Machine review of arXiv:2411.12592}
}
read the original abstract
Recent efforts in Gaussian-Splat-based Novel View Synthesis can achieve photorealistic rendering; however, such capability is limited in sparse-view scenarios due to sparse initialization and over-fitting floaters. Recent progress in depth estimation and alignment can provide dense point cloud with few views; however, the resulting pose accuracy is suboptimal. In this work, we present SPARS3R, which combines the advantages of accurate pose estimation from Structure-from-Motion and dense point cloud from depth estimation. To this end, SPARS3R first performs a Global Fusion Alignment process that maps a prior dense point cloud to a sparse point cloud from Structure-from-Motion based on triangulated correspondences. RANSAC is applied during this process to distinguish inliers and outliers. SPARS3R then performs a second, Semantic Outlier Alignment step, which extracts semantically coherent regions around the outliers and performs local alignment in these regions. Along with several improvements in the evaluation process, we demonstrate that SPARS3R can achieve photorealistic rendering with sparse images and significantly outperforms existing approaches.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
AdaptiveSplat:Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction
Texture-aware SuperCluster pruning plus an adaptive Gaussian head lets feed-forward 3DGS models hit a user budget β while outperforming post-hoc pruners on RE10K, ACID, DL3DV and DTU.
-
DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring
A pose-free deblurring 3D Gaussian Splatting pipeline using DUSt3R point clouds, confidence-balanced sampling, and event-decoded latent image supervision.
Reference graph
Works this paper leans on
-
[1]
Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields
Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 5855–5864,
-
[2]
Mip-nerf 360: Unbounded anti-aliased neural radiance fields
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5470–5479, 2022. 5, 6, 7, 8
work page 2022
-
[3]
Zip-nerf: Anti-aliased grid-based neural radiance fields
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Zip-nerf: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 19697–19705, 2023. 2
2023
-
[4]
Hexplane: A fast representa- tion for dynamic scenes
Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 130–141, 2023. 2
2023
-
[5]
Tensorf: Tensorial radiance fields
Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. InEuropean con- ference on computer vision, pages 333–350. Springer, 2022. 2
2022
-
[6]
Nerf-hugs: Improved neural radiance fields in non-static scenes using heuristics-guided segmentation
Jiahao Chen, Yipeng Qin, Lingjie Liu, Jiangbo Lu, and Guanbin Li. Nerf-hugs: Improved neural radiance fields in non-static scenes using heuristics-guided segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 19436–19446, 2024. 2
2024
-
[7]
Tianlong Chen, Peihao Wang, Zhiwen Fan, and Zhangyang Wang. Aug-nerf: Training stronger neural radiance fields with triple-level physically-grounded augmentations. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15191–15202, 2022. 2
work page 2022
-
[8]
Hallucinated neural radiance fields in the wild
Xingyu Chen, Qi Zhang, Xiaoyu Li, Yue Chen, Ying Feng, Xuan Wang, and Jue Wang. Hallucinated neural radiance fields in the wild. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 12943–12952, 2022. 2
2022
Show all 61 references
-
[9]
Gaussianpro: 3d gaussian splatting with progressive propagation
Kai Cheng, Xiaoxiao Long, Kaizhi Yang, Yao Yao, Wei Yin, Yuexin Ma, Wenping Wang, and Xuejin Chen. Gaussianpro: 3d gaussian splatting with progressive propagation. InForty- first International Conference on Machine Learning, 2024. 2
2024
-
[10]
Depth-regularized optimization for 3d gaussian splatting in few-shot images
Jaeyoung Chung, Jeongtaek Oh, and Kyoung Mu Lee. Depth-regularized optimization for 3d gaussian splatting in few-shot images. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 811–820, 2024. 7
2024
-
[11]
Depth-regularized optimization for 3d gaussian splatting in few-shot images
Jaeyoung Chung, Jeongtaek Oh, and Kyoung Mu Lee. Depth-regularized optimization for 3d gaussian splatting in few-shot images. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 811–820, 2024. 2, 7
2024
-
[12]
Depth-supervised nerf: Fewer views and faster train- ing for free
Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ra- manan. Depth-supervised nerf: Fewer views and faster train- ing for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12882– 12891, 2022. 1, 5
2022
-
[13]
Lightgaussian: Unbounded 3d gaussian compression with 15x reduction and 200+ fps
Zhiwen Fan, Kevin Wang, Kairun Wen, Zehao Zhu, De- jia Xu, and Zhangyang Wang. Lightgaussian: Unbounded 3d gaussian compression with 15x reduction and 200+ fps. arXiv preprint arXiv:2311.17245, 2023. 2
2023 arXiv
-
[14]
Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds
Zhiwen Fan, Wenyan Cong, Kairun Wen, Kevin Wang, Jian Zhang, Xinghao Ding, Danfei Xu, Boris Ivanovic, Marco Pavone, Georgios Pavlakos, et al. Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds. arXiv preprint arXiv:2403.20309, 2, 2024. 1, 2, 5, 6, 7
2024 arXiv
-
[15]
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Communications of the ACM, 24(6):381–395, 1981
Martin A Fischler and Robert C Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Communications of the ACM, 24(6):381–395, 1981. 2, 5
1981
-
[16]
Plenoxels: Radiance fields without neural networks
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5501–5510, 2022. 2
2022
-
[17]
K-planes: Explicit radiance fields in space, time, and appearance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 12479–12488, 2023. 2
2023
-
[18]
Dream- sim: Learning new dimensions of human visual similarity using synthetic data, 2023
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similarity using synthetic data, 2023. 6
2023
-
[19]
Efros, and Xiaolong Wang
Yang Fu, Sifei Liu, Amey Kulkarni, Jan Kautz, Alexei A. Efros, and Xiaolong Wang. Colmap-free 3d gaussian splat- ting. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 20796– 20805, 2024. 7
2024
-
[20]
Fastnerf: High-fidelity neu- ral rendering at 200fps
Stephan J Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien Valentin. Fastnerf: High-fidelity neu- ral rendering at 200fps. In Proceedings of the IEEE/CVF international conference on computer vision , pages 14346– 14355, 2021. 2
2021
-
[21]
Ea- gles: Efficient accelerated 3d gaussians with lightweight en- codings
Sharath Girish, Kamal Gupta, and Abhinav Shrivastava. Ea- gles: Efficient accelerated 3d gaussians with lightweight en- codings. arXiv preprint arXiv:2312.04564, 2023. 2
2023 arXiv
-
[22]
Generalized procrustes analysis
John C Gower. Generalized procrustes analysis. Psychome- trika, 40:33–51, 1975. 4, 5
1975
-
[23]
Putting nerf on a diet: Semantically consistent few-shot view synthesis
Ajay Jain, Matthew Tancik, and Pieter Abbeel. Putting nerf on a diet: Semantically consistent few-shot view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5885–5894, 2021. 1, 5
2021
-
[24]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,
-
[25]
Infonerf: Ray entropy minimization for few-shot neural volume ren- dering
Mijeong Kim, Seonguk Seo, and Bohyung Han. Infonerf: Ray entropy minimization for few-shot neural volume ren- dering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 12912– 12921, 2022. 1, 5
2022
-
[26]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023. 2, 6
2023 arXiv
-
[27]
Tanks and temples: Benchmarking large-scale scene reconstruction
Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (ToG) , 36 (4):1–13, 2017. 5, 7
2017
-
[28]
Compact 3d gaussian representation for radiance field
Joo Chan Lee, Daniel Rho, Xiangyu Sun, Jong Hwan Ko, and Eunbyung Park. Compact 3d gaussian representation for radiance field. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21719– 21728, 2024. 2
2024
-
[29]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and J´erˆome Revaud. Ground- ing image matching in 3d with mast3r. arXiv preprint arXiv:2406.09756, 2024. 2, 3, 4, 5, 6
2024 arXiv
-
[31]
Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normaliza- tion
Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu. Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normaliza- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 2...
-
[32]
Barf: Bundle-adjusting neural radiance fields
Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Si- mon Lucey. Barf: Bundle-adjusting neural radiance fields. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5741–5751, 2021. 6
2021
-
[33]
Scaffold-gs: Structured 3d gaussians for view-adaptive rendering
Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. Scaffold-gs: Structured 3d gaussians for view-adaptive rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20654–20664, 2024. 2
2024
-
[34]
Nerf in the wild: Neural radiance fields for uncon- strained photo collections
Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duck- worth. Nerf in the wild: Neural radiance fields for uncon- strained photo collections. In Proceedings of the IEEE/CVF conference on computer vision and pattern recogn...
2021
-
[35]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM, 65(1):99–106, 2021. 1, 2, 7
2021
-
[36]
Instant neural graphics primitives with a mul- tiresolution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a mul- tiresolution hash encoding. ACM transactions on graphics (TOG), 41(4):1–15, 2022. 1, 2, 7
2022
-
[37]
Compgs: Smaller and faster gaussian splatting with vector quantization
KL Navaneet, Kossar Pourahmadi Meibodi, Soroush Abbasi Koohpayegani, and Hamed Pirsiavash. Compgs: Smaller and faster gaussian splatting with vector quantization. In Euro- pean Conference on Computer Vision, 2024. 2
2024
-
[38]
Compressed 3d gaussian splatting for accelerated novel view synthesis
Simon Niedermayr, Josef Stumpfegger, and R ¨udiger West- ermann. Compressed 3d gaussian splatting for accelerated novel view synthesis. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 10349–10358, 2024. 2
2024
-
[39]
Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs
Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition,...
2022
-
[40]
Radsplat: Radiance field-informed gaussian splat- ting for robust real-time rendering with 900+ fps
Michael Niemeyer, Fabian Manhardt, Marie-Julie Rakoto- saona, Michael Oechsle, Daniel Duckworth, Rama Gosula, Keisuke Tateno, John Bates, Dominik Kaeser, and Federico Tombari. Radsplat: Radiance field-informed gaussian splat- ting for robust real-time rendering with 900+ fps. ...
2024 arXiv
-
[41]
Coherentgs: Sparse novel view synthesis with coherent 3d gaussians
Avinash Paliwal, Wei Ye, Jinhui Xiong, Dmytro Kotovenko, Rakesh Ranjan, Vikas Chandra, and Nima Khademi Kalan- tari. Coherentgs: Sparse novel view synthesis with coherent 3d gaussians. In European Conference on Computer Vision, pages 19–37. Springer, 2025. 2
2025
-
[42]
Nerfies: Deformable neural radiance fields
Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5865–5874, 2021. 2
2021
-
[43]
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 10318–10327, 2021. 2
2021
-
[44]
Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps
Christian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps. In Proceedings of the IEEE/CVF international conference on computer vision , pages 14335– 14345, 2021. 2
2021
-
[45]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Com- puter Vision and Pattern Recognition (CVPR) , 2016. 2, 5, 6
2016
-
[46]
Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering
Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu, Hongwen Zhang, and Yebin Liu. Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 166...
2023
-
[47]
Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction
Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 5459– 5469, 2022. 2
2022
-
[48]
Nerfstudio: A modular framework for neural radiance field development
Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. In ACM SIGGRAPH 2023 Conference Proceedings , pages 1–12...
2023
-
[49]
Ref-nerf: Struc- tured view-dependent appearance for neural radiance fields
Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T Barron, and Pratul P Srinivasan. Ref-nerf: Struc- tured view-dependent appearance for neural radiance fields. In 2022 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 5481–5490. IE...
2022
-
[50]
Sparsenerf: Distilling depth ranking for few-shot novel view synthesis
Guangcong Wang, Zhaoxi Chen, Chen Change Loy, and Zi- wei Liu. Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 9065–9076,
-
[51]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20697– 20709, 2024. 1, 2, 3, 4, 6, 7
2024
-
[52]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 5
2004
-
[53]
Dˆ 2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video
Tianhao Wu, Fangcheng Zhong, Andrea Tagliasacchi, For- rester Cole, and Cengiz Oztireli. Dˆ 2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video. Advances in neural information processing systems , 35:32653–32666, 2022. 2
2022
-
[54]
Sparsegs: Real- time 360 tzdegu sparse view synthesis using gaussian splat- ting
Haolin Xiong, Sairisheek Muttukuru, Rishi Upadhyay, Pradyumna Chari, and Achuta Kadambi. Sparsegs: Real- time 360 tzdegu sparse view synthesis using gaussian splat- ting. arXiv preprint arXiv:2312.00206, 2023. 1, 2, 7
2023 arXiv
-
[55]
Multi-scale 3d gaussian splatting for anti-aliased rendering
Zhiwen Yan, Weng Fei Low, Yu Chen, and Gim Hee Lee. Multi-scale 3d gaussian splatting for anti-aliased rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20923–20931, 2024. 2
2024
-
[56]
Freenerf: Im- proving few-shot neural rendering with free frequency reg- ularization
Jiawei Yang, Marco Pavone, and Yue Wang. Freenerf: Im- proving few-shot neural rendering with free frequency reg- ularization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8254–8263,
-
[57]
Cross-ray neural radiance fields for novel- view synthesis from unconstrained image collections
Yifan Yang, Shuhai Zhang, Zixiong Huang, Yubing Zhang, and Mingkui Tan. Cross-ray neural radiance fields for novel- view synthesis from unconstrained image collections. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15901–15911, 2023. 2
2023
-
[58]
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 4578–4587, 2021. 5
2021
-
[59]
Mvimgnet: A large-scale dataset of multi-view images
Xianggang Yu, Mutian Xu, Yidan Zhang, Haolin Liu, Chongjie Ye, Yushuang Wu, Zizheng Yan, Chenming Zhu, Zhangyang Xiong, Tianyou Liang, et al. Mvimgnet: A large-scale dataset of multi-view images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognit...
2023
-
[60]
Mip-splatting: Alias-free 3d gaussian splat- ting
Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. Mip-splatting: Alias-free 3d gaussian splat- ting. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 19447–19456,
-
[61]
Pixel-gs: Density control with pixel- aware gradient for 3d gaussian splatting
Zheng Zhang, Wenbo Hu, Yixing Lao, Tong He, and Hengshuang Zhao. Pixel-gs: Density control with pixel- aware gradient for 3d gaussian splatting. arXiv preprint arXiv:2403.15530, 2024. 2
2024 arXiv
-
[62]
Fsgs: Real-time few-shot view synthesis using gaussian splatting
Zehao Zhu, Zhiwen Fan, Yifan Jiang, and Zhangyang Wang. Fsgs: Real-time few-shot view synthesis using gaussian splatting. In European Conference on Computer Vision , pages 145–163. Springer, 2025. 1, 2, 5, 7
2025
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.