REVIEW 3 major objections 6 minor 2 cited by
HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Hybrid 2D and 3D Gaussians separate transient objects from static scenes in casually captured photos.
desk verdict A clean empirical contribution for removing moving distractors in 3DGS, with a clearly stated but untested assumption about single-view transients. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the hybrid Gaussian representation: per-image 2D Gaussians, whose centroid, scale, rotation, color, and opacity live in image space, and 3D Gaussians, whose parameters live in world space and are initialized from COLMAP point clouds. The 2D Gaussians render a transient image and a mask; the 3D Gaussians render the static image; the final render is a mask-weighted blend. To keep the 3D branch static-only, the paper supervises it only on the intersection of the frustums of K randomly sampled views, and restricts optimization to Gaussians whose centers fall in that co-visible region. The training alternates between freezing one branch and refining the other, then joint fine-tuning, so the 2D Gaussians learn residuals that the 3D branch cannot explain.
What would settle it
Render a scene where the same moving distractor (for example, a person or vehicle) is visible from multiple overlapping viewpoints at the same time and position: if the method renders it inconsistently across those views or folds it into the static 3D model, the single-view planar assumption is violated. A synthetic scene with a known transient object present in all K co-visible sampled views should cause the multi-view supervision to treat it as static, which can be directly checked against a ground-truth segmentation.
Extended reading notes
Core claim
The central discovery is that the static/transient decomposition can be achieved without semantic segmentation or per-image uncertainty networks by choosing the representation to match the geometric consistency of each part. Transient content is rendered by a fixed number of 2D Gaussians per image, which act as a per-view planar layer that also produces a soft transient mask via accumulated opacity. The static scene is rendered by 3D Gaussians supervised only in co-visible frustum regions across K sampled views, which prevents transient pixels from being baked into the 3D model. The two layers are combined with alpha blending and trained in three stages: warm-up, alternating refinement, and joint fine-tuning. The paper reports state-of-the-art results on both benchmark datasets and shows that the learned masks capture not just pedestrians and vehicles but also shadows and motion blur.
Load-bearing premise
Transient objects can always be treated as flat single-view objects because they never appear consistently across multiple views; if a transient is visible from several angles at the same location, the per-image 2D Gaussian representation cannot model it and the static/transient split would blur.
Editorial extensions
If this is right
- Clean static novel views can be rendered from casually captured phone photos without semantic labels or pretrained feature networks.
- Explicit transient masks are produced for free from the 2D Gaussian opacity, useful for editing or filtering.
- The method needs fewer 3D Gaussians because transients are not baked in, cutting storage and enabling faster training and rendering than plain 3DGS.
- The decomposition naturally absorbs non-semantic per-image effects like shadows and motion blur, not just recognized objects.
- The multi-view supervision stabilizes training and reduces overfitting to training views.
Reading between the lines
- The same 2D/3D split could be extended to model appearance and illumination changes in unconstrained web photo collections, since the paper notes that 2D Gaussians also capture photometric differences between images.
- The per-image planar assumption implies a testable limit: a scene where the same transient object is visible from multiple viewpoints at the same time (for example, a stationary car photographed from several angles) will force the 3D branch to absorb it, blurring the decomposition.
- A video variant could reuse 2D Gaussians across nearby frames with a small motion prior instead of allocating a full set per frame, reducing per-image storage and training cost.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes HybridGS, a hybrid scene representation for novel-view synthesis of scenes containing transient objects. The method uses per-image 2D Gaussians to model transient content and multi-view-consistent 3D Gaussians to model the static scene, together with a multi-view regulated supervision scheme (Algorithm 1) and a three-stage training procedure (warm-up, iterative training, joint fine-tuning). Experiments on NeRF On-the-go and RobustNeRF report consistent improvements over RobustNeRF, NeRF On-the-go, 3DGS, and SLS-mlp, with average gains of 1.10 dB PSNR and 5.27% SSIM on NeRF On-the-go, plus ablations and efficiency measurements.
Significance. If the reported results are reliable, HybridGS is a practically attractive contribution: it removes the need for semantic features, decouples transients and statics explicitly, and is faster and more storage-efficient than vanilla 3DGS on the reported Corner scene. The benchmark coverage is broad, with two established datasets, qualitative decomposition results, and ablations over components and 2D Gaussian counts. The main limitation of the claimed significance is that the paper frames the decomposition as fundamental ('transients are planar because they lack multi-view consistency'), yet the evaluation only concerns moving distractors and does not probe multi-view-consistent transients. The empirical case is also weakened by the absence of variance statistics and code. With those gaps addressed, the SOTA claim would be credible.
major comments (3)
- [Sec. 4.2, Sec. 4.3, Algorithm 1] The central premise that transients can be modeled as planar objects because they 'lack multi-view consistency and usually only appear in a single view' is load-bearing but untested for multi-view-consistent transients. A stationary or slowly moving distractor (e.g., a parked car visible in a contiguous block of frames) is multi-view consistent during those frames and will be captured by the 3DGS branch in the warm-up stage. Afterwards, Eq. (11) masks the 3D loss with (1-M_t), while M_t is learned from the residuals of a 3D model that already explains the distractor, so the 2D branch has no gradient signal to claim the region. Algorithm 1 further reinforces such a distractor when it falls in the intersection of the K sampled frustums. The paper should either restrict the claim to moving transients or provide an experiment that measures the method's behavior on stationary-but-transient objects and reports the fraction of transient pixels inside the co-visible frustum during training.
- [Sec. 5.1 and Figures 8/10] The training schedule is stated inconsistently. Section 5.1 says the process uses a 1k-step warm-up, then iterative training with 10k 2DGS steps and 1k 3DGS steps, and a 30k-step joint fine-tuning. Figures 8 and 10 instead state 'Warm-up: 0∼1,010, Iterative Training: 1,010∼40,400, Joint Training: 40,400∼60,600', which implies roughly 1k/39.4k/20.2k steps. These two accounts cannot both be right, and the discrepancy affects both the efficiency claim (0.18 GPU hours) and the convergence analysis in Figure 10. Please specify the exact number of iterations for each stage and for each branch, and align the figure captions.
- [Tables 1-4] The empirical evidence consists of single runs with no variance estimate or repeated-seed statistics, and no code is released. Given that several reported gaps over SLS-mlp are small (e.g., 0.10 dB on Android in Table 2), the 'sets a new standard' claim is stronger than the evidence supports. Please report mean/std over at least three runs or, at minimum, state that the numbers are from a single run and make the code available to allow reproduction.
minor comments (6)
- [Sec. 4, first paragraph] The sentence 'we introduce our hybrid representation of scenes using 3D Gaussians in Sec. 4.1 and 2D Gaussians in Sec. 8.1' refers to a supplementary section; the 2D Gaussian description belongs in the main paper or the reference should be to a main-paper section.
- [Sec. 5.3] The word 'Guassians' should be 'Gaussians' in the sentence '10k 2D Guassians achieve...'.
- [Eq. (6)] The 2D Gaussian rendering is written as a plain sum of alpha-weighted colors without transmittance terms; please clarify whether this is intentional and describe any normalization or saturation handling, since it differs from the alpha-blending formula in Eq. (4).
- [Algorithm 1] The update step is written as a generic parameter update without specifying the loss; state explicitly that the loss is Eq. (9) during warm-up and Eq. (11) during iterative training, and when the binarization threshold ε=0.1 is applied.
- [Table 2] The caption explains that Crab (1) and Crab (2) differ in the test set, but the main text never defines them; please move that explanation into Section 5.2.2.
- [Sec. 5.4 and Figure 11] The Photo Tourism results are only qualitative; a sentence clarifying that this dataset is not evaluated quantitatively would help readers interpret Figure 11.
Circularity Check
No significant circularity: the hybrid-Gaussian representation is an empirical design evaluated on held-out views, and the planar-transient assumption is an inductive bias rather than a self-fulfilling prediction.
full rationale
The paper's central claim is empirical: a hybrid representation (per-image 2D Gaussians for transients, multi-view 3D Gaussians for statics) with multi-view frustum supervision and a multi-stage training schedule is evaluated on held-out test views of NeRF On-the-go and RobustNeRF against external baselines. The decomposition into transients and statics (Eqs. 7-8, Sec. 4.2) is a modeling choice, not a first-principles prediction; the transient mask is a trainable output fitted through reconstruction losses (Eqs. 10-12), and the qualitative mask comparisons are illustrative rather than benchmarked predictions. No equation in the paper reduces a reported test metric to a fitted parameter by construction. The planar-transient assumption (Secs. 1 and 4.2) and the co-visible-frustum heuristic (Algorithm 1) are stated inductive biases whose failure modes for multi-view-consistent distractors are limitations, not circularity; the paper also explicitly acknowledges the illumination limitation in Sec. 5.4. The VastGaussian citation [17] appears only in a related-work list, is not load-bearing, and its first author (Jiaqi Lin) is not among the present paper's authors, so it is not a self-citation that affects the argument. I therefore find no significant circularity.
Assumptions & free parameters
free parameters (5)
- loss weight lambda and beta =
0.2
- multi-view batch size K =
4
- mask binarization threshold epsilon =
0.1
- number of 2D Gaussians per image =
10,000
- training step counts =
1k warm-up, 10k+1k iterative, 30k joint (text) versus 1,010/40,400/60,600 (figures)
assumptions (4)
- domain assumption Transient objects lack multi-view consistency and can be represented as planar objects per image.
- domain assumption Co-visible regions across K randomly sampled views are static and consistent.
- domain assumption COLMAP-derived initial point clouds represent mostly static structure.
- ad hoc to paper Binarizing the soft mask with threshold epsilon=0.1 improves training stability.
invented entities (1)
-
Per-image 2D Gaussian field for transients
Cite this review
Pith. "Pith review of HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting." pith.science (2026). https://pith.science/paper/SOW72K3G
@misc{pith2026241203844,
author = {Pith},
title = {Pith review of: HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting},
year = {2026},
howpublished = {\url{https://pith.science/paper/SOW72K3G}},
note = {Machine review of arXiv:2412.03844}
}
read the original abstract
Generating high-quality novel view renderings of 3D Gaussian Splatting (3DGS) in scenes featuring transient objects is challenging. We propose a novel hybrid representation, termed as HybridGS, using 2D Gaussians for transient objects per image and maintaining traditional 3D Gaussians for the whole static scenes. Note that, the 3DGS itself is better suited for modeling static scenes that assume multi-view consistency, but the transient objects appear occasionally and do not adhere to the assumption, thus we model them as planar objects from a single view, represented with 2D Gaussians. Our novel representation decomposes the scene from the perspective of fundamental viewpoint consistency, making it more reasonable. Additionally, we present a novel multi-view regulated supervision method for 3DGS that leverages information from co-visible regions, further enhancing the distinctions between the transients and statics. Then, we propose a straightforward yet effective multi-stage training strategy to ensure robust training and high-quality view synthesis across various settings. Experiments on benchmark datasets show our state-of-the-art performance of novel view synthesis in both indoor and outdoor scenes, even in the presence of distracting elements.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 2 Pith papers
-
Rectifying Mask via Entropy for Distractor-Free 3DGS in Ambiguous Scenarios
RefineSplat removes ambiguous distractors from 3DGS via entropy-aware adaptive masking and density control, releasing an 18-scene Ambiguous wild dataset and reporting SOTA metrics on multiple wild benchmarks.
-
RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGS
RobustSplat improves transient-free 3D Gaussian Splatting by postponing densification to 10,000 iterations and bootstrapping mask supervision from low to high resolution.
Reference graph
Works this paper leans on
-
[1]
Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P
Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. ICCV, 2021. 2
work page 2021
-
[2]
Barron, Ben Mildenhall, Dor Verbin, Pratul P
Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. CVPR, 2022. 2, 1
work page 2022
-
[3]
Pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. Pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 19457–19467, 2023. 1
work page 2024
-
[4]
Nerf-hugs: Improved neural radiance fields in non-static scenes using heuristics-guided segmentation
Jiahao Chen, Yipeng Qin, Lingjie Liu, Jiangbo Lu, and Guanbin Li. Nerf-hugs: Improved neural radiance fields in non-static scenes using heuristics-guided segmentation. 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR) , pages 19436–19446, 2024. 2, 3, 1
work page 2024
-
[5]
Hallucinated neural radiance fields in the wild
Xingyu Chen, Qi Zhang, Xiaoyu Li, Yue Chen, Feng Ying, Xuan Wang, and Jue Wang. Hallucinated neural radiance fields in the wild. 2022 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 12933– 12942, 2021. 3, 1
work page 2022
-
[6]
Text-to-3d us- ing gaussian splatting
Zilong Chen, Feng Wang, and Huaping Liu. Text-to-3d us- ing gaussian splatting. 2024 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 21401– 21412, 2023. 1
work page 2024
-
[7]
Swag: Splatting in the wild images with appearance-conditioned gaussians
Hiba Dahmani. Swag: Splatting in the wild images with appearance-conditioned gaussians. In ECCV, 2024. 3, 1
work page 2024
-
[8]
K-planes: Ex- plicit radiance fields in space, time, and appearance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Ex- plicit radiance fields in space, time, and appearance. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12479–12488, 2023. 3
work page 2023
Show all 54 references
-
[9]
2d gaussian splatting for geometrically accu- rate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2d gaussian splatting for geometrically accu- rate radiance fields. In SIGGRAPH 2024 Conference Papers. Association for Computing Machinery, 2024. 2, 3
2024
-
[10]
Refinedfields: Radiance fields refinement for unconstrained scenes
Karim Kassab, Antoine Schnepf, Jean-Yves Franceschi, Laurent Caraffa, Jeremie Mary, and Val ´erie Gouet-Brunet. Refinedfields: Radiance fields refinement for unconstrained scenes. arXiv preprint arXiv:2312.00639, 2023. 3
2023 arXiv
-
[11]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuehler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (TOG), 42:1 – 14, 2023. 1, 2, 6, 3
2023
-
[12]
A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets
Bernhard Kerbl, Andreas Meuleman, Georgios Kopanas, Michael Wimmer, Alexandre Lanvin, and George Drettakis. A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets. ACM Transactions on Graph- ics, 43(4), 2024. 2
2024
-
[13]
Kulh´anek and Torsten Sattler
Jon ´avs. Kulh´anek and Torsten Sattler. Tetra-nerf: Represent- ing neural radiance fields using tetrahedra. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 18412–18423, 2023. 2
2023
-
[14]
Wildgaussians: 3d gaussian splatting in the wild
Jonas Kulhanek, Songyou Peng, Zuzana Kukelova, Marc Pollefeys, and Torsten Sattler. Wildgaussians: 3d gaussian splatting in the wild. 2024. 3, 5, 1
2024
-
[15]
Art3d: 3d gaussian splatting for text-guided artistic scenes generation, 2024
Pengzhi Li, Chengshuai Tang, Qinxuan Huang, and Zhi- heng Li. Art3d: 3d gaussian splatting for text-guided artistic scenes generation, 2024. 1
2024
-
[16]
Gs-ir: 3d gaussian splatting for inverse rendering
Zhihao Liang, Qi Zhang, Yingfa Feng, Ying Shan, and Kui Jia. Gs-ir: 3d gaussian splatting for inverse rendering. 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 21644–21653, 2023. 1
2024
-
[17]
Vastgaussian: Vast 3d gaussians for large scene reconstruction
Jiaqi Lin, Zhihao Li, Xiao Tang, Jianzhuang Liu, Shiy- ong Liu, Jiayue Liu, Yangdi Lu, Xiaofei Wu, Songcen Xu, Youliang Yan, and Wenming Yang. Vastgaussian: Vast 3d gaussians for large scene reconstruction. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...
2024
-
[18]
Scaffold-gs: Structured 3d gaussians for view-adaptive rendering
Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. Scaffold-gs: Structured 3d gaussians for view-adaptive rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20654–20664, 2024. 2
2024
-
[19]
Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Saj- jadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for un- constrained photo collections. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),...
2021
-
[20]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 2
2020
-
[21]
Instant neural graphics primitives with a multires- olution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a multires- olution hash encoding. ACM Trans. Graph. , 41(4):102:1– 102:15, 2022. 2
2022
-
[22]
Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...
2024
-
[23]
Srinivasan, Jonathan T
Konstantinos Rematas, Andrew Liu, Pratul P. Srinivasan, Jonathan T. Barron, Andrea Tagliasacchi, Tom Funkhouser, and Vittorio Ferrari. Urban radiance fields. CVPR, 2022. 2
2022
-
[24]
Nerf on-the-go: Exploiting uncertainty for distractor-free nerfs in the wild
Weining Ren, Zihan Zhu, Boyang Sun, Jiaqi Chen, Marc Pollefeys, and Songyou Peng. Nerf on-the-go: Exploiting uncertainty for distractor-free nerfs in the wild. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2, 3, 5, 6, 8, 1, 4 9
2024
-
[25]
L3dg: Latent 3d gaussian diffusion
Barbara Roessle, Norman M ¨uller, Lorenzo Porzi, Samuel Rota Bul `o, Peter Kontschieder, Angela Dai, and Matthias Nießner. L3dg: Latent 3d gaussian diffusion. In SIGGRAPH Asia 2024 Conference Papers, 2024. 1
2024
-
[26]
Fleet, and Andrea Tagliasacchi
Sara Sabour, Suhani V ora, Daniel Duckworth, Ivan Krasin, David J. Fleet, and Andrea Tagliasacchi. Robustnerf: Ignor- ing distractors with robust losses. 2023 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 20626–20636, 2023. 2, 3, 5, 6, 8, 1
2023
-
[27]
Fleet, and Andrea Tagliasacchi
Sara Sabour, Lily Goli, George Kopanas, Mark Matthews, Dmitry Lagun, Leonidas Guibas, Alec Jacobson, David J. Fleet, and Andrea Tagliasacchi. SpotLessSplats: Ignoring distractors in 3d gaussian splatting.arXiv:2406.20055, 2024. 2, 3, 6, 8, 1
2024 arXiv
-
[28]
Taming 3dgs: High-quality radiance fields with limited resources
Saswat Mallick and Rahul Goel, Bernhard Kerbl, Fran- cisco Vicente Carrasco, Markus Steinberger, and Fernando De La Torre. Taming 3dgs: High-quality radiance fields with limited resources. In SIGGRAPH Asia 2024 Conference Pa- pers, 2024. 2, 5
2024
-
[29]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 4
2016
-
[30]
Pixelwise view selection for un- structured multi-view stereo
Johannes Lutz Sch ¨onberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for un- structured multi-view stereo. In European Conference on Computer Vision (ECCV), 2016. 4
2016
-
[31]
Language embedded 3d gaussians for open- vocabulary scene understanding
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao- Hua Guan. Language embedded 3d gaussians for open- vocabulary scene understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5333–5343, 2024. 1
2024
-
[32]
Seitz, and Richard Szeliski
Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo tourism: Exploring photo collections in 3d. In SIGGRAPH Conference Proceedings , pages 835–846, New York, NY , USA, 2006. ACM Press. 8, 1, 4, 5
2006
-
[33]
Srinivasan, Jonathan T
Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Prad- han, Ben Mildenhall, Pratul P. Srinivasan, Jonathan T. Bar- ron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)...
2022
-
[34]
Nerfstudio: A modu- lar framework for neural radiance field development
Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Justin Kerr, Terrance Wang, Alexander Kristof- fersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, David McAllister, and Angjoo Kanazawa. Nerfstudio: A modu- lar framework for neural radiance field development. ACM SIG...
2023
-
[35]
Dc- gaussian: Improving 3d gaussian splatting for reflective dash cam videos
Linhan Wang, Kai Cheng, Shuo Lei, Shengkun Wang, Wei Yin, Chenyang Lei, Xiaoxiao Long, and Chang-Tien Lu. Dc- gaussian: Improving 3d gaussian splatting for reflective dash cam videos. NeurIPS 2024, 2024. 1
2024
-
[36]
Ie-nerf: Inpainting enhanced neural radiance fields in the wild
Shuaixian Wang, Haoran Xu, Yaokun Li, Jiwei Chen, and Guang Tan. Ie-nerf: Inpainting enhanced neural radiance fields in the wild. ArXiv, abs/2407.10695, 2024. 3
2024
-
[37]
We-gs: An in-the-wild efficient 3d gaussian representation for unconstrained photo collections, 2024
Yuze Wang, Junyi Wang, and Yue Qi. We-gs: An in-the-wild efficient 3d gaussian representation for unconstrained photo collections, 2024. 3, 1
2024
-
[38]
Sheikh, and Eero P
Zhou Wang, Alan Conrad Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Im- age Processing, 13:600–612, 2004. 5, 6
2004
-
[39]
4d gaussian splatting for real-time dynamic scene render- ing
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene render- ing. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 2...
2024
-
[40]
Dˆ2neRF: Self- supervised decoupling of dynamic and static objects from a monocular video
Tianhao Walter Wu, Fangcheng Zhong, Andrea Tagliasac- chi, Forrester Cole, and Cengiz Oztireli. Dˆ2neRF: Self- supervised decoupling of dynamic and static objects from a monocular video. In Advances in Neural Information Pro- cessing Systems, 2022. 2
2022
-
[41]
Splatfacto-w: A nerfstudio implementation of gaussian splatting for unconstrained photo collections
Congrong Xu, Justin Kerr, and Angjoo Kanazawa. Splatfacto-w: A nerfstudio implementation of gaussian splatting for unconstrained photo collections. ArXiv, abs/2407.12306, 2024. 3
2024 arXiv
-
[42]
Cross-ray neural radiance fields for novel- view synthesis from unconstrained image collections
Yifan Yang, Shuhai Zhang, Zixiong Huang, Yubing Zhang, and Mingkui Tan. Cross-ray neural radiance fields for novel- view synthesis from unconstrained image collections. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 15855–15865, 2023. 3
2023
-
[43]
gsplat: An open-source library for Gaussian splatting
Vickie Ye, Ruilong Li, Justin Kerr, Matias Turkulainen, Brent Yi, Zhuoyang Pan, Otto Seiskari, Jianbo Ye, Jeffrey Hu, Matthew Tancik, and Angjoo Kanazawa. gsplat: An open-source library for Gaussian splatting. arXiv preprint arXiv:2409.06765, 2024. 5
2024 arXiv
-
[44]
Mip-splatting: Alias-free 3d gaussian splat- ting
Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. Mip-splatting: Alias-free 3d gaussian splat- ting. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 19447– 19456, 2024. 2
2024
-
[45]
Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes
Zehao Yu, Torsten Sattler, and Andreas Geiger. Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes. ACM Transactions on Graphics, 2024. 2
2024
-
[46]
Gaussian in the wild: 3d gaussian splatting for unconstrained image collections
Dongbin Zhang, Chuming Wang, Weitao Wang, Peihao Li, Minghan Qin, and Haoqian Wang. Gaussian in the wild: 3d gaussian splatting for unconstrained image collections. ECCV, 2024. 3, 1
2024
-
[47]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. 2018 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , pages 586–595, 2018. 6
2018
-
[48]
Gaussianimage: 1000 fps image representation and compres- sion by 2d gaussian splatting
Xinjie Zhang, Xingtong Ge, Tongda Xu, Dailan He, Yan Wang, Hongwei Qin, Guo Lu, Jing Geng, and Jun Zhang. Gaussianimage: 1000 fps image representation and compres- sion by 2d gaussian splatting. In European Conference on Computer Vision, 2024. 3
2024
-
[49]
Image-gs: Content-adaptive image representation via 2d gaussians, 2024
Yunxiang Zhang, Alexandr Kuznetsov, Akshay Jindal, Ken- neth Chen, Anton Sochenov, Anton Kaplanyan, and Qi Sun. Image-gs: Content-adaptive image representation via 2d gaussians, 2024. 3
2024
-
[50]
Drivinggaussian: 10 Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes
Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. Drivinggaussian: 10 Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pa...
2024
-
[51]
We distinguish these approaches according to the primary datasets they uti- lize and provide a detailed comparison in Tab
Positioning of Our Work In the community, there are currently two predominant ap- proaches for tackling the challenge of novel view synthesis in wild images with different complexities. We distinguish these approaches according to the primary datasets they uti- lize and provid...
-
[52]
2D Gaussians The fitting capability of 2D Gaussians is inherited from 3D Gaussians
More Discussions 8.1. 2D Gaussians The fitting capability of 2D Gaussians is inherited from 3D Gaussians. Given J, the Jacobian of the affine projective transformation, and W, the viewing transformation, the 3D Gaussians can be projected to 2D image plane and blended through a...
-
[53]
Training For the training of 3D Gaussians, we perform the densifi- cation of 3D Gaussians during the warm-up stage
More Implementation Details 9.1. Training For the training of 3D Gaussians, we perform the densifi- cation of 3D Gaussians during the warm-up stage. Then, in the subsequent stages, we maintain a constant number of existing 3D Gaussians and focus solely on optimizing their para...
-
[54]
More Visualization Results 10.1. Training Process To better demonstrate the changes during our training pro- cess, we select IMG 7195.JPG of Corner from NeRF On- the-go dataset as example, visualizing the statics and tran- sients during different training stages and comparing ...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.