REVIEW 5 major objections 5 minor 46 references
GradiSeg: Gradient-Guided Gaussian Segmentation with Enhanced 3D Boundary Precision
T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper's central claim is that the persistent blurriness at object boundaries in Gaussian-splatting segmentation comes from single Gaussians trying to represent two objects at once, and that splitting and realigning those Gaussians…
desk verdict A sensible incremental extension to Gaussian Grouping with two genuinely new modules, but the empirical payoff is under-supported by a three-scene, no-error-bar evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the per-Gaussian Identity Encoding vector and its training-time gradient. Identity Encoding is a 16-dimensional learnable vector rendered by alpha-blending like color, so each pixel's semantic code is a weighted average of the Gaussians that cover it; a 1x1 convolution turns the rendered vectors into per-class probabilities. The load-bearing mechanism is the accumulated gradient of this identity vector: boundary Gaussians accumulate high identity gradients because their single encoding is pulled in conflicting directions by the two objects on either side. IGD uses that gradient as a trigger to densify, that is, split and adjust the offending Gaussians, and LA-KNN uses the opposite of the position-gradient direction as a neighbor-selection cue, with a KL loss that pulls selected neighbors' identity encodings together.
What would settle it
A concrete check would be to record, for every boundary Gaussian, the direction of its position gradient and the ground-truth semantic labels of the Gaussians LA-KNN selects; if the selected neighbors frequently sit on the far side of a true object boundary, the direction cue fails. A companion calculation would train a scene with IGD disabled and count how many identity encodings on a known boundary end up as weighted mixtures of two object classes, which would show whether the splitting trigger actually resolves the optimization conflict.
Extended reading notes
Core claim
The central claim is that the persistent failure mode of 3DGS segmentation, ambiguous object edges, can be traced to a single Gaussian being forced to represent two objects at once, and that gradient signals already present during training reveal exactly which Gaussians sit on boundaries. GradiSeg's IGD module monitors the accumulated gradient of each Gaussian's identity encoding; when this gradient exceeds a threshold, the Gaussian is split into two sub-Gaussians placed on opposite sides of the boundary, and their positions and scales are adjusted along the boundary contour. The LA-KNN module then imposes a 3D consistency loss using neighbors selected along the direction opposite to each Gaussian's position gradient, so identity encodings propagate along a surface rather than across an edge. With both modules, the paper reports state-of-the-art multi-view segmentation on LERF-Mask and large open-vocabulary gains over Gaussian Grouping, and the ablation results attribute the gain specifically to the two modules.
Load-bearing premise
The load-bearing premise is that the opposite of a Gaussian's position-gradient direction reliably points toward the object interior and thus toward correct same-instance neighbors; the paper gives no direct evidence for that mapping, and if position gradients instead respond to color or geometry changes, LA-KNN's neighbor selection becomes arbitrary.
Editorial extensions
If this is right
- If the claims hold, the same 3DGS scene can be segmented more accurately at object boundaries without retraining reconstruction, since PSNR, SSIM, and LPIPS stay essentially unchanged (PSNR 27.05 versus 27.09 for Gaussian Grouping).
- The decoupled Identity Encoding representation means downstream edits, such as object removal and color or style swapping, can be performed by manipulating a single group of Gaussians, with sharper group boundaries than previous grouping methods.
- Boundary precision improves open-vocabulary segmentation by a large margin, up to 11.6% mIoU on a single scene, making text-prompt-driven selection in Gaussian scenes more reliable when the external prompt-to-mask matcher is correct.
- The ablation results imply that both modules matter: removing LA-KNN lowers figurines mIoU from 81.3 to 79.3, while removing IGD is more damaging on the ramen and teatime scenes, so the two mechanisms are complementary rather than redundant.
- Because the gains are reported without sacrificing reconstruction quality, the method is compatible with existing 3DGS rendering pipelines that need both editable segmentation and high-fidelity rendering.
Reading between the lines
- Beyond the paper's experiments, the identity-gradient signal itself could be reused as a boundary detector: scenes with thin or detailed objects see the largest reported gains, so monitoring per-Gaussian gradient magnitude after training might localize problematic regions without any new supervision.
- LA-KNN's direction rule is one specific choice among many possible directional neighbor-selection schemes; comparing it against neighbor selection based on local surface normals or covariance eigenvectors on the same benchmarks would reveal whether the improvement comes from direction awareness or merely from replacing global nearest neighbors with any local directional rule.
- The evaluation uses three indoor tabletop scenes; extending to outdoor or cluttered scenes with many small objects would test whether the 5–11 mIoU gains generalize, since boundary density is far higher there.
- Because IGD changes the Gaussian distribution after the initial densification phase, an untested risk is that its gradient threshold interacts with scene density; a sweep over that threshold would indicate whether the reported gains are robust or tuned to these three scenes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. GradiSeg proposes a 3D Gaussian Splatting-based semantic segmentation method that augments Gaussian Grouping with two modules: Identity Gradient Guided Densification (IGD), which splits boundary Gaussians whose identity-encoding gradients exceed a threshold, and Local Adaptive K-Nearest Neighbors (LA-KNN), which propagates identity encodings among neighbors selected along the direction opposite to each Gaussian's position gradient. The paper claims significant mIoU and mBIoU improvements over state-of-the-art methods on the three-scene LERF-Mask dataset, with maintained reconstruction quality on Mip-NeRF 360, and demonstrates editing applications.
Significance. If the reported gains are robust, GradiSeg offers a practical and conceptually simple remedy to boundary blur in 3DGS segmentation, a known weakness of methods such as Gaussian Grouping. The paper's core idea of using gradient signals to drive densification and neighbor selection is plausible and the qualitative results are visually compelling. The framework is described in sufficient detail to be reimplemented, and the ablation structure (IGD and LA-KNN) is appropriate. However, the empirical support is currently too thin to establish the headline claim: the quantitative evaluation rests on three scenes, no variance estimates, and a schedule hyperparameter selected on the same evaluation set, while several numerical and textual inconsistencies prevent the reader from fully trusting the reported numbers. The contribution is potentially valuable, but the evidence as presented does not yet meet the bar for a strong empirical claim of state-of-the-art performance.
major comments (5)
- [Table 1, Sec 4.2] The central claim of a 5.27% average mIoU improvement over Gaussian Grouping is based on only three scenes, with no standard deviations, multiple seeds, or significance tests. The per-scene gains are 11.6, 1.5, and 2.7 points for figurines, ramen, and teatime respectively, so the average is dominated by one scene. Given the known run-to-run variability in 3DGS training, the reported difference cannot currently be distinguished from noise. Please report variance over at least several seeds and, ideally, additional scenes from the LERF-Mask or related benchmarks.
- [Sec 4.3, Fig 6, Sec 4.1] The IGD starting iteration (12,000) is selected using the same three scenes on which all results are reported. Figure 6 shows a clear peak near 12k, meaning the reported numbers incorporate tuning of a schedule hyperparameter on the evaluation set. This is a form of evaluation-set overfitting. A held-out validation scene, a sensitivity analysis across a range of start iterations, or a clear statement that this is a fixed schedule without tuning would be needed to make the results trustworthy.
- [Table 2, Sec 4.1, Fig 4, Supp Sec 6] There are several internal inconsistencies that must be resolved. (1) In Table 2, the OmniSeg3D average is reported as 79.4, but the per-scene values are 69.7, 77.0, and 71.7, whose average is 72.8; either the numbers or the label are wrong. (2) Sec 4.1 says LA-KNN is applied between 12,000 and 30,000 iterations, while the Supplementary Material (Sec 6) says LA-KNN is applied from 15,000 to 30,000. (3) Figure 4 states K=2 in the LA-KNN description, whereas Sec 4.1 and the Supplementary Material set K=5. These discrepancies undermine confidence in the experimental setup and must be corrected.
- [Sec 3.5, Eq (4)] The LA-KNN module relies on the assumption that the opposite of a Gaussian's position-gradient direction points toward correct same-instance neighbors. This premise is load-bearing: if position gradients respond to appearance or geometry changes rather than semantic boundaries, the selected neighbors may lie across object boundaries, and the KL loss in Eq. (3) could align identity encodings across distinct objects. The paper provides no direct evidence for this link. Please add an analysis or a targeted ablation that validates the directional choice, for example by comparing against random-direction neighbor selection or global-KNN under the same loss.
- [Algorithm 1, Sec 3.4] The threshold tau for identity-gradient-based splitting is never specified or analyzed in the paper. Algorithm 1 uses the condition gradienti > tau, but the value of tau and its sensitivity are not reported. Since this threshold controls which Gaussians are split, it is a key hyperparameter; please give its value and, if possible, a sensitivity study or a fixed heuristic (e.g., based on gradient quantiles).
minor comments (5)
- [Abstract and Sec 4.2] The phrase 'comprehensive experiments' is used to describe an evaluation on three scenes; a more modest wording, e.g., 'evaluations on the LERF-Mask benchmark', would be more accurate and avoid overclaiming.
- [Sec 4.2, paragraph after Eq. (5)] There is a typo in the loss definition: 'Ioutput is the is the rendered RGB image' should read 'Iout is the rendered RGB image'.
- [Table 1 and Table 4] The mBIoU metric is never formally defined in the paper. Please include a definition or a citation for mean Boundary Intersection over Union so that the reader can interpret the reported numbers.
- [Sec 4.1] The implementation details state that the dataset is downsampled by a factor of 8; please clarify whether this applies to both LERF-Mask and Mip-NeRF 360 and how it affects the comparison with baselines that may use the original resolution.
- [Sec 9 (Limitation)] The limitation section correctly notes that open-vocabulary results are constrained by third-party models (DEVA, Grounding DINO). This limitation is relevant to the interpretation of Table 1 and could be mentioned earlier in the main text.
Circularity Check
No circularity found: GradiSeg's central claims are empirical benchmark comparisons, not derivations that reduce to their own inputs.
full rationale
GradiSeg's central claim is an empirical one: that its two proposed modules, IGD and LA-KNN, improve 3D semantic segmentation accuracy on LERF-Mask while preserving reconstruction quality on Mip-NeRF 360. The paper contains no fitted formula that is renamed as a prediction. The IGD module (Algorithm 1) uses accumulated Identity Encoding gradients as a heuristic signal to split and adjust Gaussians near boundaries; the LA-KNN module (Eq. 3 and Eq. 4) selects neighbors by position-gradient direction and aligns Identity Encodings via a KL loss. These are training-time design choices whose effect is measured afterward with mIoU and mBIoU computed against ground-truth masks. Neither equation encodes the reported metrics, and the ablation in Table 4 shows that removing either module changes performance, which is exactly the behavior expected of a non-circular intervention. There are no self-citations on which the argument depends; the cited methods (Gaussian Grouping, DEVA, SAM, Grounding DINO) are external baselines or tools. The limitations the paper states in Section 9, such as dependence on target tracking and third-party mask generation, weaken generalizability but do not make the derivation circular. Concerns about the three-scene evaluation, the absence of error bars, and hyperparameters tuned on the same benchmark are robustness and statistical-inference concerns, not circularity, and per the review rules they are noted but do not raise the circularity score.
Assumptions & free parameters
free parameters (6)
- IGD gradient threshold tau
- LA-KNN neighbor count K =
5 (Figure 4 caption says 2)
- Loss weights alpha and beta =
alpha=1, beta=2
- Identity Encoding dimension =
16
- IGD start iteration =
12,000
- Sampled Gaussians M for 3D loss =
1000
assumptions (5)
- standard math Differentiable alpha-blending rasterization of 3DGS (Eq 1, Eq 2) and the standard densification and pruning pipeline are valid.
- domain assumption DEVA provides reliable multi-view consistent segmentation masks as training supervision.
- domain assumption Gaussians near object boundaries exhibit unusually high accumulated Identity Encoding gradients.
- ad hoc to paper The opposite of the position-gradient direction indicates the direction toward correct same-instance neighbors.
- ad hoc to paper Splitting a boundary Gaussian along the boundary and adjusting its scale resolves the optimization conflict.
Cite this review
Pith. "Pith review of GradiSeg: Gradient-Guided Gaussian Segmentation with Enhanced 3D Boundary Precision." pith.science (2026). https://pith.science/paper/VHWUBHBX
@misc{pith2026241200392,
author = {Pith},
title = {Pith review of: GradiSeg: Gradient-Guided Gaussian Segmentation with Enhanced 3D Boundary Precision},
year = {2026},
howpublished = {\url{https://pith.science/paper/VHWUBHBX}},
note = {Machine review of arXiv:2412.00392}
}
read the original abstract
While 3D Gaussian Splatting enables high-quality real-time rendering, existing Gaussian-based frameworks for 3D semantic segmentation still face significant challenges in boundary recognition accuracy. To address this, we propose a novel 3DGS-based framework named GradiSeg, incorporating Identity Encoding to construct a deeper semantic understanding of scenes. Our approach introduces two key modules: Identity Gradient Guided Densification (IGD) and Local Adaptive K-Nearest Neighbors (LA-KNN). The IGD module supervises gradients of Identity Encoding to refine Gaussian distributions along object boundaries, aligning them closely with boundary contours. Meanwhile, the LA-KNN module employs position gradients to adaptively establish locality-aware propagation of Identity Encodings, preventing irregular Gaussian spreads near boundaries. We validate the effectiveness of our method through comprehensive experiments. Results show that GradiSeg effectively addresses boundary-related issues, significantly improving segmentation accuracy without compromising scene reconstruction quality. Furthermore, our method's robust segmentation capability and decoupled Identity Encoding representation make it highly suitable for various downstream scene editing tasks, including 3D object removal, swapping and so on.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Mip-nerf 360: Unbounded anti-aliased neural radiance fields
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5470–5479, 2022. 2, 6, 7, 1
work page 2022
-
[2]
Seg- ment anything in 3d with nerfs
Jiazhong Cen, Zanwei Zhou, Jiemin Fang, Wei Shen, Lingxi Xie, Dongsheng Jiang, Xiaopeng Zhang, Qi Tian, et al. Seg- ment anything in 3d with nerfs. Advances in Neural Infor- mation Processing Systems, 36:25971–25990, 2023. 2, 3, 6
work page 2023
-
[3]
Segment any 3d gaussians, 2024
Jiazhong Cen, Jiemin Fang, Chen Yang, Lingxi Xie, Xi- aopeng Zhang, Wei Shen, and Qi Tian. Segment any 3d gaussians, 2024. 2, 3
work page 2024
-
[4]
Tracking anything with de- coupled video segmentation, 2023
Ho Kei Cheng, Seoung Wug Oh, Brian Price, Alexander Schwing, and Joon-Young Lee. Tracking anything with de- coupled video segmentation, 2023. 4, 6, 1
work page 2023
-
[5]
Gaussianpro: 3d gaussian splatting with progressive propagation, 2024
Kai Cheng, Xiaoxiao Long, Kaizhi Yang, Yao Yao, Wei Yin, Yuexin Ma, Wenping Wang, and Xuejin Chen. Gaussianpro: 3d gaussian splatting with progressive propagation, 2024. 2
work page 2024
-
[6]
Click-gaussian: Interactive segmenta- tion to any 3d gaussians
Seokhun Choi, Hyeonseop Song, Jaechul Kim, Taehyeong Kim, and Hoseok Do. Click-gaussian: Interactive segmenta- tion to any 3d gaussians. arXiv preprint arXiv:2407.11793,
-
[7]
Plenoxels: Radiance fields without neural networks
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5501–5510, 2022. 2
2022
-
[8]
Ea- gles: Efficient accelerated 3d gaussians with lightweight en- codings, 2024
Sharath Girish, Kamal Gupta, and Abhinav Shrivastava. Ea- gles: Efficient accelerated 3d gaussians with lightweight en- codings, 2024. 2
work page 2024
Show all 46 references
-
[9]
Ges: Generalized exponential splatting for ef- ficient radiance field rendering, 2024
Abdullah Hamdi, Luke Melas-Kyriazi, Jinjie Mai, Guocheng Qian, Ruoshi Liu, Carl V ondrick, Bernard Ghanem, and An- drea Vedaldi. Ges: Generalized exponential splatting for ef- ficient radiance field rendering, 2024
2024
-
[10]
Low latency point cloud rendering with learned splatting, 2024
Yueyu Hu, Ran Gong, Qi Sun, and Yao Wang. Low latency point cloud rendering with learned splatting, 2024. 2
2024
-
[11]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,
-
[12]
Lerf: Language embedded radiance fields, 2023
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. Lerf: Language embedded radiance fields, 2023. 6
2023
-
[13]
Garfield: Group anything with radiance fields
Chung Min Kim, Mingxuan Wu, Justin Kerr, Ken Gold- berg, Matthew Tancik, and Angjoo Kanazawa. Garfield: Group anything with radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21530–21539, 2024. 3, 6
2024
-
[14]
Segment any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 2023. 1, 3
2023
-
[15]
Robo3d: Towards robust and reliable 3d perception against corruptions
Lingdong Kong, Youquan Liu, Xin Li, Runnan Chen, Wen- wei Zhang, Jiawei Ren, Liang Pan, Kai Chen, and Ziwei Liu. Robo3d: Towards robust and reliable 3d perception against corruptions. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , pages 1...
2023
-
[16]
Vastgaussian: Vast 3d gaus- sians for large scene reconstruction, 2024
Jiaqi Lin, Zhihao Li, Xiao Tang, Jianzhuang Liu, Shiyong Liu, Jiayue Liu, Yangdi Lu, Xiaofei Wu, Songcen Xu, You- liang Yan, and Wenming Yang. Vastgaussian: Vast 3d gaus- sians for large scene reconstruction, 2024. 2
2024
-
[17]
Rtgs: Enabling real- time gaussian splatting on mobile devices using efficiency- guided pruning and foveated rendering, 2024
Weikai Lin, Yu Feng, and Yuhao Zhu. Rtgs: Enabling real- time gaussian splatting on mobile devices using efficiency- guided pruning and foveated rendering, 2024. 2
2024
-
[18]
Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle, 2023
Youtian Lin, Zuozhuo Dai, Siyu Zhu, and Yao Yao. Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle, 2023. 2
2023
-
[19]
Neural sparse voxel fields
Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. Advances in Neural Information Processing Systems, 33:15651–15663,
-
[20]
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023. 6, 1
2023 arXiv
-
[21]
Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians, 2024
Yang Liu, He Guan, Chuanchen Luo, Lue Fan, Naiyan Wang, Junran Peng, and Zhaoxiang Zhang. Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians, 2024. 2
2024
-
[22]
Scaffold-gs: Structured 3d gaussians for view-adaptive rendering, 2023
Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. Scaffold-gs: Structured 3d gaussians for view-adaptive rendering, 2023. 2
2023
-
[23]
Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713, 2023. 2
2023 arXiv
-
[24]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM , 65(1):99–106, 2021. 2 9
2021
-
[25]
Instant neural graphics primitives with a mul- tiresolution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a mul- tiresolution hash encoding. ACM transactions on graphics (TOG), 41(4):1–15, 2022. 2
2022
-
[26]
Langsplat: 3d language gaussian splatting,
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. Langsplat: 3d language gaussian splatting,
-
[27]
Semantics-controlled gaussian splatting for outdoor scene reconstruction and rendering in virtual reality, 2024
Hannah Schieber, Jacob Young, Tobias Langlotz, Stefanie Zollmann, and Daniel Roth. Semantics-controlled gaussian splatting for outdoor scene reconstruction and rendering in virtual reality, 2024. 1
2024
-
[28]
Panoocc: Unified occupancy representation for camera-based 3d panoptic segmentation
Yuqi Wang, Yuntao Chen, Xingyu Liao, Lue Fan, and Zhaox- iang Zhang. Panoocc: Unified occupancy representation for camera-based 3d panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 17158–17168, 2024. 1
2024
-
[29]
4d gaussian splatting for real-time dynamic scene rendering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20310–20320, 2024. 2
2024
-
[30]
Multi-scale 3d gaussian splatting for anti-aliased rendering,
Zhiwen Yan, Weng Fei Low, Yu Chen, and Gim Hee Lee. Multi-scale 3d gaussian splatting for anti-aliased rendering,
-
[31]
Tupper-map: Temporal and uni- fied panoptic perception for 3d metric-semantic mapping
Zhiliu Yang and Chen Liu. Tupper-map: Temporal and uni- fied panoptic perception for 3d metric-semantic mapping. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 1094–1101. IEEE, 2021. 2
2021
-
[32]
Unified perception and collabora- tive mapping for connected and autonomous vehicles
Zhiliu Yang and Chen Liu. Unified perception and collabora- tive mapping for connected and autonomous vehicles. IEEE Network, 37(4):273–281, 2023. 1
2023
-
[33]
Deformable 3d gaussians for high- fidelity monocular dynamic scene reconstruction
Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3d gaussians for high- fidelity monocular dynamic scene reconstruction. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20331–20341, 2024. 2
2024
-
[34]
Gaussian grouping: Segment and edit anything in 3d scenes
Mingqiao Ye, Martin Danelljan, Fisher Yu, and Lei Ke. Gaussian grouping: Segment and edit anything in 3d scenes. arXiv preprint arXiv:2312.00732, 2023. 2, 3, 4, 6, 7, 1
2023 arXiv
-
[35]
Omniseg3d: Omniversal 3d segmentation via hierarchical contrastive learning
Haiyang Ying, Yixuan Yin, Jinzhi Zhang, Fan Wang, Tao Yu, Ruqi Huang, and Lu Fang. Omniseg3d: Omniversal 3d segmentation via hierarchical contrastive learning. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20612–20622, 2024. 2, 3, 6
2024
-
[36]
Mip-splatting: Alias-free 3d gaussian splat- ting, 2023
Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. Mip-splatting: Alias-free 3d gaussian splat- ting, 2023. 2
2023
-
[37]
Segnerf: 3d part segmentation with neural radi- ance fields
Jesus Zarzar, Sara Rojas, Silvio Giancola, and Bernard Ghanem. Segnerf: 3d part segmentation with neural radi- ance fields. arXiv preprint arXiv:2211.11215, 2022. 2
2022 arXiv
-
[38]
Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024
Hanyue Zhang, Zhiliu Yang, Xinhe Zuo, Yuxin Tong, Ying Long, and Chen Liu. Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024. 2
2024
-
[39]
Nerf++: Analyzing and improving neural radiance fields
Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020. 2
2010 arXiv
-
[40]
In-place scene labelling and understanding with implicit scene representation
Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and An- drew J Davison. In-place scene labelling and understanding with implicit scene representation. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 15838–15847, 2021. 2, 3
2021
-
[41]
Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields
Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Ze- hao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi. Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields. InPro- ceedings of the IEEE/CVF Conference on Compu...
2024
-
[42]
However, the 2D semantic masks generated by SAM from different viewpoints often have inconsistent mask IDs for the same object
More Implementation Details More Implementation Details Following [34], given a series of RGB images with associated poses, we lever- age SAM [14] to produce the corresponding segmenta- tion masks, capitalizing on its outstanding segmentation capabilities. However, the 2D sema...
-
[44]
Additionally, the Fig- ure 9 shows more visualization comparison results on the Mip-NeRF 360 dataset [1]
More Results The Figure 8 shows more visualization comparison re- sults on the LERF-Mask dataset. Additionally, the Fig- ure 9 shows more visualization comparison results on the Mip-NeRF 360 dataset [1]. Compared to Gaussian Group- ing [34], our method enhances semantic segmen...
-
[45]
To mitigate this issue, we propose splitting such Gaussians into two smaller sub-Gaussians, strategi- cally distributing them on either side of the object bound- ary
Analyses and Discussions Q1: Why is a splitting operation performed in the IGD module? A1: For Gaussians exhibiting anomalous Identity Encoding gradients, the majority are concentrated near object bound- aries, which often results in rendering inaccuracies for iden- tity encod...
-
[46]
Limitation Our method leverages a target tracking mechanism to pre- generate multi-view consistent segmentation masks before model training. However, unsuccessful target tracking may result in the incorrect assignment of segmentation mask IDs, which hinders the model’s ability...
-
[256]
Softmax is applied after the convolution output to calculate class probabilities
This design aligns with the semantic mask, where ID values are mapped to pixel values ranging from 0 to 255. Softmax is applied after the convolution output to calculate class probabilities. The entire training process is conducted using the Adam optimizer on an A100 GPU. The ...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.