REVIEW 4 major objections 5 minor 59 references
MMOne: Representing Multiple Modalities in One Scene
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read MMOne claims that modality conflicts inside a shared Gaussian scene—property and granularity differences—can be resolved with a per-modality indicator plus a gradient-difference decomposition, yielding better quality for every modality…
desk verdict MMOne is a genuinely new 3DGS-based multimodal representation with impressive compactness but a central decomposition mechanism that is underspecified and not cleanly isolated in the ablations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the multimodal decomposition step embedded in 3D Gaussian Splatting's densification loop. Each Gaussian carries modality-specific features and a modality indicator; during training, the framework accumulates per-modality gradients and compares them through the L2 norm of their difference. When that norm exceeds a fixed threshold (0.0002), the Gaussian is decomposed into single-modal Gaussians, each driven by its own modality loss. A companion 'Soft Prune' rule turns off a single modality's indicator instead of deleting the whole Gaussian, and raises the pruning threshold for single-modal Gaussians to keep the scene compact. The modality indicator thus does double duty: it weights each modality's contribution in alpha-blending and acts as a switch that freezes updates for deactivated modalities.
What would settle it
Run MMOne with the multimodal decomposition step disabled, keeping the modality indicator and soft prune: if RGB, thermal, and language quality stay within noise and the Gaussian count does not grow, the decomposition is not what carries the reported gains. Directly measuring the distribution of per-Gaussian gradient differences during training would further show whether the fixed 0.0002 threshold actually separates conflicting Gaussians from non-conflicting ones.
Extended reading notes
Core claim
The central discovery is that modality conflicts in a unified Gaussian scene can be resolved by explicitly separating shared structure from modality-specific structure. MMOne models each modality with its own feature vector plus a modality indicator $\alpha^m \in [0,1]$ that acts as a per-modality opacity and as a learnable switch; because the switch can be turned off during rasterization, gradients from one modality no longer corrupt the geometry that another modality needs. The second mechanism watches the L2 norm of the gradient difference $g^d_{ij} = \mathrm{norm}(g^{m_i} - g^{m_j})$ accumulated on each Gaussian during densification. When that difference exceeds a threshold, the mechanism decomposes the multi-modal Gaussian into single-modal copies, each optimized by its own modality loss. The experiments claim that this yields consistently better rendering and semantics for RGB, thermal, and language, and that the decomposed representation is more compact than joint training on shared Gaussians.
Load-bearing premise
The framework assumes that the L2 norm of the difference between two modalities' gradients on a Gaussian, compared against a fixed threshold of 0.0002, reliably tells when that Gaussian should be split into modality-specific copies; if that heuristic does not track the true granularity needs of each modality, the quality and compactness claims lose their foundation.
Editorial extensions
If this is right
- Adding a fourth modality, monocular depth, improves depth rendering on flat surfaces and leaves RGB, thermal, and language quality unchanged, supporting the scalability claim.
- Because the modality indicator acts as a switch, the trained scene can render or suppress individual modalities at inference time without retraining.
- The scene is more compact as well as better: the method uses about one-third of the Gaussians of joint-training baselines on RGB-thermal scenes and about one-quarter of LangSplat's count on RGB-language scenes.
- In the three-modality experiments, adding language degrades RGB and thermal by about 0.5 dB for shared-Gaussian baselines, while MMOne slightly improves them, showing the decomposition removes interference that grows with modality count.
Reading between the lines
- The gradient-difference split rule resembles gradient-conflict resolution in multi-task learning, so an adaptive per-region threshold might replace the fixed 0.0002 value; the paper does not test that variant.
- The per-modality switch suggests a cheap route to editable or privacy-filtered representations, since thermal or language information could be suppressed at inference time, an application the paper describes in passing but does not explore.
- Because the compactness gain comes partly from pruning single-modal Gaussians more aggressively, an ablation that varies only the prune threshold would separate the contributions of decomposition and pruning.
- A natural stress test of the scalability claim would be to add a modality with a very different dimensionality or physical unit, such as audio or tactile signals; the paper only demonstrates a fourth modality that is geometrically aligned with the others (monocular depth).
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MMOne proposes a general 3D Gaussian Splatting framework that represents multiple modalities (RGB, thermal, language, and optionally depth) in a single scene. The paper identifies property disparity and granularity disparity as two modality-conflict challenges. The method contributes a modality modeling module with modality-specific features and a per-modality indicator opacity, and a multimodal decomposition mechanism that splits multi-modal Gaussians into modality-specific Gaussians based on the L2 norm of per-modality gradient differences (Eq. 3). The evaluation covers RGB-thermal (RGBT-Scenes), RGB-language (LERF/LangSplat), and RGB-thermal-language (four RGBT-Scenes scenes with manually annotated masks), plus ablations. The paper claims consistent per-modality improvements and compactness: roughly one-third to one-quarter the Gaussian count of baselines while improving or matching average PSNR, SSIM, and mIoU.
Significance. If the central claim is validated, MMOne is a step toward a unified multimodal scene representation: the compactness results (Table 8 and Table 9) are substantial, the framework is designed to add modalities with modest code changes, the code is released, and the paper includes ablations and a threshold-sensitivity study in the supplementary material. These are real strengths. However, the evidence is currently mixed: the gradient-difference decomposition rule in Eq. (3) is underspecified, the ablations do not isolate the decision rule, per-scene LPIPS worsens on several scenes, and no error bars or multiple-seed results are provided. The 'consistently enhances' claim is therefore stronger than the presented numbers support.
major comments (4)
- [§4.3, Eq. (3)] The conflict metric gd = norm(g_mi - g_mj) is underspecified: the paper does not state whether the gradients are taken with respect to the shared Gaussian attributes (mean, covariance, opacity) or the modality-specific features, nor how they are accumulated across pixels or iterations. Gradient magnitudes depend on the modality loss weights (0.5/0.5/0.2 in the RGB-thermal-language experiments), learning rate, image resolution, and scene scale, so a fixed threshold of 0.0002 (Supplementary, Sec. 7) has no obvious calibration across datasets. If the gradients are with respect to modality-specific feature vectors of different dimensions, the subtraction in Eq. (3) is not even defined. Table 6 only sweeps the threshold over 0.0001-0.0004 on one aggregate setting; it does not show that the criterion tracks true granularity conflict rather than optimization dynamics. Because decomposition is the mechanism that grounds the compactness and granularity claims, this specification gap is load-bearing.
- [§5.4, Table 5] The ablation does not isolate the decomposition decision rule. The 'Decomp.' row adds the multimodal decomposition mechanism on top of the modality indicator and soft prune, so gains relative to 'Prune (S)' could come from the added per-modality capacity (separate Gaussians for each modality) or from the modified densification, rather than from the Eq. (3) gradient-difference signal. A control that replaces the gradient-difference criterion with a random split of the same fraction of Gaussians, or with splitting all multi-modal Gaussians, is needed; without such a control, the paper's central mechanistic claim is not supported.
- [§5.1, Table 1] The abstract's claim of 'consistently enhances the representation capability for each modality' is not supported by per-scene results. Against ThermalGaussian, MMOne has worse RGB LPIPS on Dim (0.203 vs 0.194), RB (0.235 vs 0.199), and Pt (0.291 vs 0.268), and worse thermal LPIPS on RB (0.213 vs 0.198) and LS (0.272 vs 0.248). The average PSNR gain is 0.5 dB for RGB and 0.4 dB for thermal; with no error bars or multiple seeds, it is unclear whether these average differences are significant relative to per-scene variance. Please either soften the consistency claim or add per-scene significance testing and variance estimates.
- [§5.3, Table 3] The three-modality evaluation rests entirely on a self-implemented baseline ('MM-J') and manually annotated ground-truth masks for language queries, as the paper acknowledges. Because no independent baseline or established dataset exists for RGB-thermal-language, the mIoU numbers in Table 3 are not directly comparable with those in Table 2, and the manual annotation (medium-level SAM masks) could bias semantic segmentation results. The paper should release the annotations, report annotation consistency, and hedge any comparative conclusion on this setting.
minor comments (5)
- [§3.1, Eq. (1) vs Eq. (2)] The notation for the per-Gaussian feature m_i is overloaded: in Eq. (1) m_i denotes the modality feature vector, whereas in Eq. (2) the modality indicator uses a superscript m (α_i^m), making the distinction between modality index and Gaussian index confusing.
- [Figure 3] The dashed/solid 'distributions' are not labeled with axes or the exact plotted quantity; please clarify the units and the definition of the difference shown (e.g., α_R - α_T).
- [Table 5] The column headed 'Num×10^4' appears without units in the table body; please state explicitly that the numbers are Gaussian counts in units of 10^4 and align this with the scene-wise counts in the supplementary tables.
- [Supplementary Sec. 11] The scalability experiment with monocular depth is performed on a single scene ('Dimsum') and is presented as 'additional'; the conclusion that adding depth does not compromise other modalities should be supported by more than one scene or explicitly labeled as a pilot result.
- [References] In the LangSurf reference, 'Dingewn Zhang' appears to be a typo; please verify the author names.
Circularity Check
No significant circularity: MMOne's claims are empirical performance claims validated against external baselines, and no fitted parameter is renamed as a prediction.
full rationale
The paper's central claims are experimental: MMOne consistently enhances per-modality rendering and segmentation quality relative to external baselines (3DGS, ThermalGaussian, LangSplat) on novel-view splits. There is no derivation chain in which a predicted quantity is defined in terms of the outcome it is said to predict. The multimodal decomposition rule in Eq. (3) is a design heuristic, and the threshold 0.0002 is a tuned hyperparameter examined in an ablation (Table 6), not a parameter fitted to data and then reported as a prediction. The modality indicator in Eq. (2) is a learned per-modality opacity used in rendering; it is not constructed from the benchmark metrics it later improves. The self-citations to the authors' prior work [10, 11] appear only in the related-work survey and are not load-bearing for the method's design or evaluation. Concerns that the gradient-difference criterion is scale-dependent, that hyperparameters may have been selected using the evaluation scenes, and that no ablation isolates the Eq. (3) signal alone are experimental-support and correctness concerns, not circularity. Accordingly, no self-definitional step, fitted-input-called-prediction step, or load-bearing self-citation chain is present.
Assumptions & free parameters
free parameters (4)
- decomposition threshold =
0.0002
- soft prune threshold =
0.5
- loss weights =
0.5 RGB, 0.5 thermal, 0.2 language (3-modality)
- thermal smoothness loss weight =
0.6
assumptions (5)
- domain assumption 3D Gaussian Splatting is an appropriate base representation for multimodal scenes.
- ad hoc to paper Gradient difference between per-modality rendering losses is a reliable measure of modality conflict.
- ad hoc to paper Modality-specific opacities can be learned to disentangle shared vs. specific geometry.
- domain assumption Language features distilled from SAM and CLIP provide a stable ground truth for supervision.
- domain assumption The RGB point cloud from COLMAP is a sufficient initialization for thermal and language modalities.
Cite this review
Pith. "Pith review of MMOne: Representing Multiple Modalities in One Scene." pith.science (2026). https://pith.science/paper/O7OJP5NK
@misc{pith2026250711129,
author = {Pith},
title = {Pith review of: MMOne: Representing Multiple Modalities in One Scene},
year = {2026},
howpublished = {\url{https://pith.science/paper/O7OJP5NK}},
note = {Machine review of arXiv:2507.11129}
}
read the original abstract
Humans perceive the world through multimodal cues to understand and interact with the environment. Learning a scene representation for multiple modalities enhances comprehension of the physical world. However, modality conflicts, arising from inherent distinctions among different modalities, present two critical challenges: property disparity and granularity disparity. To address these challenges, we propose a general framework, MMOne, to represent multiple modalities in one scene, which can be readily extended to additional modalities. Specifically, a modality modeling module with a novel modality indicator is proposed to capture the unique properties of each modality. Additionally, we design a multimodal decomposition mechanism to separate multi-modal Gaussians into single-modal Gaussians based on modality differences. We address the essential distinctions among modalities by disentangling multimodal information into shared and modality-specific components, resulting in a more compact and efficient multimodal scene representation. Extensive experiments demonstrate that our method consistently enhances the representation capability for each modality and is scalable to additional modalities. The code is available at https://github.com/Neal2020GitHub/MMOne.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields
Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields. In Proceedings of the International Conference on Computer Vision, 2021. 2
work page 2021
-
[2]
Emerg- ing Properties in Self-Supervised Vision Transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing Properties in Self-Supervised Vision Transformers. In Proceedings of the International Conference on Computer Vision, 2021. 2
work page 2021
-
[3]
PGSR: Planar-based Gaussian Splat- ting for Efficient and High-Fidelity Surface Reconstruction
Danpeng Chen, Hai Li, Weicai Ye, Yifan Wang, Weijian Xie, Shangjin Zhai, Nan Wang, Haomin Liu, Hujun Bao, and Guofeng Zhang. PGSR: Planar-based Gaussian Splat- ting for Efficient and High-Fidelity Surface Reconstruction. IEEE Transactions on Visualization and Computer Graph- ics, 2024. 2
work page 2024
-
[4]
A Survey on 3D Gaussian Splatting
Guikun Chen and Wenguan Wang. A Survey on 3D Gaussian Splatting. arXiv preprint arXiv:2401.03890, 2024. 2
arXiv 2024
-
[5]
Thermal3D-GS: Physics-induced 3D Gaussians for Thermal Infrared Novel- view Synthesis
Qian Chen, Shihao Shu, and Xiangzhi Bai. Thermal3D-GS: Physics-induced 3D Gaussians for Thermal Infrared Novel- view Synthesis. In Proceedings of the European Conference on Computer Vision, 2024. 2
work page 2024
-
[6]
Tactile-Augmented Radiance Fields
Yiming Dou, Fengyu Yang, Yi Liu, Antonio Loquercio, and Andrew Owens. Tactile-Augmented Radiance Fields. In Proceedings of the Conference on Computer Vision and Pat- tern Recognition, 2024. 2
work page 2024
-
[7]
Multimodal Sensors and ML-Based Data Fusion for Advanced Robots
Shengshun Duan, Qiongfeng Shi, and Jun Wu. Multimodal Sensors and ML-Based Data Fusion for Advanced Robots. Advanced Intelligent Systems, 4(12):2200213, 2022. 1
work page 2022
-
[8]
Privacy-Preserving Person Detection Using Low-Resolution Infrared Cameras
Thomas Dubail, Fidel Alejandro Guerrero Pe ˜na, Heitor Rapela Medeiros, Masih Aminbeidokhti, Eric Granger, and Marco Pedersoli. Privacy-Preserving Person Detection Using Low-Resolution Infrared Cameras. In Proceedings of the European Conference on Computer Vision, 2022. 2
work page 2022
Show all 59 references
-
[9]
Scene Perception in the Human Brain
Russell A Epstein and Chris I Baker. Scene Perception in the Human Brain. Annual Review of Vision Science , 5(1): 373–397, 2019. 1
2019
-
[10]
Mini-Splatting: Represent- ing Scenes with a Constrained Number of Gaussians
Guangchi Fang and Bing Wang. Mini-Splatting: Represent- ing Scenes with a Constrained Number of Gaussians. InPro- ceedings of the European Conference on Computer Vision ,
-
[11]
Mini-Splatting2: Building 360 Scenes within Minutes via Aggressive Gaussian Densi- fication
Guangchi Fang and Bing Wang. Mini-Splatting2: Building 360 Scenes within Minutes via Aggressive Gaussian Densi- fication. arXiv preprint arXiv:2411.12788, 2024. 2
2024
-
[12]
NeRF: Neural Radiance Field in 3D Vision: A Comprehensive Review
Kyle Gao, Yina Gao, Hongjie He, Dening Lu, Linlin Xu, and Jonathan Li. NeRF: Neural Radiance Field in 3D Vision: A Comprehensive Review. arXiv preprint arXiv:2210.00379 ,
-
[13]
Deep Learning for 3D Point Clouds: A Survey
Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu, Li Liu, and Mohammed Bennamoun. Deep Learning for 3D Point Clouds: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(12):4338–4364, 2020. 1
2020
-
[14]
Illuminating the Dark Spaces of Healthcare with Ambient Intelligence
Albert Haque, Arnold Milstein, and Li Fei-Fei. Illuminating the Dark Spaces of Healthcare with Ambient Intelligence. Nature, 585(7824):193–202, 2020. 2
2020
-
[15]
ThermoNeRF: Multimodal Neural Radiance Fields for Thermal Novel View Synthesis
Mariam Hassan, Florent Forest, Olga Fink, and Mal- colm Mielle. ThermoNeRF: Multimodal Neural Radiance Fields for Thermal Novel View Synthesis. arXiv preprint arXiv:2403.12154, 2024. 1, 2
2024 arXiv
-
[16]
2D Gaussian Splatting for Geometrically Ac- curate Radiance Fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian Splatting for Geometrically Ac- curate Radiance Fields . In SIGGRAPH, 2024. 2
2024
-
[17]
Rethinking Visual Scene Perception
Helene Intraub. Rethinking Visual Scene Perception. Wiley Interdisciplinary Reviews: Cognitive Science, 3(1):117–127,
-
[18]
Visible and Infrared Imaging Based Inspection of Power Installation
B Jalil, MA Pascali, GR Leone, M Martinelli, D Moroni, O Salvetti, and A Berton. Visible and Infrared Imaging Based Inspection of Power Installation. Pattern Recognition and Image Analysis, 29(1):35–41, 2019. 2
2019
-
[19]
3D Gaussian Splatting for Real-Time Radiance Field Rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics, 42(4):139–1, 2023. 1, 2, 3, 5
2023
-
[20]
LERF: Language Embed- ded Radiance Fields
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. LERF: Language Embed- ded Radiance Fields. In Proceedings of the International Conference on Computer Vision, 2023. 1, 2, 5
2023
-
[21]
Segment Any- thing
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment Any- thing. In Proceedings of the International Conference on Computer Vision, 2023. 5, 6
2023
-
[22]
LangSurf: Language-Embedded Surface Gaussians for 3D Scene Un- derstanding
Hao Li, Roy Qin, Zhengyu Zou, Diqi He, Bohan Li, Bingquan Dai, Dingewn Zhang, and Junwei Han. LangSurf: Language-Embedded Surface Gaussians for 3D Scene Un- derstanding. arXiv preprint arXiv:2412.17635, 2024. 2
2024
-
[23]
Weakly Supervised 3D Open- vocabulary Segmentation
Kunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu, Yingchen Yu, Abdulmotaleb El Saddik, Christian Theobalt, Eric Xing, and Shijian Lu. Weakly Supervised 3D Open- vocabulary Segmentation. Advances in Neural Information Processing Systems, 36:53433–53456, 2023. 1, 2
2023
-
[24]
EfficientGS: Stream- lining Gaussian Splatting for Large-Scale High-Resolution Scene Representation
Wenkai Liu, Tao Guan, Bin Zhu, Luoyuan Xu, Zikai Song, Dan Li, Yuesong Wang, and Wei Yang. EfficientGS: Stream- lining Gaussian Splatting for Large-Scale High-Resolution Scene Representation. IEEE MultiMedia, 2025. 2
2025
-
[25]
Thermal- Gaussian: Thermal 3D Gaussian Splatting
Rongfeng Lu, Hangyu Chen, Zunjie Zhu, Yuhang Qin, Ming Lu, Le Zhang, Chenggang Yan, and Anke Xue. Thermal- Gaussian: Thermal 3D Gaussian Splatting. InProceedings of the International Conference on Learning Representations ,
-
[26]
Taming 3DGS: High-Quality Radiance Fields with Limited Resources
Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl, Markus Steinberger, Francisco Vicente Carrasco, and Fer- nando De La Torre. Taming 3DGS: High-Quality Radiance Fields with Limited Resources. In SIGGRAPH Asia, 2024. 2
2024
-
[27]
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. Communications of the ACM , 65(1):99–106,
-
[28]
Recent Indus- trial Applications of Infrared Thermography: A Review
Roque Alfredo Osornio-Rios, Jose Alfonso Antonino-Daviu, and Rene de Jesus Romero-Troncoso. Recent Indus- trial Applications of Infrared Thermography: A Review. 9 IEEE Transactions on Industrial Informatics , 15(2):615– 625, 2018. 2
2018
-
[29]
Exploring Multi-modal Neural Scene Rep- resentations With Applications on Thermal Imaging
Mert ¨Ozer, Maximilian Weiherer, Martin Hundhausen, and Bernhard Egger. Exploring Multi-modal Neural Scene Rep- resentations With Applications on Thermal Imaging. In Pro- ceedings of the European Conference on Computer Vision ,
-
[30]
3D Vision-Language Gaussian Splatting
Qucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng, Anwesa Choudhuri, Terrence Chen, Chen Chen, and Ziyan Wu. 3D Vision-Language Gaussian Splatting. In Pro- ceedings of the International Conference on Learning Rep- resentations, 2025. 1, 2
2025
-
[31]
Cross- Spectral Neural Radiance Fields
Matteo Poggi, Pierluigi Zama Ramirez, Fabio Tosi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano. Cross- Spectral Neural Radiance Fields. In Proceedings of the In- ternational Conference on 3D Vision, 2022. 2
2022
-
[32]
Gaus- sianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians
Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain, and Matthias Nießner. Gaus- sianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians. In Proceedings of the Conference on Computer Vision and Pattern Recognition, 2024. 2
2024
-
[33]
3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting
Zhiyin Qian, Shaofei Wang, Marko Mihajlovic, Andreas Geiger, and Siyu Tang. 3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting. In Proceedings of the Conference on Computer Vision and Pattern Recognition,
-
[34]
LangSplat: 3D Language Gaussian Splat- ting
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. LangSplat: 3D Language Gaussian Splat- ting. In Proceedings of the Conference on Computer Vision and Pattern Recognition, 2024. 1, 2, 3, 5, 6, 7
2024
-
[35]
GLS: Geometry-aware 3D Language Gaussian Splatting
Jiaxiong Qiu, Liu Liu, Zhizhong Su, and Tianwei Lin. GLS: Geometry-aware 3D Language Gaussian Splatting. arXiv preprint arXiv:2411.18066, 2024. 2
2024 arXiv
-
[36]
GOI: Find 3D Gaussians of Interest with an Optimizable Open- vocabulary Semantic-space Hyperplane
Yansong Qu, Shaohui Dai, Xinyang Li, Jianghang Lin, Li- ujuan Cao, Shengchuan Zhang, and Rongrong Ji. GOI: Find 3D Gaussians of Interest with an Optimizable Open- vocabulary Semantic-space Hyperplane. In Proceedings of the International Conference on Multimedia, 2024. 2
2024
-
[37]
Learning Transferable Visual Models From Natural Language Super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models From Natural Language Super- vision . In Proceedings of the International Conference o...
2021
-
[38]
Graph-based Thermal-Inertial SLAM with Probabilistic Neural Networks
Muhamad Risqi U Saputra, Chris Xiaoxuan Lu, Pedro Porto B de Gusmao, Bing Wang, Andrew Markham, and Niki Trigoni. Graph-based Thermal-Inertial SLAM with Probabilistic Neural Networks. IEEE Transactions on Robotics, 38(3):1875–1893, 2021. 2
2021
-
[39]
Structure- From-Motion Revisited
Johannes L Schonberger and Jan-Michael Frahm. Structure- From-Motion Revisited. In Proceedings of the Conference on Computer Vision and Pattern Recognition, 2016. 5
2016
-
[40]
Language Embedded 3D Gaussians for Open- V ocabulary Scene Understanding
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao- Hua Guan. Language Embedded 3D Gaussians for Open- V ocabulary Scene Understanding. In Proceedings of the Conference on Computer Vision and Pattern Recognition ,
-
[41]
Splat-MOVER: Multi-Stage, Open-V ocabulary Robotic Manipulation via Editable Gaus- sian Splatting
Olaolu Shorinwa, Johnathan Tucker, Aliyah Smith, Aiden Swann, Timothy Chen, Roya Firoozi, Monroe David Kennedy, and Mac Schwager. Splat-MOVER: Multi-Stage, Open-V ocabulary Robotic Manipulation via Editable Gaus- sian Splatting. In Proceedings of the Conference on Robot Learni...
2024
-
[42]
Implicit Neural Representa- tions with Periodic Activation Functions
Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit Neural Representa- tions with Periodic Activation Functions. Advances in Neu- ral Information Processing Systems, 33:7462–7473, 2020. 1
2020
-
[43]
Touch- GS: Visual-Tactile Supervised 3D Gaussian Splatting
Aiden Swann, Matthew Strong, Won Kyung Do, Gadiel Sz- naier Camps, Mac Schwager, and Monroe Kennedy. Touch- GS: Visual-Tactile Supervised 3D Gaussian Splatting. In Proceedings of the International Conference on Intelligent Robots and Systems, 2024. 2
2024
-
[44]
Nerfstudio: A Modular Framework for Neural Radiance Field Develop- ment
Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al. Nerfstudio: A Modular Framework for Neural Radiance Field Develop- ment. In SIGGRAPH, 2023. 1, 2
2023
-
[45]
DN-Splatter: Depth and Normal Priors for Gaussian Splatting and Meshing
Matias Turkulainen, Xuqian Ren, Iaroslav Melekhov, Otto Seiskari, Esa Rahtu, and Juho Kannala. DN-Splatter: Depth and Normal Priors for Gaussian Splatting and Meshing. In Proceedings of the Winter Conference on Applications of Computer Vision, 2025. 2
2025
-
[46]
The Development of Multisensory Pro- cesses
Mark T Wallace. The Development of Multisensory Pro- cesses. Cognitive Processing, 5:69–83, 2004. 1
2004
-
[47]
4D Gaussian Splatting for Real-Time Dynamic Scene Ren- dering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4D Gaussian Splatting for Real-Time Dynamic Scene Ren- dering. In Proceedings of the Conference on Computer Vi- sion and Pattern Recognition, 2024. 2
2024
-
[48]
OpenGaussian: Towards Point- Level 3D Gaussian-based Open V ocabulary Understanding
Yanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu, Ya- hao Shi, Xinhua Cheng, Chen Zhao, Haocheng Feng, Errui Ding, Jingdong Wang, et al. OpenGaussian: Towards Point- Level 3D Gaussian-based Open V ocabulary Understanding. Advances in Neural Information Processing Systems , 37: 1...
2024
-
[49]
PhysGaussian: Physics- Integrated 3D Gaussians for Generative Dynamics
Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu Jiang. PhysGaussian: Physics- Integrated 3D Gaussians for Generative Dynamics. In Pro- ceedings of the Conference on Computer Vision and Pattern Recognition, 2024. 2
2024
-
[50]
Street Gaussians: Modeling Dynamic Ur- ban Scenes with Gaussian Splatting
Yunzhi Yan, Haotong Lin, Chenxu Zhou, Weijie Wang, Haiyang Sun, Kun Zhan, Xianpeng Lang, Xiaowei Zhou, and Sida Peng. Street Gaussians: Modeling Dynamic Ur- ban Scenes with Gaussian Splatting. In Proceedings of the European Conference on Computer Vision, 2024. 1, 2
2024
-
[51]
Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes
Zehao Yu, Torsten Sattler, and Andreas Geiger. Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes. ACM Transactions on Graphics, 43 (6):1–13, 2024. 2
2024
-
[52]
Mul- timodal Intelligence: Representation Learning, Information Fusion, and Applications
Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng. Mul- timodal Intelligence: Representation Learning, Information Fusion, and Applications. IEEE Journal of Selected Topics in Signal Processing, 14(3):478–493, 2020. 1
2020
-
[53]
Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Ze- hao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang 10 Wang, and Achuta Kadambi. Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields. In Proceedings of the Conference on Computer Vis...
2024
-
[54]
DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes
Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes. In Proceedings of the Con- ference on Computer Vision and Pattern Recognition, 2024. 1, 2 11 M...
2024
-
[55]
Soft Prune
Additional Implementation Details In this section, we introduce the implementation details of the RGB-Thermal, RGB-Language, and RGB-Thermal- Language experiments. We also introduce the hyperparam- eter settings in our proposed module. RGB-Thermal. We adhere to the experimenta...
-
[56]
MM-J”. “MM
Additional Ablation Studies To further investigate the sensitivity of the threshold setting of multimodal decomposition, we conduct an additional ab- lation study. As shown in Tab. 6, our chosen threshold con- sistently achieves superior performance across all metrics. Moreove...
-
[57]
T-GS” and “MM-J
Additional Qualitative Results To analyze the distributions of multimodal and single- modal Gaussians, we use the “Dimsum” scene as an exam- ple. As shown in Fig. 8, different modalities require varying number of Gaussians. We also present the qualitative results of modality c...
-
[58]
F-GS” and “LS-J
Additional Quantitative Results We present the number of Gaussians in the RGB-Thermal and RGB-Language experiments to further highlight the ef- fectiveness of our method in achieving a compact repre- sentation. As shown in Tab. 8, our method uses approxi- mately one-third of t...
-
[59]
T-GS” refers to ThermalGaussian and “MM-J
Additional Experiments on Scalability We further validate scalability by incorporating monocular depth in the “Dimsum” scene. The results in Fig. 9 and Tab. 12 show that the rendered depth quality is enhanced, particularly on flat surfaces, without compromising the per- forman...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.