Pith. sign in

REVIEW 4 major objections 5 minor 59 references

MMOne: Representing Multiple Modalities in One Scene

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read MMOne claims that modality conflicts inside a shared Gaussian scene—property and granularity differences—can be resolved with a per-modality indicator plus a gradient-difference decomposition, yielding better quality for every modality…

desk verdict MMOne is a genuinely new 3DGS-based multimodal representation with impressive compactness but a central decomposition mechanism that is underspecified and not cleanly isolated in the ablations. read the letter →

arxiv 2507.11129 v2 pith:O7OJP5NK submitted 2025-07-15 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords 3DGaussianSplattingmultimodalscenerepresentationmodalityconflictgranularitydisparityindicatordecompositionRGB-thermal-languagerenderingopen-vocabularysegmentation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MMOne tackles a problem that arises when one 3D scene must represent several modalities at once: RGB, thermal, and language differ in physical properties and in the level of detail they need. The paper claims that the usual approach of attaching extra feature vectors to the same set of Gaussians lets these differences fight each other, degrading every modality. MMOne instead gives each modality its own feature vector and a per-modality opacity 'indicator' that can switch a modality off during rendering. It then adds a decomposition step in training that watches the per-modality gradients on each Gaussian and, when two modalities pull in different directions, replaces the shared Gaussian with separate per-modality Gaussians. On RGB-thermal, RGB-language, and RGB-thermal-language benchmarks, the method reports improved quality for every modality while using roughly one-third of the Gaussians of joint-training baselines.

What carries the argument

The load-bearing mechanism is the multimodal decomposition step embedded in 3D Gaussian Splatting's densification loop. Each Gaussian carries modality-specific features and a modality indicator; during training, the framework accumulates per-modality gradients and compares them through the L2 norm of their difference. When that norm exceeds a fixed threshold (0.0002), the Gaussian is decomposed into single-modal Gaussians, each driven by its own modality loss. A companion 'Soft Prune' rule turns off a single modality's indicator instead of deleting the whole Gaussian, and raises the pruning threshold for single-modal Gaussians to keep the scene compact. The modality indicator thus does double duty: it weights each modality's contribution in alpha-blending and acts as a switch that freezes updates for deactivated modalities.

What would settle it

Run MMOne with the multimodal decomposition step disabled, keeping the modality indicator and soft prune: if RGB, thermal, and language quality stay within noise and the Gaussian count does not grow, the decomposition is not what carries the reported gains. Directly measuring the distribution of per-Gaussian gradient differences during training would further show whether the fixed 0.0002 threshold actually separates conflicting Gaussians from non-conflicting ones.

Watch

Extended reading notes

Core claim

The central discovery is that modality conflicts in a unified Gaussian scene can be resolved by explicitly separating shared structure from modality-specific structure. MMOne models each modality with its own feature vector plus a modality indicator $\alpha^m \in [0,1]$ that acts as a per-modality opacity and as a learnable switch; because the switch can be turned off during rasterization, gradients from one modality no longer corrupt the geometry that another modality needs. The second mechanism watches the L2 norm of the gradient difference $g^d_{ij} = \mathrm{norm}(g^{m_i} - g^{m_j})$ accumulated on each Gaussian during densification. When that difference exceeds a threshold, the mechanism decomposes the multi-modal Gaussian into single-modal copies, each optimized by its own modality loss. The experiments claim that this yields consistently better rendering and semantics for RGB, thermal, and language, and that the decomposed representation is more compact than joint training on shared Gaussians.

Load-bearing premise

The framework assumes that the L2 norm of the difference between two modalities' gradients on a Gaussian, compared against a fixed threshold of 0.0002, reliably tells when that Gaussian should be split into modality-specific copies; if that heuristic does not track the true granularity needs of each modality, the quality and compactness claims lose their foundation.

Editorial extensions

If this is right

  • Adding a fourth modality, monocular depth, improves depth rendering on flat surfaces and leaves RGB, thermal, and language quality unchanged, supporting the scalability claim.
  • Because the modality indicator acts as a switch, the trained scene can render or suppress individual modalities at inference time without retraining.
  • The scene is more compact as well as better: the method uses about one-third of the Gaussians of joint-training baselines on RGB-thermal scenes and about one-quarter of LangSplat's count on RGB-language scenes.
  • In the three-modality experiments, adding language degrades RGB and thermal by about 0.5 dB for shared-Gaussian baselines, while MMOne slightly improves them, showing the decomposition removes interference that grows with modality count.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The gradient-difference split rule resembles gradient-conflict resolution in multi-task learning, so an adaptive per-region threshold might replace the fixed 0.0002 value; the paper does not test that variant.
  • The per-modality switch suggests a cheap route to editable or privacy-filtered representations, since thermal or language information could be suppressed at inference time, an application the paper describes in passing but does not explore.
  • Because the compactness gain comes partly from pruning single-modal Gaussians more aggressively, an ablation that varies only the prune threshold would separate the contributions of decomposition and pruning.
  • A natural stress test of the scalability claim would be to add a modality with a very different dimensionality or physical unit, such as audio or tactile signals; the paper only demonstrates a fourth modality that is geometrically aligned with the others (monocular depth).
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. MMOne proposes a general 3D Gaussian Splatting framework that represents multiple modalities (RGB, thermal, language, and optionally depth) in a single scene. The paper identifies property disparity and granularity disparity as two modality-conflict challenges. The method contributes a modality modeling module with modality-specific features and a per-modality indicator opacity, and a multimodal decomposition mechanism that splits multi-modal Gaussians into modality-specific Gaussians based on the L2 norm of per-modality gradient differences (Eq. 3). The evaluation covers RGB-thermal (RGBT-Scenes), RGB-language (LERF/LangSplat), and RGB-thermal-language (four RGBT-Scenes scenes with manually annotated masks), plus ablations. The paper claims consistent per-modality improvements and compactness: roughly one-third to one-quarter the Gaussian count of baselines while improving or matching average PSNR, SSIM, and mIoU.

Significance. If the central claim is validated, MMOne is a step toward a unified multimodal scene representation: the compactness results (Table 8 and Table 9) are substantial, the framework is designed to add modalities with modest code changes, the code is released, and the paper includes ablations and a threshold-sensitivity study in the supplementary material. These are real strengths. However, the evidence is currently mixed: the gradient-difference decomposition rule in Eq. (3) is underspecified, the ablations do not isolate the decision rule, per-scene LPIPS worsens on several scenes, and no error bars or multiple-seed results are provided. The 'consistently enhances' claim is therefore stronger than the presented numbers support.

major comments (4)
  1. [§4.3, Eq. (3)] The conflict metric gd = norm(g_mi - g_mj) is underspecified: the paper does not state whether the gradients are taken with respect to the shared Gaussian attributes (mean, covariance, opacity) or the modality-specific features, nor how they are accumulated across pixels or iterations. Gradient magnitudes depend on the modality loss weights (0.5/0.5/0.2 in the RGB-thermal-language experiments), learning rate, image resolution, and scene scale, so a fixed threshold of 0.0002 (Supplementary, Sec. 7) has no obvious calibration across datasets. If the gradients are with respect to modality-specific feature vectors of different dimensions, the subtraction in Eq. (3) is not even defined. Table 6 only sweeps the threshold over 0.0001-0.0004 on one aggregate setting; it does not show that the criterion tracks true granularity conflict rather than optimization dynamics. Because decomposition is the mechanism that grounds the compactness and granularity claims, this specification gap is load-bearing.
  2. [§5.4, Table 5] The ablation does not isolate the decomposition decision rule. The 'Decomp.' row adds the multimodal decomposition mechanism on top of the modality indicator and soft prune, so gains relative to 'Prune (S)' could come from the added per-modality capacity (separate Gaussians for each modality) or from the modified densification, rather than from the Eq. (3) gradient-difference signal. A control that replaces the gradient-difference criterion with a random split of the same fraction of Gaussians, or with splitting all multi-modal Gaussians, is needed; without such a control, the paper's central mechanistic claim is not supported.
  3. [§5.1, Table 1] The abstract's claim of 'consistently enhances the representation capability for each modality' is not supported by per-scene results. Against ThermalGaussian, MMOne has worse RGB LPIPS on Dim (0.203 vs 0.194), RB (0.235 vs 0.199), and Pt (0.291 vs 0.268), and worse thermal LPIPS on RB (0.213 vs 0.198) and LS (0.272 vs 0.248). The average PSNR gain is 0.5 dB for RGB and 0.4 dB for thermal; with no error bars or multiple seeds, it is unclear whether these average differences are significant relative to per-scene variance. Please either soften the consistency claim or add per-scene significance testing and variance estimates.
  4. [§5.3, Table 3] The three-modality evaluation rests entirely on a self-implemented baseline ('MM-J') and manually annotated ground-truth masks for language queries, as the paper acknowledges. Because no independent baseline or established dataset exists for RGB-thermal-language, the mIoU numbers in Table 3 are not directly comparable with those in Table 2, and the manual annotation (medium-level SAM masks) could bias semantic segmentation results. The paper should release the annotations, report annotation consistency, and hedge any comparative conclusion on this setting.
minor comments (5)
  1. [§3.1, Eq. (1) vs Eq. (2)] The notation for the per-Gaussian feature m_i is overloaded: in Eq. (1) m_i denotes the modality feature vector, whereas in Eq. (2) the modality indicator uses a superscript m (α_i^m), making the distinction between modality index and Gaussian index confusing.
  2. [Figure 3] The dashed/solid 'distributions' are not labeled with axes or the exact plotted quantity; please clarify the units and the definition of the difference shown (e.g., α_R - α_T).
  3. [Table 5] The column headed 'Num×10^4' appears without units in the table body; please state explicitly that the numbers are Gaussian counts in units of 10^4 and align this with the scene-wise counts in the supplementary tables.
  4. [Supplementary Sec. 11] The scalability experiment with monocular depth is performed on a single scene ('Dimsum') and is presented as 'additional'; the conclusion that adding depth does not compromise other modalities should be supported by more than one scene or explicitly labeled as a pilot result.
  5. [References] In the LangSurf reference, 'Dingewn Zhang' appears to be a typo; please verify the author names.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: MMOne's claims are empirical performance claims validated against external baselines, and no fitted parameter is renamed as a prediction.

full rationale

The paper's central claims are experimental: MMOne consistently enhances per-modality rendering and segmentation quality relative to external baselines (3DGS, ThermalGaussian, LangSplat) on novel-view splits. There is no derivation chain in which a predicted quantity is defined in terms of the outcome it is said to predict. The multimodal decomposition rule in Eq. (3) is a design heuristic, and the threshold 0.0002 is a tuned hyperparameter examined in an ablation (Table 6), not a parameter fitted to data and then reported as a prediction. The modality indicator in Eq. (2) is a learned per-modality opacity used in rendering; it is not constructed from the benchmark metrics it later improves. The self-citations to the authors' prior work [10, 11] appear only in the related-work survey and are not load-bearing for the method's design or evaluation. Concerns that the gradient-difference criterion is scale-dependent, that hyperparameters may have been selected using the evaluation scenes, and that no ablation isolates the Eq. (3) signal alone are experimental-support and correctness concerns, not circularity. Accordingly, no self-definitional step, fitted-input-called-prediction step, or load-bearing self-citation chain is present.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The method rests on standard 3DGS machinery plus several hyperparameters (decomposition threshold, soft-prune threshold, loss weights). The decomposition rule itself is an ad hoc heuristic. No new physical entities are introduced.

free parameters (4)
  • decomposition threshold = 0.0002
    Threshold for gradient difference between modalities to trigger Gaussian decomposition; selected via ablation on the evaluation scenes (Table 6).
  • soft prune threshold = 0.5
    Threshold for pruning single-modal Gaussians; set manually, controls compactness and performance (Sec. 7).
  • loss weights = 0.5 RGB, 0.5 thermal, 0.2 language (3-modality)
    Weights for modality losses; set manually and differ between RGB-Thermal and RGB-Thermal-Language experiments (Sec. 7).
  • thermal smoothness loss weight = 0.6
    Weight for the smoothness loss on thermal images, inherited from ThermalGaussian (Sec. 7).
assumptions (5)
  • domain assumption 3D Gaussian Splatting is an appropriate base representation for multimodal scenes.
    Section 3.1 assumes the 3DGS formulation and alpha-blending as the foundation for all modalities.
  • ad hoc to paper Gradient difference between per-modality rendering losses is a reliable measure of modality conflict.
    Eq. 3 proposes gd = norm(g_mi - g_mj) without theoretical justification; assumed to indicate when to split Gaussians.
  • ad hoc to paper Modality-specific opacities can be learned to disentangle shared vs. specific geometry.
    Section 4.2 asserts that a per-modality indicator 'switch' captures unique properties and granularities, without a formal model.
  • domain assumption Language features distilled from SAM and CLIP provide a stable ground truth for supervision.
    Section 5 describes obtaining ground-truth semantic features using SAM ViT-H and OpenCLIP ViT-B/16, assuming these features are reliable targets.
  • domain assumption The RGB point cloud from COLMAP is a sufficient initialization for thermal and language modalities.
    Section 5.1 states all methods use the same sparse point cloud and camera poses obtained from RGB images, implicitly assuming multi-modal alignment.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MMOne: Representing Multiple Modalities in One Scene." pith.science (2026). https://pith.science/paper/O7OJP5NK

@misc{pith2026250711129,
  author       = {Pith},
  title        = {Pith review of: MMOne: Representing Multiple Modalities in One Scene},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O7OJP5NK}},
  note         = {Machine review of arXiv:2507.11129}
}
read the original abstract

Humans perceive the world through multimodal cues to understand and interact with the environment. Learning a scene representation for multiple modalities enhances comprehension of the physical world. However, modality conflicts, arising from inherent distinctions among different modalities, present two critical challenges: property disparity and granularity disparity. To address these challenges, we propose a general framework, MMOne, to represent multiple modalities in one scene, which can be readily extended to additional modalities. Specifically, a modality modeling module with a novel modality indicator is proposed to capture the unique properties of each modality. Additionally, we design a multimodal decomposition mechanism to separate multi-modal Gaussians into single-modal Gaussians based on modality differences. We address the essential distinctions among modalities by disentangling multimodal information into shared and modality-specific components, resulting in a more compact and efficient multimodal scene representation. Extensive experiments demonstrate that our method consistently enhances the representation capability for each modality and is scalable to additional modalities. The code is available at https://github.com/Neal2020GitHub/MMOne.

Figures

Figures reproduced from arXiv: 2507.11129 by the authors.

Figure 1
Figure 1. MMOne Overview. MMOne is a general framework designed to represent multiple modalities in one scene. By disentangling multimodal information based on the inherent differences among modalities, we achieve enhanced performance across all modalities. Abstract Humans perceive the world through multimodal cues to understand and interact with the environment. Learn￾ing a scene representation for multiple modalities enhanc… view at source ↗
Figure 2
Figure 2. The General Framework of MMOne. Given multi-view multimodal inputs of the scene, we progressively construct a multimodal scene representation. Each modality is represented by our modality modeling module, which includes modality-specific features and a modality indicator. The densification process is integrated with our multimodal decomposition mechanism, which disentangles multi￾modal information based on gradient … view at source ↗
Figure 3
Figure 3. Distribution differences among different modality in [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Multimodal Decomposition Mechanism. (a) Multi￾modal Prune: “Hard Prune” refers to directly pruning the Gaus￾sian, while “Soft Prune” involves pruning modality “A” but re￾taining modality “B”. (b) Multimodal Decomposition: Instead of cloning the Gaussian, multimodal dec…
Figure 5
Figure 5. Figure 5: Qualitative results of RGB-Thermal among our MMOne, 3DGS, and ThermalGaussian (abbreviated as “T-GS”). [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Qualitative results of RGB-Language among our MMOne, LangSplat (“LS*”), and joint training version of LangSplat (“LS-J”). [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Qualitative results of RGB-Thermal-Language between [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Gaussian Distributions. The total number of Gaussians is 89.5K. “T-L Gaussians” are omitted for visual clarity. position mechanism, its performance remains suboptimal without the disentangling of modalities, due to the varying levels of granularity among them [PITH_FU…
Figure 9
Figure 9. Figure 9: Qualitative results of incorporating monocular depth. [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Qualitative results of the modality conflicts caused by the introduction of language. “T-GS” refers to ThermalGaussian and [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 53 canonical work pages

  1. [1]

    Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields

    Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields. In Proceedings of the International Conference on Computer Vision, 2021. 2

  2. [2]

    Emerg- ing Properties in Self-Supervised Vision Transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing Properties in Self-Supervised Vision Transformers. In Proceedings of the International Conference on Computer Vision, 2021. 2

  3. [3]

    PGSR: Planar-based Gaussian Splat- ting for Efficient and High-Fidelity Surface Reconstruction

    Danpeng Chen, Hai Li, Weicai Ye, Yifan Wang, Weijian Xie, Shangjin Zhai, Nan Wang, Haomin Liu, Hujun Bao, and Guofeng Zhang. PGSR: Planar-based Gaussian Splat- ting for Efficient and High-Fidelity Surface Reconstruction. IEEE Transactions on Visualization and Computer Graph- ics, 2024. 2

  4. [4]

    A Survey on 3D Gaussian Splatting

    Guikun Chen and Wenguan Wang. A Survey on 3D Gaussian Splatting. arXiv preprint arXiv:2401.03890, 2024. 2

  5. [5]

    Thermal3D-GS: Physics-induced 3D Gaussians for Thermal Infrared Novel- view Synthesis

    Qian Chen, Shihao Shu, and Xiangzhi Bai. Thermal3D-GS: Physics-induced 3D Gaussians for Thermal Infrared Novel- view Synthesis. In Proceedings of the European Conference on Computer Vision, 2024. 2

  6. [6]

    Tactile-Augmented Radiance Fields

    Yiming Dou, Fengyu Yang, Yi Liu, Antonio Loquercio, and Andrew Owens. Tactile-Augmented Radiance Fields. In Proceedings of the Conference on Computer Vision and Pat- tern Recognition, 2024. 2

  7. [7]

    Multimodal Sensors and ML-Based Data Fusion for Advanced Robots

    Shengshun Duan, Qiongfeng Shi, and Jun Wu. Multimodal Sensors and ML-Based Data Fusion for Advanced Robots. Advanced Intelligent Systems, 4(12):2200213, 2022. 1

  8. [8]

    Privacy-Preserving Person Detection Using Low-Resolution Infrared Cameras

    Thomas Dubail, Fidel Alejandro Guerrero Pe ˜na, Heitor Rapela Medeiros, Masih Aminbeidokhti, Eric Granger, and Marco Pedersoli. Privacy-Preserving Person Detection Using Low-Resolution Infrared Cameras. In Proceedings of the European Conference on Computer Vision, 2022. 2

Show all 59 references
  1. [9]

    Scene Perception in the Human Brain

    Russell A Epstein and Chris I Baker. Scene Perception in the Human Brain. Annual Review of Vision Science , 5(1): 373–397, 2019. 1

  2. [10]

    Mini-Splatting: Represent- ing Scenes with a Constrained Number of Gaussians

    Guangchi Fang and Bing Wang. Mini-Splatting: Represent- ing Scenes with a Constrained Number of Gaussians. InPro- ceedings of the European Conference on Computer Vision ,

  3. [11]

    Mini-Splatting2: Building 360 Scenes within Minutes via Aggressive Gaussian Densi- fication

    Guangchi Fang and Bing Wang. Mini-Splatting2: Building 360 Scenes within Minutes via Aggressive Gaussian Densi- fication. arXiv preprint arXiv:2411.12788, 2024. 2

  4. [12]

    NeRF: Neural Radiance Field in 3D Vision: A Comprehensive Review

    Kyle Gao, Yina Gao, Hongjie He, Dening Lu, Linlin Xu, and Jonathan Li. NeRF: Neural Radiance Field in 3D Vision: A Comprehensive Review. arXiv preprint arXiv:2210.00379 ,

  5. [13]

    Deep Learning for 3D Point Clouds: A Survey

    Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu, Li Liu, and Mohammed Bennamoun. Deep Learning for 3D Point Clouds: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(12):4338–4364, 2020. 1

  6. [14]

    Illuminating the Dark Spaces of Healthcare with Ambient Intelligence

    Albert Haque, Arnold Milstein, and Li Fei-Fei. Illuminating the Dark Spaces of Healthcare with Ambient Intelligence. Nature, 585(7824):193–202, 2020. 2

  7. [15]

    ThermoNeRF: Multimodal Neural Radiance Fields for Thermal Novel View Synthesis

    Mariam Hassan, Florent Forest, Olga Fink, and Mal- colm Mielle. ThermoNeRF: Multimodal Neural Radiance Fields for Thermal Novel View Synthesis. arXiv preprint arXiv:2403.12154, 2024. 1, 2

  8. [16]

    2D Gaussian Splatting for Geometrically Ac- curate Radiance Fields

    Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2D Gaussian Splatting for Geometrically Ac- curate Radiance Fields . In SIGGRAPH, 2024. 2

  9. [17]

    Rethinking Visual Scene Perception

    Helene Intraub. Rethinking Visual Scene Perception. Wiley Interdisciplinary Reviews: Cognitive Science, 3(1):117–127,

  10. [18]

    Visible and Infrared Imaging Based Inspection of Power Installation

    B Jalil, MA Pascali, GR Leone, M Martinelli, D Moroni, O Salvetti, and A Berton. Visible and Infrared Imaging Based Inspection of Power Installation. Pattern Recognition and Image Analysis, 29(1):35–41, 2019. 2

  11. [19]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics, 42(4):139–1, 2023. 1, 2, 3, 5

  12. [20]

    LERF: Language Embed- ded Radiance Fields

    Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik. LERF: Language Embed- ded Radiance Fields. In Proceedings of the International Conference on Computer Vision, 2023. 1, 2, 5

  13. [21]

    Segment Any- thing

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment Any- thing. In Proceedings of the International Conference on Computer Vision, 2023. 5, 6

  14. [22]

    LangSurf: Language-Embedded Surface Gaussians for 3D Scene Un- derstanding

    Hao Li, Roy Qin, Zhengyu Zou, Diqi He, Bohan Li, Bingquan Dai, Dingewn Zhang, and Junwei Han. LangSurf: Language-Embedded Surface Gaussians for 3D Scene Un- derstanding. arXiv preprint arXiv:2412.17635, 2024. 2

  15. [23]

    Weakly Supervised 3D Open- vocabulary Segmentation

    Kunhao Liu, Fangneng Zhan, Jiahui Zhang, Muyu Xu, Yingchen Yu, Abdulmotaleb El Saddik, Christian Theobalt, Eric Xing, and Shijian Lu. Weakly Supervised 3D Open- vocabulary Segmentation. Advances in Neural Information Processing Systems, 36:53433–53456, 2023. 1, 2

  16. [24]

    EfficientGS: Stream- lining Gaussian Splatting for Large-Scale High-Resolution Scene Representation

    Wenkai Liu, Tao Guan, Bin Zhu, Luoyuan Xu, Zikai Song, Dan Li, Yuesong Wang, and Wei Yang. EfficientGS: Stream- lining Gaussian Splatting for Large-Scale High-Resolution Scene Representation. IEEE MultiMedia, 2025. 2

  17. [25]

    Thermal- Gaussian: Thermal 3D Gaussian Splatting

    Rongfeng Lu, Hangyu Chen, Zunjie Zhu, Yuhang Qin, Ming Lu, Le Zhang, Chenggang Yan, and Anke Xue. Thermal- Gaussian: Thermal 3D Gaussian Splatting. InProceedings of the International Conference on Learning Representations ,

  18. [26]

    Taming 3DGS: High-Quality Radiance Fields with Limited Resources

    Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl, Markus Steinberger, Francisco Vicente Carrasco, and Fer- nando De La Torre. Taming 3DGS: High-Quality Radiance Fields with Limited Resources. In SIGGRAPH Asia, 2024. 2

  19. [27]

    NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. Communications of the ACM , 65(1):99–106,

  20. [28]

    Recent Indus- trial Applications of Infrared Thermography: A Review

    Roque Alfredo Osornio-Rios, Jose Alfonso Antonino-Daviu, and Rene de Jesus Romero-Troncoso. Recent Indus- trial Applications of Infrared Thermography: A Review. 9 IEEE Transactions on Industrial Informatics , 15(2):615– 625, 2018. 2

  21. [29]

    Exploring Multi-modal Neural Scene Rep- resentations With Applications on Thermal Imaging

    Mert ¨Ozer, Maximilian Weiherer, Martin Hundhausen, and Bernhard Egger. Exploring Multi-modal Neural Scene Rep- resentations With Applications on Thermal Imaging. In Pro- ceedings of the European Conference on Computer Vision ,

  22. [30]

    3D Vision-Language Gaussian Splatting

    Qucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng, Anwesa Choudhuri, Terrence Chen, Chen Chen, and Ziyan Wu. 3D Vision-Language Gaussian Splatting. In Pro- ceedings of the International Conference on Learning Rep- resentations, 2025. 1, 2

  23. [31]

    Cross- Spectral Neural Radiance Fields

    Matteo Poggi, Pierluigi Zama Ramirez, Fabio Tosi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano. Cross- Spectral Neural Radiance Fields. In Proceedings of the In- ternational Conference on 3D Vision, 2022. 2

  24. [32]

    Gaus- sianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians

    Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain, and Matthias Nießner. Gaus- sianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians. In Proceedings of the Conference on Computer Vision and Pattern Recognition, 2024. 2

  25. [33]

    3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting

    Zhiyin Qian, Shaofei Wang, Marko Mihajlovic, Andreas Geiger, and Siyu Tang. 3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting. In Proceedings of the Conference on Computer Vision and Pattern Recognition,

  26. [34]

    LangSplat: 3D Language Gaussian Splat- ting

    Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. LangSplat: 3D Language Gaussian Splat- ting. In Proceedings of the Conference on Computer Vision and Pattern Recognition, 2024. 1, 2, 3, 5, 6, 7

  27. [35]

    GLS: Geometry-aware 3D Language Gaussian Splatting

    Jiaxiong Qiu, Liu Liu, Zhizhong Su, and Tianwei Lin. GLS: Geometry-aware 3D Language Gaussian Splatting. arXiv preprint arXiv:2411.18066, 2024. 2

  28. [36]

    GOI: Find 3D Gaussians of Interest with an Optimizable Open- vocabulary Semantic-space Hyperplane

    Yansong Qu, Shaohui Dai, Xinyang Li, Jianghang Lin, Li- ujuan Cao, Shengchuan Zhang, and Rongrong Ji. GOI: Find 3D Gaussians of Interest with an Optimizable Open- vocabulary Semantic-space Hyperplane. In Proceedings of the International Conference on Multimedia, 2024. 2

  29. [37]

    Learning Transferable Visual Models From Natural Language Super- vision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning Transferable Visual Models From Natural Language Super- vision . In Proceedings of the International Conference o...

  30. [38]

    Graph-based Thermal-Inertial SLAM with Probabilistic Neural Networks

    Muhamad Risqi U Saputra, Chris Xiaoxuan Lu, Pedro Porto B de Gusmao, Bing Wang, Andrew Markham, and Niki Trigoni. Graph-based Thermal-Inertial SLAM with Probabilistic Neural Networks. IEEE Transactions on Robotics, 38(3):1875–1893, 2021. 2

  31. [39]

    Structure- From-Motion Revisited

    Johannes L Schonberger and Jan-Michael Frahm. Structure- From-Motion Revisited. In Proceedings of the Conference on Computer Vision and Pattern Recognition, 2016. 5

  32. [40]

    Language Embedded 3D Gaussians for Open- V ocabulary Scene Understanding

    Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao- Hua Guan. Language Embedded 3D Gaussians for Open- V ocabulary Scene Understanding. In Proceedings of the Conference on Computer Vision and Pattern Recognition ,

  33. [41]

    Splat-MOVER: Multi-Stage, Open-V ocabulary Robotic Manipulation via Editable Gaus- sian Splatting

    Olaolu Shorinwa, Johnathan Tucker, Aliyah Smith, Aiden Swann, Timothy Chen, Roya Firoozi, Monroe David Kennedy, and Mac Schwager. Splat-MOVER: Multi-Stage, Open-V ocabulary Robotic Manipulation via Editable Gaus- sian Splatting. In Proceedings of the Conference on Robot Learni...

  34. [42]

    Implicit Neural Representa- tions with Periodic Activation Functions

    Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit Neural Representa- tions with Periodic Activation Functions. Advances in Neu- ral Information Processing Systems, 33:7462–7473, 2020. 1

  35. [43]

    Touch- GS: Visual-Tactile Supervised 3D Gaussian Splatting

    Aiden Swann, Matthew Strong, Won Kyung Do, Gadiel Sz- naier Camps, Mac Schwager, and Monroe Kennedy. Touch- GS: Visual-Tactile Supervised 3D Gaussian Splatting. In Proceedings of the International Conference on Intelligent Robots and Systems, 2024. 2

  36. [44]

    Nerfstudio: A Modular Framework for Neural Radiance Field Develop- ment

    Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al. Nerfstudio: A Modular Framework for Neural Radiance Field Develop- ment. In SIGGRAPH, 2023. 1, 2

  37. [45]

    DN-Splatter: Depth and Normal Priors for Gaussian Splatting and Meshing

    Matias Turkulainen, Xuqian Ren, Iaroslav Melekhov, Otto Seiskari, Esa Rahtu, and Juho Kannala. DN-Splatter: Depth and Normal Priors for Gaussian Splatting and Meshing. In Proceedings of the Winter Conference on Applications of Computer Vision, 2025. 2

  38. [46]

    The Development of Multisensory Pro- cesses

    Mark T Wallace. The Development of Multisensory Pro- cesses. Cognitive Processing, 5:69–83, 2004. 1

  39. [47]

    4D Gaussian Splatting for Real-Time Dynamic Scene Ren- dering

    Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4D Gaussian Splatting for Real-Time Dynamic Scene Ren- dering. In Proceedings of the Conference on Computer Vi- sion and Pattern Recognition, 2024. 2

  40. [48]

    OpenGaussian: Towards Point- Level 3D Gaussian-based Open V ocabulary Understanding

    Yanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu, Ya- hao Shi, Xinhua Cheng, Chen Zhao, Haocheng Feng, Errui Ding, Jingdong Wang, et al. OpenGaussian: Towards Point- Level 3D Gaussian-based Open V ocabulary Understanding. Advances in Neural Information Processing Systems , 37: 1...

  41. [49]

    PhysGaussian: Physics- Integrated 3D Gaussians for Generative Dynamics

    Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu Jiang. PhysGaussian: Physics- Integrated 3D Gaussians for Generative Dynamics. In Pro- ceedings of the Conference on Computer Vision and Pattern Recognition, 2024. 2

  42. [50]

    Street Gaussians: Modeling Dynamic Ur- ban Scenes with Gaussian Splatting

    Yunzhi Yan, Haotong Lin, Chenxu Zhou, Weijie Wang, Haiyang Sun, Kun Zhan, Xianpeng Lang, Xiaowei Zhou, and Sida Peng. Street Gaussians: Modeling Dynamic Ur- ban Scenes with Gaussian Splatting. In Proceedings of the European Conference on Computer Vision, 2024. 1, 2

  43. [51]

    Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes

    Zehao Yu, Torsten Sattler, and Andreas Geiger. Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes. ACM Transactions on Graphics, 43 (6):1–13, 2024. 2

  44. [52]

    Mul- timodal Intelligence: Representation Learning, Information Fusion, and Applications

    Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng. Mul- timodal Intelligence: Representation Learning, Information Fusion, and Applications. IEEE Journal of Selected Topics in Signal Processing, 14(3):478–493, 2020. 1

  45. [53]

    Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields

    Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Ze- hao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang 10 Wang, and Achuta Kadambi. Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields. In Proceedings of the Conference on Computer Vis...

  46. [54]

    DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes

    Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes. In Proceedings of the Con- ference on Computer Vision and Pattern Recognition, 2024. 1, 2 11 M...

  47. [55]

    Soft Prune

    Additional Implementation Details In this section, we introduce the implementation details of the RGB-Thermal, RGB-Language, and RGB-Thermal- Language experiments. We also introduce the hyperparam- eter settings in our proposed module. RGB-Thermal. We adhere to the experimenta...

  48. [56]

    MM-J”. “MM

    Additional Ablation Studies To further investigate the sensitivity of the threshold setting of multimodal decomposition, we conduct an additional ab- lation study. As shown in Tab. 6, our chosen threshold con- sistently achieves superior performance across all metrics. Moreove...

  49. [57]

    T-GS” and “MM-J

    Additional Qualitative Results To analyze the distributions of multimodal and single- modal Gaussians, we use the “Dimsum” scene as an exam- ple. As shown in Fig. 8, different modalities require varying number of Gaussians. We also present the qualitative results of modality c...

  50. [58]

    F-GS” and “LS-J

    Additional Quantitative Results We present the number of Gaussians in the RGB-Thermal and RGB-Language experiments to further highlight the ef- fectiveness of our method in achieving a compact repre- sentation. As shown in Tab. 8, our method uses approxi- mately one-third of t...

  51. [59]

    T-GS” refers to ThermalGaussian and “MM-J

    Additional Experiments on Scalability We further validate scalability by incorporating monocular depth in the “Dimsum” scene. The results in Fig. 9 and Tab. 12 show that the rendered depth quality is enhanced, particularly on flat surfaces, without compromising the per- forman...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.