REVIEW 5 cited by
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines.
Forward citations
Cited by 5 Pith papers
-
E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding
3D Gaussian view-dependent colors are repacked as 3×3 matrices so geometry and color rotate together, giving exact rotation-equivariant recognition and world modeling in 3DGS.
-
MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction
Semantically enriched MASt3R correspondences plus a multi-attribute 3D consistency loss raise sparse-view ScanNet++ PSNR by >4.5 dB over Splatt3R and preserve quality under wide baselines.
-
Rectifying Mask via Entropy for Distractor-Free 3DGS in Ambiguous Scenarios
RefineSplat removes ambiguous distractors from 3DGS via entropy-aware adaptive masking and density control, releasing an 18-scene Ambiguous wild dataset and reporting SOTA metrics on multiple wild benchmarks.
-
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
GaussianVLM embeds language features per Gaussian splat, sparsifies them by task and location, and reports state-of-the-art results on embodied 3D reasoning benchmarks.
-
Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding
A superpoint-guided, scale-normalized tokenizer lets a frozen CLIP model perform 3D segmentation and classification without fine-tuning.
Discussion (0). Sign in to comment.