Pith. sign in

REVIEW 10 cited by

Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.15624 v2 pith:MJC45WJR submitted 2024-03-22 cs.CV

Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting

classification cs.CV
keywords semanticgaussiansscenesegmentationunderstandingmethodsopen-vocabularyapplications
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Open-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neurel rendering methods as 3D representations and jointly optimize color and semantic features to achieve rendering and scene understanding simultaneously. In this paper, we introduce Semantic Gaussians, a novel open-vocabulary scene understanding approach based on 3D Gaussian Splatting. Our key idea is to distill knowledge from 2D pre-trained models to 3D Gaussians. Unlike existing methods, we design a versatile projection approach that maps various 2D semantic features from pre-trained image encoders into a novel semantic component of 3D Gaussians, which is based on spatial relationship and need no additional training. We further build a 3D semantic network that directly predicts the semantic component from raw 3D Gaussians for fast inference. The quantitative results on ScanNet segmentation and LERF object localization demonstates the superior performance of our method. Additionally, we explore several applications of Semantic Gaussians including object part segmentation, instance segmentation, scene editing, and spatiotemporal segmentation with better qualitative results over 2D and 3D baselines, highlighting its versatility and effectiveness on supporting diverse downstream tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splatting

    cs.CV 2026-02 unverdicted novelty 7.0

    BrepGaussian reconstructs B-Rep CAD models from multi-view images using a two-stage Gaussian Splatting method that separates geometry and edge capture from patch feature refinement.

  2. E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding

    cs.CV 2026-07 conditional novelty 6.0

    3D Gaussian view-dependent colors are repacked as 3×3 matrices so geometry and color rotate together, giving exact rotation-equivariant recognition and world modeling in 3DGS.

  3. Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images

    cs.CV 2026-07 unverdicted novelty 6.0

    COVScene is a pose-free framework that lifts semantic Gaussians into a volumetric occupancy field during training to jointly support novel view synthesis, open-vocabulary segmentation, and semantic occupancy prediction.

  4. EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation

    cs.CV 2026-06 unverdicted novelty 6.0

    EPS3D is an end-to-end architecture for 3D panoptic segmentation from multi-view images that uses distillation and semantic-instance mutual enhancement to achieve higher benchmark performance and speed than prior methods.

  5. STaR-Quant: State-Time Consistent Post-Training Quantization for Diffusion Large Language Models

    cs.LG 2026-06 unverdicted novelty 6.0

    STaR-Quant provides a state-time consistent PTQ framework for DLLMs using SGAT and TAC to improve low-bit weight-activation quantization.

  6. TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing

    cs.CV 2026-06 unverdicted novelty 6.0

    TASE introduces truncation-aware semantic embeddings from 2D features for controllable text-driven 3D scene editing with added multi-view consistency via equivariance loss and diffusion finetuning.

  7. Part-Level 3D Gaussian Vehicle Generation with Joint and Hinge Axis Estimation

    cs.AI 2026-04 unverdicted novelty 6.0

    A new framework generates part-level animatable 3D Gaussian vehicles from images by adding modules for exclusive part ownership and kinematic joint/axis prediction.

  8. SAD-GS: Learning Reliable 3D Semantic Gaussian Fields via Dynamic Geo-Semantic Anchoring

    cs.CV 2026-06 unverdicted novelty 5.0

    SAD-GS proposes dynamic geo-semantic anchoring via SAD and GSFL to learn reliable 3D semantic Gaussian fields, reporting best performance on LERF-OVS, 3D-OVS, and Mip-NeRF360 for open-vocabulary localization and segmentation.

  9. Enhancing LLM Training via Spectral Clipping

    cs.LG 2026-03 unverdicted novelty 5.0

    SPECTRA improves LLM pretraining via post-clipping of update spectral norms and optional pre-clipping of gradient spikes, framed as Composite Frank-Wolfe regularization.

  10. A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation

    cs.CV 2025-08 unverdicted novelty 3.0

    A survey that categorizes and summarizes methods applying 3D Gaussian Splatting to segmentation, editing, generation, and related tasks, including datasets and evaluation protocols.