Intrinsic-GS recovers object-level segmentation in 4D Gaussian scenes from intrinsic cues alone via affinity graph and Leiden partitioning, reaching 0.746 mIoU on Neu3D and 0.575 on HyperNeRF without mask supervision.
Segment any 3d gaussians
11 Pith papers cite this work, alongside 3 external citations. Polarity classification is still indexing.
abstract
This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching an scale-gated affinity feature to each 3D Gaussian to endow it a new property towards multi-granularity segmentation. Specifically, a scale-aware contrastive training strategy is proposed for the scale-gated affinity feature learning. It 1) distills the segmentation capability of the Segment Anything Model (SAM) from 2D masks into the affinity features and 2) employs a soft scale gate mechanism to deal with multi-granularity ambiguity in 3D segmentation through adjusting the magnitude of each feature channel according to a specified 3D physical scale. Evaluations demonstrate that SAGA achieves real-time multi-granularity segmentation with quality comparable to state-of-the-art methods. As one of the first methods addressing promptable segmentation in 3D-GS, the simplicity and effectiveness of SAGA pave the way for future advancements in this field. Our code will be released.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.
FalconTrack automates photorealistic dataset creation via Gaussian Splatting and achieves high zero-shot sim-to-real performance in vision-based aerial tracking using multi-head perception and class-conditioned EKF.
Diffusion-based per-view harmonization for lighting-consistent object transfer between 3DGS scenes, using heterogeneous training data and final 3D consolidation.
EPS3D is an end-to-end architecture for 3D panoptic segmentation from multi-view images that uses distillation and semantic-instance mutual enhancement to achieve higher benchmark performance and speed than prior methods.
A 3D object codebook leveraging mask semantics and Gaussian spatial information enables multi-view mask association for indoor asset detection in 3DGS scenes, yielding 65% F1 and 11% mAP gains on two large indoor scenes.
TrianguLang achieves state-of-the-art feed-forward text-guided 3D localization and segmentation by using predicted geometry to gate cross-view semantic correspondences without ground-truth poses.
Semantic Foam extends Radiant Foam with explicit cell-level semantic feature fields and spatial regularization to improve object segmentation consistency in scene representations.
FF3R unifies geometric and semantic 3D reconstruction in a single annotation-free feed-forward network trained solely via RGB and feature rendering supervision.
A survey compiling principles, applications, benchmarks, and challenges of 3D Gaussian Splatting for explicit 3D scene representation.
citing papers explorer
-
Intrinsic 4D Gaussian Segmentation from Scene Cues
Intrinsic-GS recovers object-level segmentation in 4D Gaussian scenes from intrinsic cues alone via affinity graph and Leiden partitioning, reaching 0.746 mIoU on Neu3D and 0.575 on HyperNeRF without mask supervision.
-
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.
-
FalconTrack: Photorealistic Auto-Labeled Perception and Physics-Aware Vision-Based Aerial Tracking
FalconTrack automates photorealistic dataset creation via Gaussian Splatting and achieves high zero-shot sim-to-real performance in vision-based aerial tracking using multi-head perception and class-conditioned EKF.
-
Lighting-Consistent Object Transfer Across Radiance Fields
Diffusion-based per-view harmonization for lighting-consistent object transfer between 3DGS scenes, using heterogeneous training data and final 3D consolidation.
-
EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation
EPS3D is an end-to-end architecture for 3D panoptic segmentation from multi-view images that uses distillation and semantic-instance mutual enhancement to achieve higher benchmark performance and speed than prior methods.
-
Indoor Asset Detection in Large Scale 360{\deg} Drone-Captured Imagery via 3D Gaussian Splatting
A 3D object codebook leveraging mask semantics and Gaussian spatial information enables multi-view mask association for indoor asset detection in 3DGS scenes, yielding 65% F1 and 11% mAP gains on two large indoor scenes.
-
TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization
TrianguLang achieves state-of-the-art feed-forward text-guided 3D localization and segmentation by using predicted geometry to gate cross-view semantic correspondences without ground-truth poses.
-
Semantic Foam: Unifying Spatial and Semantic Scene Decomposition
Semantic Foam extends Radiant Foam with explicit cell-level semantic feature fields and spatial regularization to improve object segmentation consistency in scene representations.
-
FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views
FF3R unifies geometric and semantic 3D reconstruction in a single annotation-free feed-forward network trained solely via RGB and feature rendering supervision.
-
A Survey on 3D Gaussian Splatting
A survey compiling principles, applications, benchmarks, and challenges of 3D Gaussian Splatting for explicit 3D scene representation.
- TranSplat: Instant Object Relighting in Gaussian Splatting via Spherical Harmonic Radiance Transfer