Pith. sign in

REVIEW 5 cited by

GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.13193 v2 pith:DRULDJNL submitted 2024-12-17 cs.CV

classification cs.CV
keywords gausstrspatialunderstandingfoundationgaussiangaussiansoccupancyprediction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D Semantic Occupancy Prediction is fundamental for spatial understanding, yet existing approaches face challenges in scalability and generalization due to their reliance on extensive labeled data and computationally intensive voxel-wise representations. In this paper, we introduce GaussTR, a novel Gaussian-based Transformer framework that unifies sparse 3D modeling with foundation model alignment through Gaussian representations to advance 3D spatial understanding. GaussTR predicts sparse sets of Gaussians in a feed-forward manner to represent 3D scenes. By splatting the Gaussians into 2D views and aligning the rendered features with foundation models, GaussTR facilitates self-supervised 3D representation learning and enables open-vocabulary semantic occupancy prediction without requiring explicit annotations. Empirical experiments on the Occ3D-nuScenes dataset demonstrate GaussTR's state-of-the-art zero-shot performance of 12.27 mIoU, along with a 40% reduction in training time. These results highlight the efficacy of GaussTR for scalable and holistic 3D spatial understanding, with promising implications in autonomous driving and embodied agents. The code is available at https://github.com/hustvl/GaussTR.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Factorized Dense Routing approximates unconstrained 2D-to-3D feature mixing by hierarchical tensor contractions, yielding global-context occupancy prediction that remains robust without camera extrinsics.

  2. UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene via Rendering Fusion

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    UnsOcc proposes RenderFusion and GSRefinement to improve 3D semantic occupancy prediction in unstructured scenes by enhancing cross-modal fusion and long-tail supervision, outperforming SOTA on a new mine dataset and ...

  3. Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A class-conditional gradient loss (Causal Loss) plus channel-grouped lifting, learnable camera offsets, and normalized convolution raises Occ3D mIoU by 1.2/0.8 points and cuts the camera-noise mIoU drop from 32% to 7%.

  4. GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A camera-only Gaussian-surfel pipeline reconstructs full Waymo scenes, converts them to binary occupancy labels, and trains CVT-Occ to generalize on Occ3D-Waymo and Occ3D-nuScenes at a level close to or above LiDAR-la...

  5. Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SceneDINO performs semantic scene completion from a single image in a fully unsupervised way by lifting self-supervised DINO features into a 3D feature field trained with multi-view consistency.

Pith tools