Pith. sign in

REVIEW 9 cited by

RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.14604 v1 pith:FPFWUYVQ submitted 2025-04-20 cs.RO

RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots

classification cs.RO
keywords scenegaussiansrobooccgeometricrobotssemanticfine-grainedgeometry
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

3D occupancy prediction enables the robots to obtain spatial fine-grained geometry and semantics of the surrounding scene, and has become an essential task for embodied perception. Existing methods based on 3D Gaussians instead of dense voxels do not effectively exploit the geometry and opacity properties of Gaussians, which limits the network's estimation of complex environments and also limits the description of the scene by 3D Gaussians. In this paper, we propose a 3D occupancy prediction method which enhances the geometric and semantic scene understanding for robots, dubbed RoboOcc. It utilizes the Opacity-guided Self-Encoder (OSE) to alleviate the semantic ambiguity of overlapping Gaussians and the Geometry-aware Cross-Encoder (GCE) to accomplish the fine-grained geometric modeling of the surrounding scene. We conduct extensive experiments on Occ-ScanNet and EmbodiedOcc-ScanNet datasets, and our RoboOcc achieves state-of the-art performance in both local and global camera settings. Further, in ablation studies of Gaussian parameters, the proposed RoboOcc outperforms the state-of-the-art methods by a large margin of (8.47, 6.27) in IoU and mIoU metric, respectively. The codes will be released soon.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception

    cs.RO 2026-04 unverdicted novelty 7.0

    RobotPan predicts metric-scaled compact 3D Gaussians from calibrated multi-view inputs via spherical coordinates and hierarchical voxel priors for real-time 360° robotic perception and reconstruction.

  2. O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Urban Autonomous Agents

    cs.CV 2026-03 conditional novelty 6.5

    O3N is the first end-to-end pure-vision framework for open-vocabulary 3D occupancy from a single omnidirectional RGB image, using polar-spiral Mamba, cost aggregation, and gradient-free modality alignment.

  3. GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory

    cs.RO 2026-07 conditional novelty 6.0

    GEM-Occ converts transient visual geometry into semantic Gaussian and free-space ray evidence, fuses them into hierarchical occupancy memory, and beats prior indoor occupancy baselines on the new HIOcc benchmark.

  4. Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images

    cs.CV 2026-07 unverdicted novelty 6.0

    COVScene is a pose-free framework that lifts semantic Gaussians into a volumetric occupancy field during training to jointly support novel view synthesis, open-vocabulary segmentation, and semantic occupancy prediction.

  5. VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding

    cs.CV 2026-05 unverdicted novelty 6.0

    VEOcc is a voxel-based online semantic occupancy prediction method using recursive assimilation and three update modules (TLA, RCM, CSU) that reports new SOTA results on Occ-ScanNet and EmbodiedOcc-ScanNet.

  6. FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction

    cs.RO 2026-04 unverdicted novelty 6.0

    FreeOcc enables training-free open-vocabulary 3D occupancy prediction from RGB-D sequences by combining SLAM, dense Gaussian maps, off-the-shelf vision-language models, and probabilistic projection, achieving over 2x ...

  7. O3N: Omnidirectional Open-Vocabulary Occupancy Prediction for Urban Autonomous Agents

    cs.CV 2026-03 conditional novelty 6.0

    O3N is the first open-vocabulary occupancy prediction method that takes a single omnidirectional RGB image and labels 3D voxels with both seen and unseen semantic classes.

  8. Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes

    cs.CV 2026-02 unverdicted novelty 6.0

    A 3D Language-Embedded Gaussians framework with opacity-aware Poisson volumetric aggregation and progressive temperature decay achieves 59.50 IoU and 21.05 mIoU on Occ-ScanNet for open-vocabulary indoor occupancy.

  9. GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

    cs.CV 2026-07 conditional novelty 5.0

    A unified framework converts surface geometry priors into sparse Gaussian occupancy predictions and extends it to multi-view and temporal inputs.