Pith. sign in

REVIEW 12 cited by

Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13959 v6 pith:6OL5FPBW submitted 2023-03-24 cs.CV

classification cs.CV
keywords semanticscenestereocompletionreliablerepresentationdensegeometry
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to predict accurate semantic scenes due to inherent geometric ambiguity and incomplete observations. In this paper, we resort to stereo matching technique and bird's-eye-view (BEV) representation learning to address such issues in SSC. Complementary to each other, stereo matching mitigates geometric ambiguity with epipolar constraint while BEV representation enhances the hallucination ability for invisible regions with global semantic context. However, due to the inherent representation gap between stereo geometry and BEV features, it is non-trivial to bridge them for dense prediction task of SSC. Therefore, we further develop a unified occupancy-based framework dubbed BRGScene, which effectively bridges these two representations with dense 3D volumes for reliable semantic scene completion. Specifically, we design a novel Mutual Interactive Ensemble (MIE) block for pixel-level reliable aggregation of stereo geometry and BEV features. Within the MIE block, a Bi-directional Reliable Interaction (BRI) module, enhanced with confidence re-weighting, is employed to encourage fine-grained interaction through mutual guidance. Besides, a Dual Volume Ensemble (DVE) module is introduced to facilitate complementary aggregation through channel-wise recalibration and multi-group voting. Our method outperforms all published camera-based methods on SemanticKITTI for semantic scene completion. Our code is available on https://github.com/Arlo0o/StereoScene.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Step Closer: Creating the Future to Boost Monocular Semantic Scene Completion

    cs.CV 2025-07 conditional novelty 7.0 of 10

    CF-SSC predicts pseudo-future frames from past and current monocular images and fuses them in 3D, achieving state-of-the-art semantic scene completion on SemanticKITTI and SSCBench-KITTI-360.

  2. OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A joint diffusion framework trains a Stable Diffusion generator and a semantic occupancy perception model together, so each task improves the other, producing text-conditional RGB-occupancy pairs.

  3. RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion

    cs.CV 2026-08 conditional novelty 6.0 of 10

    RayLift lifts complementary ray-wise evidence from stereo and a 3D foundation model to improve camera-based 3D semantic scene completion.

  4. Geospatial-Prior Guidance for 3D Semantic Scene Completion

    cs.CV 2026-08 conditional novelty 6.0 of 10

    GeoScene uses weighted fusion of satellite imagery and OpenStreetMap priors to improve camera-based 3D semantic scene completion on SemanticKITTI and SSCBench-KITTI-360.

  5. GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory

    cs.RO 2026-07 conditional novelty 6.0 of 10

    GEM-Occ converts transient visual geometry into semantic Gaussian and free-space ray evidence, fuses them into hierarchical occupancy memory, and beats prior indoor occupancy baselines on the new HIOcc benchmark.

  6. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  7. OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

    cs.AI 2025-09 conditional novelty 6.0 of 10

    OccVLA trains a vision-language-action model to predict 3D occupancy as an auxiliary output, improving nuScenes trajectory planning and 3D VQA from camera images only, with the occupancy branch disabled at inference.

  8. VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible Regions

    cs.CV 2025-07 conditional novelty 6.0 of 10

    VisHall3D splits monocular 3D scene completion into visible-region reconstruction (VisFrontierNet) and invisible-region hallucination (OcclusionMAE), reporting SOTA mIoU of 17.46 and 20.95 on SemanticKITTI and SSCBenc...

  9. Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dual-stream BEV architecture that separates instance and scene class queries achieves state-of-the-art mIoU of 17.35 on SemanticKITTI and 20.55 on SSCBench-KITTI-360.

  10. VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection

    cs.GR 2025-06 conditional novelty 6.0 of 10

    VoxDet reformulates 3D semantic occupancy prediction as dense object detection by deriving instance-boundary offsets from voxel class labels, and reports new state-of-the-art results on camera and LiDAR benchmarks.

  11. Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...

  12. Challenger: Affordable Adversarial Driving Video Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A framework for automatic generation of photorealistic adversarial driving videos, shown to sharply increase collision rates of end-to-end autonomous driving models.

Pith tools