REVIEW 9 cited by
OccDepth: A Depth-Aware Method for 3D Semantic Scene Completion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
3D Semantic Scene Completion (SSC) can provide dense geometric and semantic scene representations, which can be applied in the field of autonomous driving and robotic systems. It is challenging to estimate the complete geometry and semantics of a scene solely from visual images, and accurate depth information is crucial for restoring 3D geometry. In this paper, we propose the first stereo SSC method named OccDepth, which fully exploits implicit depth information from stereo images (or RGBD images) to help the recovery of 3D geometric structures. The Stereo Soft Feature Assignment (Stereo-SFA) module is proposed to better fuse 3D depth-aware features by implicitly learning the correlation between stereo images. In particular, when the input are RGBD image, a virtual stereo images can be generated through original RGB image and depth map. Besides, the Occupancy Aware Depth (OAD) module is used to obtain geometry-aware 3D features by knowledge distillation using pre-trained depth models. In addition, a reformed TartanAir benchmark, named SemanticTartanAir, is provided in this paper for further testing our OccDepth method on SSC task. Compared with the state-of-the-art RGB-inferred SSC method, extensive experiments on SemanticKITTI show that our OccDepth method achieves superior performance with improving +4.82% mIoU, of which +2.49% mIoU comes from stereo images and +2.33% mIoU comes from our proposed depth-aware method. Our code and trained models are available at https://github.com/megvii-research/OccDepth.
Forward citations
Cited by 9 Pith papers
-
PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and Consistency
PacGDC synthesizes diverse pseudo training geometries by rescaling depth predictions from foundation models, improving zero-shot and few-shot generalization of depth completion.
-
Any to Full: Prompting Depth Anything for Depth Completion in One Stage
Any2Full reformulates depth completion as one-stage scale-prompting of a pretrained monocular depth estimator, yielding domain-general, pattern-agnostic dense metric depth with lower error and higher speed than prior methods.
-
QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
Query-based 4D supervision with a contractive BEV representation sets a new state of the art for self-supervised 3D semantic occupancy from cameras.
-
Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion
A dual-stream BEV architecture that separates instance and scene class queries achieves state-of-the-art mIoU of 17.35 on SemanticKITTI and 20.55 on SSCBench-KITTI-360.
-
Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion
SceneDINO performs semantic scene completion from a single image in a fully unsupervised way by lifting self-supervised DINO features into a 3D feature field trained with multi-view consistency.
-
VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
A training-only Gaussian splatting loss, which renders predicted 3D semantics and motion into 2D camera views, improves semantic occupancy and scene flow prediction across several camera-based models.
-
GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy Prediction
GTAD combines an in-model latent denoising network with global temporal interaction to improve camera-based 3D semantic occupancy prediction, reporting 40.76 mIoU on Occ3D-nuScenes at 12 epochs.
-
GaussianFusionOcc: A Seamless Sensor Fusion Approach for 3D Occupancy Prediction Using 3D Gaussians
GaussianFusionOcc fuses camera, LiDAR, and radar features through deformable attention to refine semantic 3D Gaussians, improving 3D occupancy prediction on nuScenes while lowering memory use and latency.
-
SHTOcc: Effective 3D Occupancy Prediction with Sparse Head and Tail Voxels
SHTOcc combines attention-based sparse voxel selection with decoupled classifier retraining for 3D occupancy prediction, reporting efficiency gains and small, partly inconsistent accuracy improvements.
Discussion (0). Sign in to comment.