CF-SSC predicts pseudo-future frames from past and current monocular images and fuses them in 3D, achieving state-of-the-art semantic scene completion on SemanticKITTI and SSCBench-KITTI-360.
Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Camera-based 3D semantic scene completion (SSC) is pivotal for predicting complicated 3D layouts with limited 2D image observations. The existing mainstream solutions generally leverage temporal information by roughly stacking history frames to supplement the current frame, such straightforward temporal modeling inevitably diminishes valid clues and increases learning difficulty. To address this problem, we present HTCL, a novel Hierarchical Temporal Context Learning paradigm for improving camera-based semantic scene completion. The primary innovation of this work involves decomposing temporal context learning into two hierarchical steps: (a) cross-frame affinity measurement and (b) affinity-based dynamic refinement. Firstly, to separate critical relevant context from redundant information, we introduce the pattern affinity with scale-aware isolation and multiple independent learners for fine-grained contextual correspondence modeling. Subsequently, to dynamically compensate for incomplete observations, we adaptively refine the feature sampling locations based on initially identified locations with high affinity and their neighboring relevant regions. Our method ranks $1^{st}$ on the SemanticKITTI benchmark and even surpasses LiDAR-based methods in terms of mIoU on the OpenOccupancy benchmark. Our code is available on https://github.com/Arlo0o/HTCL.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
One Step Closer: Creating the Future to Boost Monocular Semantic Scene Completion
CF-SSC predicts pseudo-future frames from past and current monocular images and fuses them in 3D, achieving state-of-the-art semantic scene completion on SemanticKITTI and SSCBench-KITTI-360.