Pith. sign in

REVIEW 1 cited by

RadOcc: Learning Cross-Modality Occupancy Knowledge through Rendering Assisted Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11829 v1 pith:LIVBT5JX submitted 2023-12-19 cs.CV

classification cs.CV
keywords occupancypredictionconsistencydistillationproposedrenderingassisteddepth
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D occupancy prediction is an emerging task that aims to estimate the occupancy states and semantics of 3D scenes using multi-view images. However, image-based scene perception encounters significant challenges in achieving accurate prediction due to the absence of geometric priors. In this paper, we address this issue by exploring cross-modal knowledge distillation in this task, i.e., we leverage a stronger multi-modal model to guide the visual model during training. In practice, we observe that directly applying features or logits alignment, proposed and widely used in bird's-eyeview (BEV) perception, does not yield satisfactory results. To overcome this problem, we introduce RadOcc, a Rendering assisted distillation paradigm for 3D Occupancy prediction. By employing differentiable volume rendering, we generate depth and semantic maps in perspective views and propose two novel consistency criteria between the rendered outputs of teacher and student models. Specifically, the depth consistency loss aligns the termination distributions of the rendered rays, while the semantic consistency loss mimics the intra-segment similarity guided by vision foundation models (VLMs). Experimental results on the nuScenes dataset demonstrate the effectiveness of our proposed method in improving various 3D occupancy prediction approaches, e.g., our proposed methodology enhances our baseline by 2.2% in the metric of mIoU and achieves 50% in Occ3D benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An end-to-end, non-autoregressive 3D occupancy world model warps dynamic voxels via predicted flow, moves static voxels by pose, and uses image-based rendering supervision, achieving state-of-the-art results on three ...

Pith tools