REVIEW 6 cited by
BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. Owing to its low cost and high efficiency, multi-view 3D object detection has demonstrated promising application prospects. However, accurately detecting objects through perspective views is extremely difficult due to the lack of depth information. Current approaches tend to adopt heavy backbones for image encoders, making them inapplicable for real-world deployment. Different from the images, LiDAR points are superior in providing spatial cues, resulting in highly precise localization. In this paper, we explore the incorporation of LiDAR-based detectors for multi-view 3D object detection. Instead of directly training a depth prediction network, we unify the image and LiDAR features in the Bird-Eye-View (BEV) space and adaptively transfer knowledge across non-homogenous representations in a teacher-student paradigm. To this end, we propose \textbf{BEVDistill}, a cross-modal BEV knowledge distillation (KD) framework for multi-view 3D object detection. Extensive experiments demonstrate that the proposed method outperforms current KD approaches on a highly-competitive baseline, BEVFormer, without introducing any extra cost in the inference phase. Notably, our best model achieves 59.4 NDS on the nuScenes test leaderboard, achieving new state-of-the-art in comparison with various image-based detectors. Code will be available at https://github.com/zehuichen123/BEVDistill.
Forward citations
Cited by 6 Pith papers
-
SCKD: Semi-Supervised Cross-Modality Knowledge Distillation for 4D Radar Object Detection
SCKD uses a LiDAR-radar fusion teacher and semi-supervised output distillation to train a radar-only student that outperforms prior radar-only methods on the VoD and ZJUODset benchmarks.
-
PromptDet: A Lightweight 3D Object Detection Framework with LiDAR Prompts
PromptDet is a single-stage framework whose LiDAR prompter improves both fusion and camera-only 3D detection on nuScenes with under 2% additional parameters.
-
DSRC: Learning Density-insensitive and Semantic-aware Collaborative Representation against Corruptions
DSRC combines distillation and point cloud reconstruction to outperform prior collaborative perception models on clean and six simulated corruption settings on two datasets.
-
Self-Supervised Pre-training with Combined Datasets for 3D Perception in Autonomous Driving
Pre-training on combined unlabeled NuScenes, Lyft, and ONCE data with BEV contrastive learning, image MAE, and dataset prompts improves downstream 3D perception tasks.
-
TiGDistill-BEV: Multi-view BEV 3D Object Detection via Target Inner-Geometry Learning Distillation
A LiDAR-to-camera distillation method that supervises relative depth inside object foregrounds and distills BEV feature relationships to boost camera-only 3D object detection.
-
Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities
Foundation-model perception for autonomous driving is surveyed through four capability lenses: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding.
Discussion (0). Continue with ORCID to comment.