REVIEW 5 cited by
GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Integrating LiDAR and camera information into Bird's-Eye-View (BEV) representation has emerged as a crucial aspect of 3D object detection in autonomous driving. However, existing methods are susceptible to the inaccurate calibration relationship between LiDAR and the camera sensor. Such inaccuracies result in errors in depth estimation for the camera branch, ultimately causing misalignment between LiDAR and camera BEV features. In this work, we propose a robust fusion framework called Graph BEV. Addressing errors caused by inaccurate point cloud projection, we introduce a Local Align module that employs neighbor-aware depth features via Graph matching. Additionally, we propose a Global Align module to rectify the misalignment between LiDAR and camera BEV features. Our Graph BEV framework achieves state-of-the-art performance, with an mAP of 70.1\%, surpassing BEV Fusion by 1.6\% on the nuscenes validation set. Importantly, our Graph BEV outperforms BEV Fusion by 8.3\% under conditions with misalignment noise.
Forward citations
Cited by 5 Pith papers
-
SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction
SDGOCC improves multimodal 3D occupancy prediction by using LiDAR depth and semantic masks to guide camera-to-BEV transformation, achieving state-of-the-art mIoU on Occ3D-nuScenes.
-
RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection
RCTrans achieves new state-of-the-art radar-camera 3D detection on nuScenes by densifying radar BEV features and using a pruning sequential decoder for query-based fusion.
-
TiGDistill-BEV: Multi-view BEV 3D Object Detection via Target Inner-Geometry Learning Distillation
A LiDAR-to-camera distillation method that supervises relative depth inside object foregrounds and distills BEV feature relationships to boost camera-only 3D object detection.
-
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
TimeAlign uses Swin-LSTM prediction and camera-guided combination of past and observed LiDAR BEV features to partially recover 3D detection accuracy under LiDAR time lag.
- CoreNet: Conflict Resolution Network for Point-Pixel Misalignment and Sub-Task Suppression of 3D LiDAR-Camera Object Detection
Discussion (0). Continue with ORCID to comment.