REVIEW 7 cited by
BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-sensor fusion is essential for an accurate and reliable autonomous driving system. Recent approaches are based on point-level fusion: augmenting the LiDAR point cloud with camera features. However, the camera-to-LiDAR projection throws away the semantic density of camera features, hindering the effectiveness of such methods, especially for semantic-oriented tasks (such as 3D scene segmentation). In this paper, we break this deeply-rooted convention with BEVFusion, an efficient and generic multi-task multi-sensor fusion framework. It unifies multi-modal features in the shared bird's-eye view (BEV) representation space, which nicely preserves both geometric and semantic information. To achieve this, we diagnose and lift key efficiency bottlenecks in the view transformation with optimized BEV pooling, reducing latency by more than 40x. BEVFusion is fundamentally task-agnostic and seamlessly supports different 3D perception tasks with almost no architectural changes. It establishes the new state of the art on nuScenes, achieving 1.3% higher mAP and NDS on 3D object detection and 13.6% higher mIoU on BEV map segmentation, with 1.9x lower computation cost. Code to reproduce our results is available at https://github.com/mit-han-lab/bevfusion.
Forward citations
Cited by 7 Pith papers
-
DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR
A multi-sensor driving dataset (stereo event-RGB-thermal, 4D radar, dual LiDAR) under diverse weather and lighting, with 2D/3D benchmarks and a fusion method that improves 3D detection robustness.
-
SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction
SparseOcc++ decouples geometry completion (via orthogonal SCF regression on sparse anchors) from semantics, improving IoU 2.3 points and running 3.9 imes faster than SparseOcc on nuScenes while 5.9 imes faster than Oc...
-
DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception
Learning which MIMO radar receivers to activate jointly with camera–LiDAR fusion lets fewer receivers match or exceed full-array 3D detection on RADIal, with the best budget depending on the sensor stack.
-
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather
Weakly supervised per-pixel reliability gating of camera features improves LiDAR–4D RADAR 3D detection under adverse weather by up to +6.5 AP_BEV / +7.4 AP_3D on K-Radar.
-
CLIFE: Camera-LiDAR Fusion Framework for Edge-Deployable Roadside VRU Perception
An edge-deployed camera–LiDAR late-fusion system with targetless online calibration achieves real-time VRU tracking on a single Jetson, but its robustness claims are only partially supported by the experiments.
-
Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
A humanoid-specific multimodal occupancy perception system with a new dataset, sensor layout, and a fusion network that claims state-of-the-art results on its own benchmark.
-
Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices
A single image-pretrained VQ-VAE encoder, applied to spectrograms of six physiological signals, matches a modality-specific fusion baseline on WESAD stress classification while using less compute and memory.
Discussion (0). Sign in to comment.