Pith. sign in

REVIEW 7 cited by

BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.13542 v3 pith:6PMVQS4Q submitted 2022-05-26 cs.CV

classification cs.CV
keywords bevfusionfusionfeaturesmulti-sensorviewbirdcamerahigher
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-sensor fusion is essential for an accurate and reliable autonomous driving system. Recent approaches are based on point-level fusion: augmenting the LiDAR point cloud with camera features. However, the camera-to-LiDAR projection throws away the semantic density of camera features, hindering the effectiveness of such methods, especially for semantic-oriented tasks (such as 3D scene segmentation). In this paper, we break this deeply-rooted convention with BEVFusion, an efficient and generic multi-task multi-sensor fusion framework. It unifies multi-modal features in the shared bird's-eye view (BEV) representation space, which nicely preserves both geometric and semantic information. To achieve this, we diagnose and lift key efficiency bottlenecks in the view transformation with optimized BEV pooling, reducing latency by more than 40x. BEVFusion is fundamentally task-agnostic and seamlessly supports different 3D perception tasks with almost no architectural changes. It establishes the new state of the art on nuScenes, achieving 1.3% higher mAP and NDS on 3D object detection and 13.6% higher mIoU on BEV map segmentation, with 1.9x lower computation cost. Code to reproduce our results is available at https://github.com/mit-han-lab/bevfusion.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR

    cs.CV 2026-04 accept novelty 7.0 of 10

    A multi-sensor driving dataset (stereo event-RGB-thermal, 4D radar, dual LiDAR) under diverse weather and lighting, with 2D/3D benchmarks and a fusion method that improves 3D detection robustness.

  2. SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction

    cs.CV 2026-07 accept novelty 6.5 of 10

    SparseOcc++ decouples geometry completion (via orthogonal SCF regression on sparse anchors) from semantics, improving IoU 2.3 points and running 3.9 imes faster than SparseOcc on nuScenes while 5.9 imes faster than Oc...

  3. DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Learning which MIMO radar receivers to activate jointly with camera–LiDAR fusion lets fewer receivers match or exceed full-array 3D detection on RADIal, with the best budget depending on the sensor stack.

  4. RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Weakly supervised per-pixel reliability gating of camera features improves LiDAR–4D RADAR 3D detection under adverse weather by up to +6.5 AP_BEV / +7.4 AP_3D on K-Radar.

  5. CLIFE: Camera-LiDAR Fusion Framework for Edge-Deployable Roadside VRU Perception

    cs.CV 2026-07 conditional novelty 5.0 of 10

    An edge-deployed camera–LiDAR late-fusion system with targetless online calibration achieves real-time VRU tracking on a single Jetson, but its robustness claims are only partially supported by the experiments.

  6. Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A humanoid-specific multimodal occupancy perception system with a new dataset, sensor layout, and a fusion network that claims state-of-the-art results on its own benchmark.

  7. Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices

    eess.SP 2025-07 conditional novelty 4.0 of 10

    A single image-pretrained VQ-VAE encoder, applied to spectrograms of six physiological signals, matches a modality-specific fusion baseline on WESAD stress classification while using less compute and memory.

Pith tools