Pith. sign in

REVIEW 3 cited by

BEVFusion4D: Learning LiDAR-Camera Fusion Under Bird's-Eye-View via Cross-Modality Guidance and Temporal Aggregation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.17099 v1 pith:A65BCQ4E submitted 2023-03-30 cs.CV

classification cs.CV
keywords cameraframeworkfusioninformationlidartemporalbevfusion4dbird
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Integrating LiDAR and Camera information into Bird's-Eye-View (BEV) has become an essential topic for 3D object detection in autonomous driving. Existing methods mostly adopt an independent dual-branch framework to generate LiDAR and camera BEV, then perform an adaptive modality fusion. Since point clouds provide more accurate localization and geometry information, they could serve as a reliable spatial prior to acquiring relevant semantic information from the images. Therefore, we design a LiDAR-Guided View Transformer (LGVT) to effectively obtain the camera representation in BEV space and thus benefit the whole dual-branch fusion system. LGVT takes camera BEV as the primitive semantic query, repeatedly leveraging the spatial cue of LiDAR BEV for extracting image features across multiple camera views. Moreover, we extend our framework into the temporal domain with our proposed Temporal Deformable Alignment (TDA) module, which aims to aggregate BEV features from multiple historical frames. Including these two modules, our framework dubbed BEVFusion4D achieves state-of-the-art results in 3D object detection, with 72.0% mAP and 73.5% NDS on the nuScenes validation set, and 73.3% mAP and 74.7% NDS on nuScenes test set, respectively.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Supervised Sparse Sensor Fusion for Long Range Perception

    cs.CV 2025-08 conditional novelty 6.0 of 10

    LRS4Fusion fuses cameras and LiDAR in a fully sparse voxel representation with self-supervised temporal pre-training, achieving 52.61 mAP for detection out to 250 meters.

  2. SAMFusion: Sensor-Adaptive Multimodal Fusion for 3D Object Detection in Adverse Weather

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SAMFusion improves 3D object detection in fog, snow, and night by adaptively fusing RGB camera, LiDAR, gated NIR, and radar features in Bird's Eye View.

  3. Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A safety-critical survey that organizes BEV perception into single-modality, multimodal, and collaborative stages and consolidates robustness evidence that multimodal fusion degrades far less than single-modality perc...

Pith tools