REVIEW 3 cited by
FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multi-sensor modal fusion has demonstrated strong advantages in 3D object detection tasks. However, existing methods that fuse multi-modal features require transforming features into the bird's eye view space and may lose certain information on Z-axis, thus leading to inferior performance. To this end, we propose a novel end-to-end multi-modal fusion transformer-based framework, dubbed FusionFormer, that incorporates deformable attention and residual structures within the fusion encoding module. Specifically, by developing a uniform sampling strategy, our method can easily sample from 2D image and 3D voxel features spontaneously, thus exploiting flexible adaptability and avoiding explicit transformation to the bird's eye view space during the feature concatenation process. We further implement a residual structure in our feature encoder to ensure the model's robustness in case of missing an input modality. Through extensive experiments on a popular autonomous driving benchmark dataset, nuScenes, our method achieves state-of-the-art single model performance of 72.6% mAP and 75.1% NDS in the 3D object detection task without test time augmentation.
Forward citations
Cited by 3 Pith papers
-
LiDAR-based End-to-end Temporal Perception for Vehicle-Infrastructure Cooperation
LET-VIC is an end-to-end lidar framework for vehicle-infrastructure cooperative detection and tracking that fuses temporal and multi-view features and learns to compensate calibration errors, outperforming the tested ...
-
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
EVT achieves state-of-the-art 75.3% NDS on the nuScenes test set by using LiDAR-guided adaptive sampling and projection for image-to-BEV transformation, along with a geometry-aware transformer decoder.
-
Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey
A safety-critical survey that organizes BEV perception into single-modality, multimodal, and collaborative stages and consolidates robustness evidence that multimodal fusion degrades far less than single-modality perc...
Discussion (0). Continue with ORCID to comment.