REVIEW 4 cited by
Fast-BEV: Towards Real-time On-vehicle Bird's-Eye View Perception
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recently, the pure camera-based Bird's-Eye-View (BEV) perception removes expensive Lidar sensors, making it a feasible solution for economical autonomous driving. However, most existing BEV solutions either suffer from modest performance or require considerable resources to execute on-vehicle inference. This paper proposes a simple yet effective framework, termed Fast-BEV, which is capable of performing real-time BEV perception on the on-vehicle chips. Towards this goal, we first empirically find that the BEV representation can be sufficiently powerful without expensive view transformation or depth representation. Starting from M2BEV baseline, we further introduce (1) a strong data augmentation strategy for both image and BEV space to avoid over-fitting (2) a multi-frame feature fusion mechanism to leverage the temporal information (3) an optimized deployment-friendly view transformation to speed up the inference. Through experiments, we show Fast-BEV model family achieves considerable accuracy and efficiency on edge. In particular, our M1 model (R18@256x704) can run over 50FPS on the Tesla T4 platform, with 47.0% NDS on the nuScenes validation set. Our largest model (R101@900x1600) establishes a new state-of-the-art 53.5% NDS on the nuScenes validation set. The code is released at: https://github.com/Sense-GVT/Fast-BEV.
Forward citations
Cited by 4 Pith papers
-
GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy Prediction
GTAD combines an in-model latent denoising network with global temporal interaction to improve camera-based 3D semantic occupancy prediction, reporting 40.76 mIoU on Occ3D-nuScenes at 12 epochs.
-
MambaDETR: Query-based Temporal Modeling using State Space Model for Multi-View 3D Object Detection
MambaDETR applies a Mamba/SSM sequence model to temporal fusion of 3D detection queries, achieving 50.8 mAP on nuScenes val with 8 frames and linear memory scaling.
-
Driver2Map: Imitating Human Driving for Online High-Definition Map Construction
Driver2Map reports state-of-the-art nuScenes accuracy for online HD map construction by fusing camera images, SD maps, and satellite imagery with a pose-weighted BEV fusion and a pretrained map-refinement module.
-
Reflective Teacher: Semi-Supervised Multimodal 3D Object Detection in Bird's-Eye-View via Uncertainty Measure
A semi-supervised teacher-student detector that applies a memory-aware regularizer and uncertainty weighting achieves near-full-supervised 3D detection accuracy with 22 to 25 percent of labels.
Discussion (0). Continue with ORCID to comment.