Pith. sign in

hub Canonical reference

BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Canonical reference. 71% of citing Pith papers cite this work as background.

39 Pith papers citing it
Background 71% of classified citations
abstract

Autonomous driving perceives its surroundings for decision making, which is one of the most complex scenarios in visual perception. The success of paradigm innovation in solving the 2D object detection task inspires us to seek an elegant, feasible, and scalable paradigm for fundamentally pushing the performance boundary in this area. To this end, we contribute the BEVDet paradigm in this paper. BEVDet performs 3D object detection in Bird-Eye-View (BEV), where most target values are defined and route planning can be handily performed. We merely reuse existing modules to build its framework but substantially develop its performance by constructing an exclusive data augmentation strategy and upgrading the Non-Maximum Suppression strategy. In the experiment, BEVDet offers an excellent trade-off between accuracy and time-efficiency. As a fast version, BEVDet-Tiny scores 31.2% mAP and 39.2% NDS on the nuScenes val set. It is comparable with FCOS3D, but requires just 11% computational budget of 215.3 GFLOPs and runs 9.2 times faster at 15.6 FPS. Another high-precision version dubbed BEVDet-Base scores 39.3% mAP and 47.2% NDS, significantly exceeding all published results. With a comparable inference speed, it surpasses FCOS3D by a large margin of +9.8% mAP and +10.0% NDS. The source code is publicly available for further research at https://github.com/HuangJunJie2017/BEVDet .

hub tools

citation-role summary

background 5 baseline 1 method 1

citation-polarity summary

years

2026 33 2025 6

representative citing papers

TRIG: Trajectory-Rig Decoupled Metric Geometry Learning

cs.CV · 2026-07-07 · unverdicted · novelty 6.0

TRIG factorizes multi-camera poses into ego-trajectory and static rig geometry, with decoupled supervision and sparse temporal-spatial attention, claiming SOTA metric depth, pose, and 3D reconstruction on five driving benchmarks.

Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints

cs.CV · 2026-06-23 · conditional · novelty 6.0

OVBEVSeg produces open-vocabulary bird's-eye-view semantic maps on nuScenes by projecting CLIP labels through 3D detections, constraining Gaussian splats with BEV occupancy, and distilling the geometry into a real-time student (15.3 mIoU novel, 0% novel GT).

DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale

cs.CV · 2026-04-01 · unverdicted · novelty 6.0

DVGT-2 is a streaming vision-geometry-action model that jointly reconstructs dense 3D geometry and plans trajectories online, achieving better reconstruction than prior batch methods while transferring directly to planning benchmarks without fine-tuning.

citing papers explorer

Showing 39 of 39 citing papers.