REVIEW 23 cited by
BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection
read the original abstract
Single frame data contains finite information which limits the performance of the existing vision-based multi-camera 3D object detection paradigms. For fundamentally pushing the performance boundary in this area, a novel paradigm dubbed BEVDet4D is proposed to lift the scalable BEVDet paradigm from the spatial-only 3D space to the spatial-temporal 4D space. We upgrade the naive BEVDet framework with a few modifications just for fusing the feature from the previous frame with the corresponding one in the current frame. In this way, with negligible additional computing budget, we enable BEVDet4D to access the temporal cues by querying and comparing the two candidate features. Beyond this, we simplify the task of velocity prediction by removing the factors of ego-motion and time in the learning target. As a result, BEVDet4D with robust generalization performance reduces the velocity error by up to -62.9%. This makes the vision-based methods, for the first time, become comparable with those relied on LiDAR or radar in this aspect. On challenge benchmark nuScenes, we report a new record of 54.5% NDS with the high-performance configuration dubbed BEVDet4D-Base, which surpasses the previous leading method BEVDet-Base by +7.3% NDS. The source code is publicly available for further research at https://github.com/HuangJunJie2017/BEVDet .
Forward citations
Cited by 23 Pith papers
-
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction
Factorized Dense Routing approximates unconstrained 2D-to-3D feature mixing by hierarchical tensor contractions, yielding global-context occupancy prediction that remains robust without camera extrinsics.
-
GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception
GaussianFusion presents a 3D Gaussian-based framework that unifies multi-modal features in continuous space for 3D object detection and semantic occupancy, reporting gains over BEVFusion and GaussFormer on nuScenes.
-
Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking
Introduces a streaming Gaussian encoder maintaining persistent volumetric representations via ego-motion compensation and confidence-guided updates for improved 4D panoptic occupancy tracking from cameras.
-
Distortion-Aware PETR for BEV Object Detection with Mixed Pinhole-Fisheye Cameras
DAPETR adds two learned adaptive modules to PETR for superior fisheye BEV detection on converted KITTI-360 data, outperforming PolarPETR while revealing negative interaction when both are combined.
-
Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction
Cross-View Supervision transfers geometric and topological priors from ego-aligned overhead views into camera-based BEV encoders via shared feature alignment, yielding +3.9 mAP and +9.9 mAP gains on nuScenes with 44% ...
-
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
Block-wise Modal Joint Attention over image, LiDAR, and diffusion action tokens yields 88.9 PDMS / 88.4 EPDMS on NAVSIM without anchors or auxiliary supervision.
-
RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection
RECO adds learnable near/far 6-DoF pose corrections to roadside BEV detectors, smoothly blended by a sigmoid gate, improving 3D detection under camera extrinsic jitter and drift.
-
Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction
Cross-View Supervision transfers geometric and topological priors from ego-aligned overhead perspectives into camera-based BEV encoders via feature-space alignment, yielding up to 44% relative mAP gains at long range ...
-
SimPB++: Simultaneously Detecting 2D and 3D Objects from Multiple Cameras
SimPB++ unifies multi-view 2D perspective and 3D BEV object detection in one model via an interactive hybrid decoder, reporting state-of-the-art results on nuScenes and long-range detection up to 150 m on Argoverse2.
-
CAM3DNet: Comprehensively mining the multi-scale features for 3D Object Detection with Multi-View Cameras
CAM3DNet outperforms prior camera-based 3D detectors on nuScenes, Waymo and Argoverse by using three new modules to better mine multi-scale spatiotemporal features from 2D queries and pyramid maps.
-
Kerr-Schild Double Copy of the Randall-Sundrum Black String
Kerr-Schild double copy of the RS II black string produces a sourceless Maxwell single copy and a warp-induced massive scalar zeroth copy, with an alternative splitting giving inequivalent gauge and scalar fields.
-
TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving
TaCarla releases 2.85M CARLA Leaderboard 2.0 frames with nuScenes-style sensors, multi-task annotations, planning baselines, and a text-based rarity score.
-
Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking
Latent Gaussian Splatting (LaGS) replaces dense voxel feature encoders with sparse feature-bearing Gaussians and achieves state-of-the-art 4D panoptic occupancy tracking on nuScenes and Waymo.
-
Semantic Causality-Aware Vision-Based 3D Occupancy Prediction
A class-conditional gradient loss (Causal Loss) plus channel-grouped lifting, learnable camera offsets, and normalized convolution raises Occ3D mIoU by 1.2/0.8 points and cuts the camera-noise mIoU drop from 32% to 7%.
-
RQR3D: Reparametrizing the regression targets for BEV-based 3D object detection
RQR3D reparametrizes oriented bounding box regression in BEV 3D detection as regressing a horizontal box plus corner offsets and achieves SOTA camera-radar performance on nuScenes with 67.5 NDS and 59.7 mAP.
-
SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
SemLT3D introduces semantic-guided expert distillation with a language MoE module and CLIP projection to enrich features for long-tailed classes in camera-only 3D detection.
-
Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
SEPatch3D accelerates ViT-based 3D object detectors up to 57% faster than StreamPETR via dynamic patch sizing and cross-granularity enhancement while keeping comparable accuracy on nuScenes and Argoverse 2.
-
Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning
GameAD models autonomous driving as a risk-prioritized game among agents via Risk-Aware Topology Anchoring, Minimax Risk-Aware Sparse Attention and related components, yielding safer trajectories than prior end-to-end...
-
Kerr-Schild Double Copy of the Randall-Sundrum Black String
Kerr-Schild double copy of the RSII black string gives a holographic-coordinate-independent sourceless single-copy gauge field and a zeroth copy with warp-induced mass m²=12/l², while an alternative split is inequivalent.
-
Multi-Modal Sensor Fusion using Hybrid Attention for Autonomous Driving
MMF-BEV fuses camera and radar branches with deformable self- and cross-attention, outperforming unimodal baselines on the VoD 4D radar dataset through a two-stage training process.
-
BEVPredFormer: Spatio-temporal Attention for BEV Instance Prediction in Autonomous Driving
BEVPredFormer uses attention-based temporal processing and 3D camera projection to match or exceed prior methods on nuScenes for BEV instance prediction.
-
Fast-BEV++: Fast by Algorithm, Deployable by Design
Fast-BEV++ achieves at least 3x speedup over Fast-BEV, a new SOTA of 0.488 NDS on nuScenes 3D detection, and over 134 FPS inference by redesigning the core transformation pipeline and adding a learnable depth module.
-
SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation
SliceSemOcc improves 3D semantic occupancy prediction by slicing voxel features into global and local height bands and applying per-height channel attention, yielding modest mIoU gains on nuScenes benchmarks.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.