Pith. sign in

REVIEW 23 cited by

BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.17054 v3 pith:LD6PNGFR submitted 2022-03-31 cs.CV

BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection

classification cs.CV
keywords bevdet4dbevdetframeperformancecuesdetectiondubbedmulti-camera
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Single frame data contains finite information which limits the performance of the existing vision-based multi-camera 3D object detection paradigms. For fundamentally pushing the performance boundary in this area, a novel paradigm dubbed BEVDet4D is proposed to lift the scalable BEVDet paradigm from the spatial-only 3D space to the spatial-temporal 4D space. We upgrade the naive BEVDet framework with a few modifications just for fusing the feature from the previous frame with the corresponding one in the current frame. In this way, with negligible additional computing budget, we enable BEVDet4D to access the temporal cues by querying and comparing the two candidate features. Beyond this, we simplify the task of velocity prediction by removing the factors of ego-motion and time in the learning target. As a result, BEVDet4D with robust generalization performance reduces the velocity error by up to -62.9%. This makes the vision-based methods, for the first time, become comparable with those relied on LiDAR or radar in this aspect. On challenge benchmark nuScenes, we report a new record of 54.5% NDS with the high-performance configuration dubbed BEVDet4D-Base, which surpasses the previous leading method BEVDet-Base by +7.3% NDS. The source code is publicly available for further research at https://github.com/HuangJunJie2017/BEVDet .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

    cs.CV 2026-07 conditional novelty 7.0

    Factorized Dense Routing approximates unconstrained 2D-to-3D feature mixing by hierarchical tensor contractions, yielding global-context occupancy prediction that remains robust without camera extrinsics.

  2. GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception

    cs.CV 2026-07 unverdicted novelty 7.0

    GaussianFusion presents a 3D Gaussian-based framework that unifies multi-modal features in continuous space for 3D object detection and semantic occupancy, reporting gains over BEVFusion and GaussFormer on nuScenes.

  3. Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces a streaming Gaussian encoder maintaining persistent volumetric representations via ego-motion compensation and confidence-guided updates for improved 4D panoptic occupancy tracking from cameras.

  4. Distortion-Aware PETR for BEV Object Detection with Mixed Pinhole-Fisheye Cameras

    cs.CV 2026-06 unverdicted novelty 7.0

    DAPETR adds two learned adaptive modules to PETR for superior fisheye BEV detection on converted KITTI-360 data, outperforming PolarPETR while revealing negative interaction when both are combined.

  5. Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction

    cs.CV 2026-05 unverdicted novelty 7.0

    Cross-View Supervision transfers geometric and topological priors from ego-aligned overhead views into camera-based BEV encoders via shared feature alignment, yielding +3.9 mAP and +9.9 mAP gains on nuScenes with 44% ...

  6. MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0

    Block-wise Modal Joint Attention over image, LiDAR, and diffusion action tokens yields 88.9 PDMS / 88.4 EPDMS on NAVSIM without anchors or auxiliary supervision.

  7. RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection

    cs.CV 2026-07 conditional novelty 6.0

    RECO adds learnable near/far 6-DoF pose corrections to roadside BEV detectors, smoothly blended by a sigmoid gate, improving 3D detection under camera extrinsic jitter and drift.

  8. Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction

    cs.CV 2026-05 unverdicted novelty 6.0

    Cross-View Supervision transfers geometric and topological priors from ego-aligned overhead perspectives into camera-based BEV encoders via feature-space alignment, yielding up to 44% relative mAP gains at long range ...

  9. SimPB++: Simultaneously Detecting 2D and 3D Objects from Multiple Cameras

    cs.CV 2026-05 unverdicted novelty 6.0

    SimPB++ unifies multi-view 2D perspective and 3D BEV object detection in one model via an interactive hybrid decoder, reporting state-of-the-art results on nuScenes and long-range detection up to 150 m on Argoverse2.

  10. CAM3DNet: Comprehensively mining the multi-scale features for 3D Object Detection with Multi-View Cameras

    cs.CV 2026-04 unverdicted novelty 6.0

    CAM3DNet outperforms prior camera-based 3D detectors on nuScenes, Waymo and Argoverse by using three new modules to better mine multi-scale spatiotemporal features from 2D queries and pyramid maps.

  11. Kerr-Schild Double Copy of the Randall-Sundrum Black String

    hep-th 2026-04 unverdicted novelty 6.0

    Kerr-Schild double copy of the RS II black string produces a sourceless Maxwell single copy and a warp-induced massive scalar zeroth copy, with an alternative splitting giving inequivalent gauge and scalar fields.

  12. TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

    cs.RO 2026-02 conditional novelty 6.0

    TaCarla releases 2.85M CARLA Leaderboard 2.0 frames with nuScenes-style sensors, multi-task annotations, planning baselines, and a text-based rarity score.

  13. Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking

    cs.CV 2026-02 conditional novelty 6.0

    Latent Gaussian Splatting (LaGS) replaces dense voxel feature encoders with sparse feature-bearing Gaussians and achieves state-of-the-art 4D panoptic occupancy tracking on nuScenes and Waymo.

  14. Semantic Causality-Aware Vision-Based 3D Occupancy Prediction

    cs.CV 2025-09 conditional novelty 6.0

    A class-conditional gradient loss (Causal Loss) plus channel-grouped lifting, learnable camera offsets, and normalized convolution raises Occ3D mIoU by 1.2/0.8 points and cuts the camera-noise mIoU drop from 32% to 7%.

  15. RQR3D: Reparametrizing the regression targets for BEV-based 3D object detection

    cs.CV 2025-05 unverdicted novelty 6.0

    RQR3D reparametrizes oriented bounding box regression in BEV 3D detection as regressing a horizontal box plus corner offsets and achieves SOTA camera-radar performance on nuScenes with 67.5 NDS and 59.7 mAP.

  16. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

    cs.CV 2026-04 unverdicted novelty 5.0

    SemLT3D introduces semantic-guided expert distillation with a language MoE module and CLIP projection to enrich features for long-tailed classes in camera-only 3D detection.

  17. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors

    cs.CV 2026-04 conditional novelty 5.0

    SEPatch3D accelerates ViT-based 3D object detectors up to 57% faster than StreamPETR via dynamic patch sizing and cross-granularity enhancement while keeping comparable accuracy on nuScenes and Argoverse 2.

  18. Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning

    cs.CV 2026-04 unverdicted novelty 5.0

    GameAD models autonomous driving as a risk-prioritized game among agents via Risk-Aware Topology Anchoring, Minimax Risk-Aware Sparse Attention and related components, yielding safer trajectories than prior end-to-end...

  19. Kerr-Schild Double Copy of the Randall-Sundrum Black String

    hep-th 2026-04 unverdicted novelty 5.0

    Kerr-Schild double copy of the RSII black string gives a holographic-coordinate-independent sourceless single-copy gauge field and a zeroth copy with warp-induced mass m²=12/l², while an alternative split is inequivalent.

  20. Multi-Modal Sensor Fusion using Hybrid Attention for Autonomous Driving

    cs.CV 2026-04 unverdicted novelty 5.0

    MMF-BEV fuses camera and radar branches with deformable self- and cross-attention, outperforming unimodal baselines on the VoD 4D radar dataset through a two-stage training process.

  21. BEVPredFormer: Spatio-temporal Attention for BEV Instance Prediction in Autonomous Driving

    cs.CV 2026-04 unverdicted novelty 5.0

    BEVPredFormer uses attention-based temporal processing and 3D camera projection to match or exceed prior methods on nuScenes for BEV instance prediction.

  22. Fast-BEV++: Fast by Algorithm, Deployable by Design

    cs.CV 2025-12 unverdicted novelty 5.0

    Fast-BEV++ achieves at least 3x speedup over Fast-BEV, a new SOTA of 0.488 NDS on nuScenes 3D detection, and over 134 FPS inference by redesigning the core transformation pipeline and adding a learnable depth module.

  23. SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation

    cs.CV 2025-09 conditional novelty 5.0

    SliceSemOcc improves 3D semantic occupancy prediction by slicing voxel features into global and local height bands and applying per-height channel attention, yielding modest mIoU gains on nuScenes benchmarks.