Pith. sign in

REVIEW 11 cited by

BEVerse: Unified Perception and Prediction in Birds-Eye-View for Vision-Centric Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.09743 v1 pith:DXKHRFAC submitted 2022-05-19 cs.CV

classification cs.CV
keywords beversepredictionautonomousbirds-eye-viewconstructiondecodersdetectiondifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present BEVerse, a unified framework for 3D perception and prediction based on multi-camera systems. Unlike existing studies focusing on the improvement of single-task approaches, BEVerse features in producing spatio-temporal Birds-Eye-View (BEV) representations from multi-camera videos and jointly reasoning about multiple tasks for vision-centric autonomous driving. Specifically, BEVerse first performs shared feature extraction and lifting to generate 4D BEV representations from multi-timestamp and multi-view images. After the ego-motion alignment, the spatio-temporal encoder is utilized for further feature extraction in BEV. Finally, multiple task decoders are attached for joint reasoning and prediction. Within the decoders, we propose the grid sampler to generate BEV features with different ranges and granularities for different tasks. Also, we design the method of iterative flow for memory-efficient future prediction. We show that the temporal information improves 3D object detection and semantic map construction, while the multi-task learning can implicitly benefit motion prediction. With extensive experiments on the nuScenes dataset, we show that the multi-task BEVerse outperforms existing single-task methods on 3D object detection, semantic map construction, and motion prediction. Compared with the sequential paradigm, BEVerse also favors in significantly improved efficiency. The code and trained models will be released at https://github.com/zhangyp15/BEVerse.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Class-Incremental Motion Forecasting

    cs.CV 2026-03 conditional novelty 7.0 of 10

    OMEN is the first end-to-end class-incremental motion forecaster that retains old-class accuracy via VLM-filtered future-detection pseudo-labels and variance-based sequence replay.

  2. VOIC: Visible-Occluded Integrated Guidance for 3D Semantic Scene Completion

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    VOIC decouples monocular 3D scene completion into visible semantic perception and occluded reasoning via VRLE and a dual-decoder architecture, achieving state-of-the-art results on SemanticKITTI and SSCBench-KITTI360.

  3. BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

    cs.CV 2026-03 conditional novelty 6.5 of 10

    BEV tokens give LLMs stronger cross-view spatial reasoning than multi-view image tokens, and reverse-distilling LLM semantics into BEV encoders measurably improves closed-loop safety-critical driving.

  4. To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A taxonomy that classifies unified perception methods in autonomous driving into Early, Late, and Full Unified Perception based on task integration, tracking formulation, and representation flow.

  5. SafeMap: Robust HD Map Construction from Incomplete Observations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SafeMap improves HD map construction accuracy under missing camera views by reconstructing the missing perspective features with Gaussian-sampled attention and correcting the BEV features through distillation.

  6. TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Auxiliary CLIP-derived BEV semantic supervision during training improves nuScenes end-to-end vehicle instance prediction over a geometric-only baseline, with the semantic head removed at inference.

  7. Radar-Camera BEV Multi-Task Learning with Cross-Task Attention Bridge for Joint 3D Detection and Segmentation

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    CTAB exchanges features between detection and segmentation via multi-scale deformable attention in BEV space, yielding segmentation gains on 7 nuScenes classes at neutral detection cost.

  8. Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A safety-critical survey that organizes BEV perception into single-modality, multimodal, and collaborative stages and consolidates robustness evidence that multimodal fusion degrades far less than single-modality perc...

  9. Rethink 3D Object Detection from Physical World

    cs.RO 2025-06 conditional novelty 5.0 of 10

    Latency-aware and planning-aware AP metrics re-rank 3D object detectors for autonomous driving, showing that faster, safer models can beat higher-mAP ones.

  10. Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance

    cs.AI 2025-08 reject novelty 4.0 of 10

    A lightweight vision-language model on an edge device fuses roadside hazard alerts with onboard camera views to adjust trajectories, and the authors report a 77% simulated collision reduction over a vision-only baseline.

  11. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

Pith tools