Introduces a synchronized cross-view urban traffic dataset with drone ground truth for identity matching and monocular BEV localization tasks.
Beverse: Unified perception and prediction in birds-eye-view for vision-centric autonomous driving
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
SceneMiner shows that identity-preserving multi-task fine-tuning removes cross-task interference by zero-initializing new heads and freezing shared-stream parameters, enabling unified BEV scene mining with preserved original heads.
LIE delivers LiDAR-only HD map segmentation via online knowledge distillation that fuses intensity maps, beating the best camera-only model by 8.2% mIoU on nuScenes while adapting quickly to new datasets.
GOLD-BEV learns dense BEV semantic maps including dynamic agents from ego-centric sensors by using synchronized aerial imagery for training supervision and pseudo-label generation.
A visible-occluded dual-decoder network with offline visibility labels reaches SOTA on SemanticKITTI and SSCBench-KITTI360, assuming its monocular framing holds.
VERDI aligns perception, prediction, and planning outputs of end-to-end AD models with VLM-generated text features at training time to embed structured reasoning, yielding up to 11% better l2 distance and 10% higher non-collision rate in closed-loop tests.
VADv2 introduces a probabilistic planning model that discretizes the high-dimensional action space into tokens, interacts them with scene tokens to predict action distributions, and reports SOTA closed-loop results on CARLA Town05 and Bench2Drive.
CTAB bidirectional deformable attention between detection and segmentation BEV branches improves 7-class mIoU by 0.6 at neutral detection on nuScenes radar-camera multi-task learning.
BEVPredFormer uses attention-based temporal processing and 3D camera projection to match or exceed prior methods on nuScenes for BEV instance prediction.
citing papers explorer
-
Cross-View Urban Traffic Dataset: Drone-Supervised Ground Truth for Monocular Bird's-Eye View Localization
Introduces a synchronized cross-view urban traffic dataset with drone ground truth for identity matching and monocular BEV localization tasks.
-
SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining
SceneMiner shows that identity-preserving multi-task fine-tuning removes cross-task interference by zero-initializing new heads and freezing shared-stream parameters, enabling unified BEV scene mining with preserved original heads.
-
LIE: LiDAR-only HD Map Construction with Intensity Enhancement via Online Knowledge Distillation
LIE delivers LiDAR-only HD map segmentation via online knowledge distillation that fuses intensity maps, beating the best camera-only model by 8.2% mIoU on nuScenes while adapting quickly to new datasets.
-
GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes
GOLD-BEV learns dense BEV semantic maps including dynamic agents from ego-centric sensors by using synchronized aerial imagery for training supervision and pseudo-label generation.
-
VOIC: Visible-Occluded Integrated Guidance for 3D Semantic Scene Completion
A visible-occluded dual-decoder network with offline visibility labels reaches SOTA on SemanticKITTI and SSCBench-KITTI360, assuming its monocular framing holds.
-
VERDI: VLM-Embedded Reasoning for Autonomous Driving
VERDI aligns perception, prediction, and planning outputs of end-to-end AD models with VLM-generated text features at training time to embed structured reasoning, yielding up to 11% better l2 distance and 10% higher non-collision rate in closed-loop tests.
-
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
VADv2 introduces a probabilistic planning model that discretizes the high-dimensional action space into tokens, interacts them with scene tokens to predict action distributions, and reports SOTA closed-loop results on CARLA Town05 and Bench2Drive.
-
Radar-Camera BEV Multi-Task Learning with Cross-Task Attention Bridge for Joint 3D Detection and Segmentation
CTAB bidirectional deformable attention between detection and segmentation BEV branches improves 7-class mIoU by 0.6 at neutral detection on nuScenes radar-camera multi-task learning.
-
BEVPredFormer: Spatio-temporal Attention for BEV Instance Prediction in Autonomous Driving
BEVPredFormer uses attention-based temporal processing and 3D camera projection to match or exceed prior methods on nuScenes for BEV instance prediction.
- Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey