REVIEW 16 cited by
FusionAD: Multi-modality Fusion for Prediction and Planning Tasks of Autonomous Driving
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Building a multi-modality multi-task neural network toward accurate and robust performance is a de-facto standard in perception task of autonomous driving. However, leveraging such data from multiple sensors to jointly optimize the prediction and planning tasks remains largely unexplored. In this paper, we present FusionAD, to the best of our knowledge, the first unified framework that fuse the information from two most critical sensors, camera and LiDAR, goes beyond perception task. Concretely, we first build a transformer based multi-modality fusion network to effectively produce fusion based features. In constrast to camera-based end-to-end method UniAD, we then establish a fusion aided modality-aware prediction and status-aware planning modules, dubbed FMSPnP that take advantages of multi-modality features. We conduct extensive experiments on commonly used benchmark nuScenes dataset, our FusionAD achieves state-of-the-art performance and surpassing baselines on average 15% on perception tasks like detection and tracking, 10% on occupancy prediction accuracy, reducing prediction error from 0.708 to 0.389 in ADE score and reduces the collision rate from 0.31% to only 0.12%.
Forward citations
Cited by 16 Pith papers
-
Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives
Fusing collaborative sensor data at detection or tracking level beats fusing predicted trajectories, and a compressed LiDAR-sharing prototype improves forecasting over a single vehicle.
-
PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving
Decoupling sensor-agnostic 2D trajectory planning from deterministic 3D lifting, plus dense GRPO rewards on perception-to-planning, yields competitive open- and closed-loop driving VLA results.
-
To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software
A taxonomy that classifies unified perception methods in autonomous driving into Early, Late, and Full Unified Perception based on task integration, tracking formulation, and representation flow.
-
MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection
A camera-LiDAR 3D detector built around a hybrid local-global Mamba block with height-fidelity LiDAR encoding reports 75.0 NDS on nuScenes validation, outperforming prior transformer-based fusion methods.
-
Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving
R2SE refines pretrained end-to-end driving policies on hard cases via residual LoRA reinforcement learning and switches between specialist and generalist policies using GPD-based uncertainty.
-
Int2Planner: An Intention-based Multi-modal Motion Planner for Integrated Prediction and Planning
Int2Planner samples intention points from the ego vehicle's route path to generate multi-modal planning trajectories, improving integrated prediction and planning on nuPlan and a private dataset.
-
Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality
A point transformer that combines perfect spatial hashing with FlashAttention to align geometric neighborhoods with GPU memory tiles achieves 2.25x faster inference and better semantic segmentation than PTv3.
-
Doe-1: Closed-Loop Autonomous Driving with Large World Model
Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.
-
LiDAR-based End-to-end Temporal Perception for Vehicle-Infrastructure Cooperation
LET-VIC is an end-to-end lidar framework for vehicle-infrastructure cooperative detection and tracking that fuses temporal and multi-view features and learns to compensate calibration errors, outperforming the tested ...
-
LADY: Linear Attention for Autonomous Driving Efficiency without Transformers
LADY shows that an end-to-end driving model using only linear attention can match transformer-based planners on NAVSIM/Bench2Drive while fusing arbitrary-length historical sensor frames at constant per-frame cost.
-
CogAD: Cognitive-Hierarchy Guided End-to-End Autonomous Driving
CogAD reports state-of-the-art open-loop and closed-loop planning results by combining hierarchical scene-to-instance perception with intent-to-trajectory planning and dual-level uncertainty.
-
GaussianAD: Gaussian-Centric End-to-End Autonomous Driving
GaussianAD uses sparse 3D semantic Gaussians as the intermediate representation for camera-only end-to-end driving, adding Gaussian flow prediction and future-scene supervision to achieve strong open-loop planning res...
-
LeAD: The LLM Enhanced Planning System Converged with End-to-end Autonomous Driving
LeAD adds a low-frequency large-language-model planner that takes over when a high-frequency end-to-end driving model gets stuck, and reports improved CARLA benchmark scores.
-
A Survey on Vision-Language-Action Models for Autonomous Driving
A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.
-
Joint Perception and Prediction for Autonomous Driving: A Survey
This survey organizes 55 joint perception and prediction methods for autonomous driving into a taxonomy based on input representation, scene context modeling, and output representation, and compares their reported per...
-
Monocular Lane Detection Based on Deep Learning: A Survey
A structured review of 2D and 3D monocular lane detection methods, with a new four-axis taxonomy and unified FPS comparisons.
Discussion (0). Continue with ORCID to comment.