Pith. sign in

REVIEW 16 cited by

FusionAD: Multi-modality Fusion for Prediction and Planning Tasks of Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01006 v4 pith:EHBJXQVH submitted 2023-08-02 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords predictionfusionmulti-modalityfusionadperceptionplanningtasksautonomous
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Building a multi-modality multi-task neural network toward accurate and robust performance is a de-facto standard in perception task of autonomous driving. However, leveraging such data from multiple sensors to jointly optimize the prediction and planning tasks remains largely unexplored. In this paper, we present FusionAD, to the best of our knowledge, the first unified framework that fuse the information from two most critical sensors, camera and LiDAR, goes beyond perception task. Concretely, we first build a transformer based multi-modality fusion network to effectively produce fusion based features. In constrast to camera-based end-to-end method UniAD, we then establish a fusion aided modality-aware prediction and status-aware planning modules, dubbed FMSPnP that take advantages of multi-modality features. We conduct extensive experiments on commonly used benchmark nuScenes dataset, our FusionAD achieves state-of-the-art performance and surpassing baselines on average 15% on perception tasks like detection and tracking, 10% on occupancy prediction accuracy, reducing prediction error from 0.708 to 0.389 in ADE score and reduces the collision rate from 0.31% to only 0.12%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Fusing collaborative sensor data at detection or tracking level beats fusing predicted trajectories, and a compressed LiDAR-sharing prototype improves forecasting over a single vehicle.

  2. PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Decoupling sensor-agnostic 2D trajectory planning from deterministic 3D lifting, plus dense GRPO rewards on perception-to-planning, yields competitive open- and closed-loop driving VLA results.

  3. To New Beginnings: A Survey of Unified Perception in Autonomous Vehicle Software

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A taxonomy that classifies unified perception methods in autonomous driving into Early, Late, and Full Unified Perception based on task integration, tracking formulation, and representation flow.

  4. MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A camera-LiDAR 3D detector built around a hybrid local-global Mamba block with height-fidelity LiDAR encoding reports 75.0 NDS on nuScenes validation, outperforming prior transformer-based fusion methods.

  5. Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving

    cs.RO 2025-06 reject novelty 6.0 of 10

    R2SE refines pretrained end-to-end driving policies on hard cases via residual LoRA reinforcement learning and switches between specialist and generalist policies using GPD-based uncertainty.

  6. Int2Planner: An Intention-based Multi-modal Motion Planner for Integrated Prediction and Planning

    cs.RO 2025-01 conditional novelty 6.0 of 10

    Int2Planner samples intention points from the ego vehicle's route path to generate multi-modal planning trajectories, improving integrated prediction and planning on nuPlan and a private dataset.

  7. Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A point transformer that combines perfect spatial hashing with FlashAttention to align geometric neighborhoods with GPU memory tiles achieves 2.25x faster inference and better semantic segmentation than PTv3.

  8. Doe-1: Closed-Loop Autonomous Driving with Large World Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.

  9. LiDAR-based End-to-end Temporal Perception for Vehicle-Infrastructure Cooperation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LET-VIC is an end-to-end lidar framework for vehicle-infrastructure cooperative detection and tracking that fuses temporal and multi-view features and learns to compensate calibration errors, outperforming the tested ...

  10. LADY: Linear Attention for Autonomous Driving Efficiency without Transformers

    cs.AI 2025-12 conditional novelty 5.0 of 10

    LADY shows that an end-to-end driving model using only linear attention can match transformer-based planners on NAVSIM/Bench2Drive while fusing arbitrary-length historical sensor frames at constant per-frame cost.

  11. CogAD: Cognitive-Hierarchy Guided End-to-End Autonomous Driving

    cs.RO 2025-05 conditional novelty 5.0 of 10

    CogAD reports state-of-the-art open-loop and closed-loop planning results by combining hierarchical scene-to-instance perception with intent-to-trajectory planning and dual-level uncertainty.

  12. GaussianAD: Gaussian-Centric End-to-End Autonomous Driving

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GaussianAD uses sparse 3D semantic Gaussians as the intermediate representation for camera-only end-to-end driving, adding Gaussian flow prediction and future-scene supervision to achieve strong open-loop planning res...

  13. LeAD: The LLM Enhanced Planning System Converged with End-to-end Autonomous Driving

    cs.RO 2025-07 conditional novelty 4.0 of 10

    LeAD adds a low-frequency large-language-model planner that takes over when a high-frequency end-to-end driving model gets stuck, and reports improved CARLA benchmark scores.

  14. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

  15. Joint Perception and Prediction for Autonomous Driving: A Survey

    cs.CV 2024-12 conditional novelty 4.0 of 10

    This survey organizes 55 joint perception and prediction methods for autonomous driving into a taxonomy based on input representation, scene context modeling, and output representation, and compares their reported per...

  16. Monocular Lane Detection Based on Deep Learning: A Survey

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A structured review of 2D and 3D monocular lane detection methods, with a new four-axis taxonomy and unified FPS comparisons.

Pith tools