Pith. sign in

REVIEW 18 cited by

DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18607 v1 pith:ILGD3JUL submitted 2024-12-24 cs.CV

DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers

classification cs.CV
keywords planningdrivingmodelingworlddrivinggptactionautoregressivegeneration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual generation but lack the flexibility to incorporate other modalities like action. In contrast, autoregressive transformers have demonstrated exceptional capability in modeling multimodal data. Our work aims to unify both driving model simulation and trajectory planning into a single sequence modeling problem. We introduce a multimodal driving language based on interleaved image and action tokens, and develop DrivingGPT to learn joint world modeling and planning through standard next-token prediction. Our DrivingGPT demonstrates strong performance in both action-conditioned video generation and end-to-end planning, outperforming strong baselines on large-scale nuPlan and NAVSIM benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

    cs.CV 2026-03 conditional novelty 6.5

    BEV tokens give LLMs stronger cross-view spatial reasoning than multi-view image tokens, and reverse-distilling LLM semantics into BEV encoders measurably improves closed-loop safety-critical driving.

  2. G2DP: Diffusion Planning with Spatio-Temporal Grid Guidance

    cs.RO 2026-06 conditional novelty 6.0

    A diffusion-based autonomous driving planner that injects gradients from a spatio-temporal occupancy-and-route cost volume into late denoising steps improves closed-loop safety and progress across three benchmarks.

  3. G2DP: Diffusion Planning with Spatio-Temporal Grid Guidance

    cs.RO 2026-06 unverdicted novelty 6.0

    G2DP adds dense spatio-temporal grid guidance to diffusion-based motion planning, reporting +7.2 reactive score gains on nuPlan and improved collision avoidance in zero-shot transfers.

  4. G2DP: Diffusion Planning with Spatio-Temporal Grid Guidance

    cs.RO 2026-06 unverdicted novelty 6.0

    G2DP constructs a differentiable spatio-temporal cost volume from occupancy and route maps to guide diffusion denoising for collision-free trajectories, reporting SOTA closed-loop scores on nuPlan.

  5. CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning

    cs.RO 2026-05 conditional novelty 6.0

    CLOVER is a closed-loop generator-scorer framework that expands proposal coverage with pseudo-expert trajectories and performs conservative self-distillation to achieve state-of-the-art planning scores on NAVSIM and nuScenes.

  6. OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models

    cs.CV 2026-04 unverdicted novelty 6.0

    OneDrive unifies heterogeneous decoding in a single VLM transformer decoder for end-to-end driving, achieving 0.28 L2 error and 0.18 collision rate on nuScenes plus 86.8 PDMS on NAVSIM.

  7. Human Cognition in Machines: A Unified Perspective of World Models

    cs.RO 2026-04 unverdicted novelty 6.0

    The paper introduces a unified framework for world models that fully incorporates all cognitive functions from Cognitive Architecture Theory, highlights under-researched areas in motivation and meta-cognition, and pro...

  8. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

    cs.CV 2025-12 conditional novelty 6.0

    A five-aspect, 24-metric benchmark, a 26K human-annotated dataset, and an AI evaluator show that today's driving world models cannot simultaneously look real, respect geometry, and behave safely.

  9. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  10. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  11. DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

    cs.CV 2025-10 unverdicted novelty 6.0

    DriveVLA-W0 adds world modeling to predict future images in VLA models, overcoming sparse action supervision and amplifying data scaling laws on NAVSIM benchmarks and a large in-house dataset.

  12. ReSim: Reliable World Simulation for Autonomous Driving

    cs.CV 2025-06 unverdicted novelty 6.0

    ReSim is a controllable video world model trained on heterogeneous real and simulated driving data that achieves higher fidelity and controllability for both expert and non-expert actions, plus a Video2Reward module f...

  13. ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving

    cs.RO 2026-06 unverdicted novelty 5.0

    ASSCG is an RWKV-based adaptive gate trained with SFT and GRPO-style RL that makes Query/Cache/Drop decisions for slow LLM guidance in fast-slow autonomous driving planners, improving scores and cutting latency on nuP...

  14. HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models

    cs.RO 2026-05 unverdicted novelty 5.0

    HEAT uses a trajectory-driven learning paradigm and a world model predicting future latent features from ego actions to enable a single unified end-to-end autonomous driving model to perform well across heterogeneous ...

  15. 3D and 4D World Modeling: A Survey

    cs.CV 2025-09 conditional novelty 5.0

    A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.

  16. DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment

    cs.RO 2025-04 unverdicted novelty 5.0

    DriVerse is a generative model that simulates driving scenes from an image and trajectory using multimodal prompting and motion alignment, achieving better performance on nuScenes and Waymo datasets with minimal training.

  17. DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

    cs.CV 2026-05 unverdicted novelty 4.0

    DeepSight uses parallel latent feature prediction in BEV for long-horizon world modeling and adaptive text reasoning to reach state-of-the-art closed-loop performance on the Bench2drive benchmark.

  18. Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-09 conditional novelty 4.0

    A structured survey of LLM-based trajectory prediction methods, organized into trajectory-language mapping, multimodal fusion, and constraint-based reasoning, with benchmarks, metrics, and future directions.