Pith. sign in

REVIEW 10 cited by

Doe-1: Closed-Loop Autonomous Driving with Large World Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09627 v1 pith:2IAUZRHS submitted 2024-12-12 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords drivingautonomousplanningtokensdoe-1largeperceptionclosed-loop
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end autonomous driving has received increasing attention due to its potential to learn from large amounts of data. However, most existing methods are still open-loop and suffer from weak scalability, lack of high-order interactions, and inefficient decision-making. In this paper, we explore a closed-loop framework for autonomous driving and propose a large Driving wOrld modEl (Doe-1) for unified perception, prediction, and planning. We formulate autonomous driving as a next-token generation problem and use multi-modal tokens to accomplish different tasks. Specifically, we use free-form texts (i.e., scene descriptions) for perception and generate future predictions directly in the RGB space with image tokens. For planning, we employ a position-aware tokenizer to effectively encode action into discrete tokens. We train a multi-modal transformer to autoregressively generate perception, prediction, and planning tokens in an end-to-end and unified manner. Experiments on the widely used nuScenes dataset demonstrate the effectiveness of Doe-1 in various tasks including visual question-answering, action-conditioned video generation, and motion planning. Code: https://github.com/wzzheng/Doe.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A new benchmark, STSnu, uses 971 verified multiple-choice questions from NuScenes to test driving vision-language models' spatio-temporal reasoning, and shows they lag far behind text-only LLMs given perfect trajectories.

  2. Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features

    cs.RO 2026-08 conditional novelty 6.0 of 10

    An early-exit, quality-guided planner on a Wan2.2-5B video-diffusion backbone reaches 90.8 PDMS on NAVSIM in about 170 ms, without generating future video at deployment.

  3. HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hybrid world model that combines pixel-token prediction with latent prediction beats both pixel-only and latent-only world models on NAVSIM and is more robust to scene noise.

  4. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  5. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  6. Epona: Autoregressive Diffusion World Model for Autonomous Driving

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive diffusion world model jointly generates the next camera frame and a multi-step trajectory, enabling long videos and real-time planning for autonomous driving.

  7. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.

  8. UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A single mask-modulated DiT that co-trains future video and trajectories yields stronger autonomous-driving action generalization and 4.3× faster trajectory-only inference than dual-DiT designs.

  9. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

  10. LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

    cs.CV 2025-06 reject novelty 5.0 of 10

    A hierarchical coarse-to-fine diffusion transformer with cross-granularity distillation improves long-term driving video prediction, but the reported gains may be inflated by future-derived text prompts and a selected...

Pith tools