Pith. sign in

REVIEW 8 cited by

Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17918 v1 pith:7BCN6M2Y submitted 2023-11-29 cs.CV

classification cs.CV
keywords drivingmodelplanningautonomousmultiviewworlddrive-wmfirst
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In autonomous driving, predicting future events in advance and evaluating the foreseeable risks empowers autonomous vehicles to better plan their actions, enhancing safety and efficiency on the road. To this end, we propose Drive-WM, the first driving world model compatible with existing end-to-end planning models. Through a joint spatial-temporal modeling facilitated by view factorization, our model generates high-fidelity multiview videos in driving scenes. Building on its powerful generation ability, we showcase the potential of applying the world model for safe driving planning for the first time. Particularly, our Drive-WM enables driving into multiple futures based on distinct driving maneuvers, and determines the optimal trajectory according to the image-based rewards. Evaluation on real-world driving datasets verifies that our method could generate high-quality, consistent, and controllable multiview videos, opening up possibilities for real-world simulations and safe planning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A driving planner that predicts a future-ego-trajectory latent and retrieves executable trajectories from a fixed memory reaches 91.3 PDMS on NAVSIM v1 without reconstructing the scene.

  2. Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.

  3. Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Post-trained Cosmos world models generate controllable multi-view driving videos and LiDAR; augmenting real AV training data with these synthetic clips improves downstream perception and policy metrics, especially in ...

  4. ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.

  5. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

  6. LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

    cs.CV 2025-06 reject novelty 5.0 of 10

    A hierarchical coarse-to-fine diffusion transformer with cross-granularity distillation improves long-term driving video prediction, but the reported gains may be inflated by future-derived text prompts and a selected...

  7. 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model

    cs.CV 2025-09 conditional novelty 3.0 of 10

    A single-camera vision-language-model system scored 0.8747 on the CVPR 2024 E2E driving benchmark, the best camera-only result.

  8. Generative AI for Autonomous Driving: A Review

    cs.CV 2025-05 conditional novelty 2.0 of 10

    A review of generative models (VAEs, GANs, diffusion, transformers, LLMs) applied to map generation, scenario generation, trajectory prediction, and motion planning for autonomous driving.

Pith tools