Pith. sign in

REVIEW 48 cited by

DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09777 v2 pith:VYS6N5T2 submitted 2023-09-18 cs.CV

classification cs.CV
keywords drivingdrivedreamerworldmodelreal-worldscenariosgenerationautonomous
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the generation of high-quality driving videos, and driving policies for safe maneuvering. However, a critical limitation in relevant research lies in its predominant focus on gaming environments or simulated settings, thereby lacking the representation of real-world driving scenarios. Therefore, we introduce DriveDreamer, a pioneering world model entirely derived from real-world driving scenarios. Regarding that modeling the world in intricate driving scenes entails an overwhelming search space, we propose harnessing the powerful diffusion model to construct a comprehensive representation of the complex environment. Furthermore, we introduce a two-stage training pipeline. In the initial phase, DriveDreamer acquires a deep understanding of structured traffic constraints, while the subsequent stage equips it with the ability to anticipate future states. The proposed DriveDreamer is the first world model established from real-world driving scenarios. We instantiate DriveDreamer on the challenging nuScenes benchmark, and extensive experiments verify that DriveDreamer empowers precise, controllable video generation that faithfully captures the structural constraints of real-world traffic scenarios. Additionally, DriveDreamer enables the generation of realistic and reasonable driving policies, opening avenues for interaction and practical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 48 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

    cs.RO 2026-07 conditional novelty 7.0 of 10

    DreamerV3's imagined rollouts are insensitive to friction changes that cause real gait collapse, revealing that world models extrapolate kinematically rather than dynamically.

  2. LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding

    cs.CV 2025-08 conditional novelty 7.0 of 10

    LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.

  3. OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A joint diffusion framework trains a Stable Diffusion generator and a semantic occupancy perception model together, so each task improves the other, producing text-conditional RGB-occupancy pairs.

  4. Population-Scalable Multi-Agent World Modeling

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Khora decouples world-state evolution from visual rendering through a shared STBoard and fixed-dimensional per-view renderers, enabling inference-time addition and removal of agents without retraining.

  5. GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes

    cs.CV 2026-08 conditional novelty 6.0 of 10

    GSRAIN fuses measured-raindrop-calibrated Gaussian streaks/haze with a diffusion-based rainy-appearance transfer in 3D Gaussian Splatting scenes, enabling 0–13 mm/h rainfall control for autonomous-driving testing.

  6. Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A driving planner that predicts a future-ego-trajectory latent and retrieves executable trajectories from a fixed memory reaches 91.3 PDMS on NAVSIM v1 without reconstructing the scene.

  7. Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adding persistently updated, supervised world-state register tokens to streaming multi-agent diffusion improves cross-agent consistency and visual quality in two-agent Minecraft generation.

  8. Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.

  9. VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A vision-language model writes a simulation program—grounded by segmentation and 3D tools—that predicts physically plausible futures from an image and caption, outperforming video generators on modified benchmarks.

  10. Thinking Ahead: Foresight Intelligence in MLLMs and World Model

    cs.CV 2025-11 conditional novelty 6.0 of 10

    FSU-QA is a new nuScenes-based VQA benchmark for foresight reasoning; current VLMs score about 41-49% accuracy, and fine-tuning an 8B model on it outperforms GPT-5 and Gemini on the benchmark.

  11. Epona: Autoregressive Diffusion World Model for Autonomous Driving

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive diffusion world model jointly generates the next camera frame and a multi-step trajectory, enabling long videos and real-time planning for autonomous driving.

  12. WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A pipeline that restores corrupted novel-view videos with a video diffusion model and jointly denoises multiple viewpoints to improve 3D scene exploration from a single image.

  13. ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A hierarchical leader-follower Gaussian world model improves multi-task bimanual manipulation success rates over prior single-arm-based methods.

  14. TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy

    cs.CV 2025-06 reject novelty 6.0 of 10

    TARDIS is a transformer world model trained on STRIDE, a graph-structured street-view dataset, with claimed abilities in controllable image generation, georeferencing, self-driving actions, and temporal simulation.

  15. Dreamland: Controllable World Creation with Simulator and Generative Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A three-stage hybrid pipeline uses an intermediate layered world representation to refine simulator-rendered driving scenes into realistic, controllable images and videos.

  16. Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MeWM combines a GPT-style policy, a diffusion tumor dynamics model, and a survival analysis heuristic to simulate post-treatment tumor appearance and select TACE treatment plans, improving physician F1-score by 13 points.

  17. ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.

  18. PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A pose-control module using self-supervised depth and forward plus inverse warping losses improves camera alignment in diffusion and autoregressive world models.

  19. NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration

    cs.CV 2025-04 conditional novelty 6.0 of 10

    NoiseController decomposes initial diffusion noise into scene-level foreground/background and shared/residual components, then collaborates them across views and frames, improving multi-view video consistency on nuScenes.

  20. Self-Consistent Model-based Adaptation for Visual Reinforcement Learning

    cs.CV 2025-02 conditional novelty 6.0 of 10

    SCMA trains a policy-agnostic observation denoiser, using a pre-trained world model as a clean-distribution reference, and shows improved visual RL performance under distractions.

  21. AdaWM: Adaptive World Model based Planning for Autonomous Driving

    cs.RO 2025-01 conditional novelty 6.0 of 10

    AdaWM selectively finetunes either the dynamics model or the policy of a pretrained world-model driving agent according to which mismatch dominates, improving success rates in CARLA.

  22. DreamDrive: Generative 4D Scene Modeling from Street View Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.

  23. DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 1B-parameter autoregressive world model generates over 40 seconds of controllable driving video with near-SOTA FVD, but key claims rest on inconsistent and cross-paper comparisons.

  24. DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DriveEditor uses depth-aware 3D bounding box projection and single-reference appearance cues to reposition, replace, remove, and insert objects in driving videos with a single diffusion framework.

  25. VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

    cs.CV 2024-12 conditional novelty 6.0 of 10

    VLM-AD uses GPT-4o-generated reasoning and action annotations as auxiliary supervision to improve end-to-end autonomous driving planning without VLM inference.

  26. An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An end-to-end, non-autoregressive 3D occupancy world model warps dynamic voxels via predicted flow, moves static voxels by pose, and uses image-based rendering supervision, achieving state-of-the-art results on three ...

  27. StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    StreetCrafter conditions a video diffusion model on LiDAR point cloud renderings to synthesize controllable street views, and distills it into a real-time 3D Gaussian representation.

  28. GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GEM generates controllable future RGB and depth ego-vision frames, conditioned on ego-trajectories, sparse object tokens, and human poses, across driving, egocentric, and drone domains.

  29. Doe-1: Closed-Loop Autonomous Driving with Large World Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.

  30. ARCON: Advancing Auto-Regressive Continuation for Driving Videos

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Alternating semantic-map tokens and RGB tokens during autoregressive video continuation improves long-term consistency and FVD for driving videos.

  31. FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A generation-reconstruction pipeline with a diffusion enhancer trained on simulated degradations enables off-trajectory camera rendering in driving scenes.

  32. HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A framework that jointly generates multi-view camera images and LiDAR point clouds for driving scenes, with cross-modal consistency, improving on prior single-modality generators.

  33. MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control

    cs.CV 2024-11 conditional novelty 6.0 of 10

    MagicDrive-V2 shows that matching geometric-control encoders to a 3D VAE's 4x temporal compression enables controllable, high-resolution, long multi-view driving video generation.

  34. Generative World Explorer

    cs.CV 2024-11 conditional novelty 6.0 of 10

    GenEx generates consistent 360-degree exploration videos from a single egocentric panorama and shows that LLM agents using these imagined views make more accurate decisions.

  35. DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation

    cs.RO 2024-11 conditional novelty 6.0 of 10

    DrivingSphere combines occupancy-based 4D world generation with video diffusion to create a closed-loop simulation environment for autonomous driving evaluation.

  36. EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    EgoVid-5M is a 5M-clip curated egocentric video dataset with action and kinematic annotations, and EgoDreamer generates egocentric videos from text and camera control signals.

  37. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  38. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

  39. MobiWorld: World Models for Mobile Wireless Network

    cs.NI 2025-07 conditional novelty 5.0 of 10

    The paper proposes MobiWorld, a diffusion-based controllable world model for mobile networks, and reports a preliminary energy-saving optimization case study.

  40. EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A Real2Sim2Real framework that aligns simulator dynamics via differentiable parameter fitting and renders photorealistic policy-training videos with a diffusion model, improving real-world manipulation success.

  41. LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

    cs.CV 2025-06 reject novelty 5.0 of 10

    A hierarchical coarse-to-fine diffusion transformer with cross-granularity distillation improves long-term driving video prediction, but the reported gains may be inflated by future-derived text prompts and a selected...

  42. Dream to Drive with Predictive Individual World Model

    cs.RO 2025-01 conditional novelty 5.0 of 10

    PIWM, an individual-vehicle world model with self-attention interaction modeling and trajectory-prediction representation learning, beats DreamerV3 and model-free RL on INTERACTION-based driving benchmarks.

  43. GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A world model operating on 3D Gaussians forecasts the current occupancy from the previous frame and current RGB, improving mIoU by about 2 points on nuScenes without meaningful added latency.

  44. Physical Informed Driving World Model

    cs.CV 2024-12 conditional novelty 5.0 of 10

    DrivePhysica adds coordinate alignment, 3D instance flow, and box-coordinate guidance to a diffusion world model, achieving state-of-the-art FID/FVD on nuScenes and improving StreamPETR NDS by 3.6 points when mixed wi...

  45. Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A reactive closed-loop driving simulator that uses a diffusion renderer with retrieval from real recordings, plus a nuPlan behavioral controller, to generate sensor images in response to an end-to-end driving model's actions.

  46. ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration

    cs.CV 2024-11 conditional novelty 5.0 of 10

    ReconDreamer fine-tunes a driving world model as an online restorer and progressively expands novel-trajectory training data, reporting first-time effective rendering of multi-lane shifts in driving scenes.

  47. Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A diffusion transformer with joint attention over all camera views and frames, a lightweight BEV controller, and a re-weighted object loss reports state-of-the-art multi-view driving video generation (FVD 37.8 on nuScenes).

  48. What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality

    cs.CV 2024-11 reject novelty 4.0 of 10

    VAMP scores generated videos by combining per-object color, shape, and texture consistency with centroid velocity and acceleration smoothness, but its physics-based and human-aligned claims are not supported by the re...

Pith tools