REVIEW 48 cited by
DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the generation of high-quality driving videos, and driving policies for safe maneuvering. However, a critical limitation in relevant research lies in its predominant focus on gaming environments or simulated settings, thereby lacking the representation of real-world driving scenarios. Therefore, we introduce DriveDreamer, a pioneering world model entirely derived from real-world driving scenarios. Regarding that modeling the world in intricate driving scenes entails an overwhelming search space, we propose harnessing the powerful diffusion model to construct a comprehensive representation of the complex environment. Furthermore, we introduce a two-stage training pipeline. In the initial phase, DriveDreamer acquires a deep understanding of structured traffic constraints, while the subsequent stage equips it with the ability to anticipate future states. The proposed DriveDreamer is the first world model established from real-world driving scenarios. We instantiate DriveDreamer on the challenging nuScenes benchmark, and extensive experiments verify that DriveDreamer empowers precise, controllable video generation that faithfully captures the structural constraints of real-world traffic scenarios. Additionally, DriveDreamer enables the generation of realistic and reasonable driving policies, opening avenues for interaction and practical applications.
Forward citations
Cited by 48 Pith papers
-
Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure
DreamerV3's imagined rollouts are insensitive to friction changes that cause real gait collapse, revealing that world models extrapolate kinematically rather than dynamically.
-
LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
LSD-3D generates explicit, 3D-consistent driving scenes by combining a generated proxy mesh with geometry-grounded distillation from a 2D diffusion model.
-
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
A joint diffusion framework trains a Stable Diffusion generator and a semantic occupancy perception model together, so each task improves the other, producing text-conditional RGB-occupancy pairs.
-
Population-Scalable Multi-Agent World Modeling
Khora decouples world-state evolution from visual rendering through a shared STBoard and fixed-dimensional per-view renderers, enabling inference-time addition and removal of agents without retraining.
-
GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes
GSRAIN fuses measured-raindrop-calibrated Gaussian streaks/haze with a diffusion-based rainy-appearance transfer in 3D Gaussian Splatting scenes, enabling 0–13 mm/h rainfall control for autonomous-driving testing.
-
Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving
A driving planner that predicts a future-ego-trajectory latent and retrieves executable trajectories from a fixed memory reaches 91.3 PDMS on NAVSIM v1 without reconstructing the scene.
-
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
Adding persistently updated, supervised world-state register tokens to streaming multi-agent diffusion improves cross-agent consistency and visual quality in two-agent Minecraft generation.
-
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation
A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.
-
VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
A vision-language model writes a simulation program—grounded by segmentation and 3D tools—that predicts physically plausible futures from an image and caption, outperforming video generators on modified benchmarks.
-
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
FSU-QA is a new nuScenes-based VQA benchmark for foresight reasoning; current VLMs score about 41-49% accuracy, and fine-tuning an 8B model on it outperforms GPT-5 and Gemini on the benchmark.
-
Epona: Autoregressive Diffusion World Model for Autonomous Driving
An autoregressive diffusion world model jointly generates the next camera frame and a multi-step trajectory, enabling long videos and real-time planning for autonomous driving.
-
WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration
A pipeline that restores corrupted novel-view videos with a video diffusion model and jointly denoises multiple viewpoints to improve 3D scene exploration from a single image.
-
ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
A hierarchical leader-follower Gaussian world model improves multi-task bimanual manipulation success rates over prior single-arm-based methods.
-
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
TARDIS is a transformer world model trained on STRIDE, a graph-structured street-view dataset, with claimed abilities in controllable image generation, georeferencing, self-driving actions, and temporal simulation.
-
Dreamland: Controllable World Creation with Simulator and Generative Models
A three-stage hybrid pipeline uses an intermediate layered world representation to refine simulator-rendered driving scenes into realistic, controllable images and videos.
-
Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
MeWM combines a GPT-style policy, a diffusion tumor dynamics model, and a survival analysis heuristic to simulate post-treatment tumor appearance and select TACE treatment plans, improving physician F1-score by 13 points.
-
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.
-
PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
A pose-control module using self-supervised depth and forward plus inverse warping losses improves camera alignment in diffusion and autoregressive world models.
-
NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration
NoiseController decomposes initial diffusion noise into scene-level foreground/background and shared/residual components, then collaborates them across views and frames, improving multi-view video consistency on nuScenes.
-
Self-Consistent Model-based Adaptation for Visual Reinforcement Learning
SCMA trains a policy-agnostic observation denoiser, using a pre-trained world model as a clean-distribution reference, and shows improved visual RL performance under distractions.
-
AdaWM: Adaptive World Model based Planning for Autonomous Driving
AdaWM selectively finetunes either the dynamics model or the policy of a pretrained world-model driving agent according to which mismatch dominates, improving success rates in CARLA.
-
DreamDrive: Generative 4D Scene Modeling from Street View Images
DreamDrive generates 3D-consistent driving videos from a single image by lifting diffusion-generated reference frames into a hybrid static and dynamic 4D Gaussian scene.
-
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
A 1B-parameter autoregressive world model generates over 40 seconds of controllable driving video with near-SOTA FVD, but key claims rest on inconsistent and cross-paper comparisons.
-
DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes
DriveEditor uses depth-aware 3D bounding box projection and single-reference appearance cues to reposition, replace, remove, and insert objects in driving videos with a single diffusion framework.
-
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
VLM-AD uses GPT-4o-generated reasoning and action annotations as auxiliary supervision to improve end-to-end autonomous driving planning without VLM inference.
-
An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training
An end-to-end, non-autoregressive 3D occupancy world model warps dynamic voxels via predicted flow, moves static voxels by pose, and uses image-based rendering supervision, achieving state-of-the-art results on three ...
-
StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
StreetCrafter conditions a video diffusion model on LiDAR point cloud renderings to synthesize controllable street views, and distills it into a real-time 3D Gaussian representation.
-
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
GEM generates controllable future RGB and depth ego-vision frames, conditioned on ego-trajectories, sparse object tokens, and human poses, across driving, egocentric, and drone domains.
-
Doe-1: Closed-Loop Autonomous Driving with Large World Model
Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.
-
ARCON: Advancing Auto-Regressive Continuation for Driving Videos
Alternating semantic-map tokens and RGB tokens during autoregressive video continuation improves long-term consistency and FVD for driving videos.
-
FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes
A generation-reconstruction pipeline with a diffusion enhancer trained on simulated degradations enables off-trajectory camera rendering in driving scenes.
-
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
A framework that jointly generates multi-view camera images and LiDAR point clouds for driving scenes, with cross-modal consistency, improving on prior single-modality generators.
-
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
MagicDrive-V2 shows that matching geometric-control encoders to a 3D VAE's 4x temporal compression enables controllable, high-resolution, long multi-view driving video generation.
-
Generative World Explorer
GenEx generates consistent 360-degree exploration videos from a single egocentric panorama and shows that LLM agents using these imagined views make more accurate decisions.
-
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
DrivingSphere combines occupancy-based 4D world generation with video diffusion to create a closed-loop simulation environment for autonomous driving evaluation.
-
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
EgoVid-5M is a 5M-clip curated egocentric video dataset with action and kinematic annotations, and EgoDreamer generates egocentric videos from text and camera control signals.
-
Quo Vadis, World Modeling?
An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.
-
UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving
A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.
-
MobiWorld: World Models for Mobile Wireless Network
The paper proposes MobiWorld, a diffusion-based controllable world model for mobile networks, and reports a preliminary energy-saving optimization case study.
-
EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
A Real2Sim2Real framework that aligns simulator dynamics via differentiable parameter fitting and renders photorealistic policy-training videos with a diffusion model, improving real-world manipulation success.
-
LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
A hierarchical coarse-to-fine diffusion transformer with cross-granularity distillation improves long-term driving video prediction, but the reported gains may be inflated by future-derived text prompts and a selected...
-
Dream to Drive with Predictive Individual World Model
PIWM, an individual-vehicle world model with self-attention interaction modeling and trajectory-prediction representation learning, beats DreamerV3 and model-free RL on INTERACTION-based driving benchmarks.
-
GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction
A world model operating on 3D Gaussians forecasts the current occupancy from the previous frame and current RGB, improving mIoU by about 2 points on nuScenes without meaningful added latency.
-
Physical Informed Driving World Model
DrivePhysica adds coordinate alignment, 3D instance flow, and box-coordinate guidance to a diffusion world model, achieving state-of-the-art FID/FVD on nuScenes and improving StreamPETR NDS by 3.6 points when mixed wi...
-
Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model
A reactive closed-loop driving simulator that uses a diffusion renderer with retrieval from real recordings, plus a nuPlan behavioral controller, to generate sensor images in response to an end-to-end driving model's actions.
-
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
ReconDreamer fine-tunes a driving world model as an online restorer and progressively expands novel-trajectory training data, reporting first-time effective rendering of multi-lane shifts in driving scenes.
-
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
A diffusion transformer with joint attention over all camera views and frames, a lightweight BEV controller, and a re-weighted object loss reports state-of-the-art multi-view driving video generation (FVD 37.8 on nuScenes).
-
What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality
VAMP scores generated videos by combining per-object color, shape, and texture consistency with centroid velocity and acceleration smoothness, but its physics-based and human-aligned claims are not supported by the re...
Discussion (0). Continue with ORCID to comment.