Pith. sign in

REVIEW 13 cited by

DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06845 v2 pith:L4TTM5LC submitted 2024-03-11 cs.CV

classification cs.CV
keywords drivingvideosdrivedreamer-2generategeneratedgenerationmodelworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

World models have demonstrated superiority in autonomous driving, particularly in the generation of multi-view driving videos. However, significant challenges still exist in generating customized driving videos. In this paper, we propose DriveDreamer-2, which builds upon the framework of DriveDreamer and incorporates a Large Language Model (LLM) to generate user-defined driving videos. Specifically, an LLM interface is initially incorporated to convert a user's query into agent trajectories. Subsequently, a HDMap, adhering to traffic regulations, is generated based on the trajectories. Ultimately, we propose the Unified Multi-View Model to enhance temporal and spatial coherence in the generated driving videos. DriveDreamer-2 is the first world model to generate customized driving videos, it can generate uncommon driving videos (e.g., vehicles abruptly cut in) in a user-friendly manner. Besides, experimental results demonstrate that the generated videos enhance the training of driving perception methods (e.g., 3D detection and tracking). Furthermore, video generation quality of DriveDreamer-2 surpasses other state-of-the-art methods, showcasing FID and FVD scores of 11.2 and 55.7, representing relative improvements of 30% and 50%.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

    cs.RO 2026-05 unverdicted novelty 7.0 of 10

    Driver-WM is a driver-centric latent world model for causal rollout of in-cabin dynamics conditioned on out-cabin traffic, unifying kinematics forecasting with behavioral and emotional recognition via dual-stream arch...

  2. Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.

  3. AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    AnyScene is an occupancy-centric framework using a Spatial-Temporal Occupancy Diffusion Transformer and Geometry-Grounded View Expansion to generate controllable driving scenes and videos from BEV layouts.

  4. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  5. ReSim: Reliable World Simulation for Autonomous Driving

    cs.CV 2025-06 unverdicted novelty 6.0 of 10

    ReSim is a controllable video world model trained on heterogeneous real and simulated driving data that achieves higher fidelity and controllability for both expert and non-expert actions, plus a Video2Reward module f...

  6. InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    InfiniVerse reconstructs 3D occupancy from one frame, extends scenes autoregressively, converts to video via diffusion, and uses re-projection feedback to achieve SOTA FID 6.4 and FVD 67.97 on Waymo and nuScenes.

  7. Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    SteinsGateDrive decouples LLM inference latency from vehicle control by pre-selecting alpha, beta, and gamma worldline futures that a runtime validates against safety contracts until abort conditions trigger.

  8. Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    Driver-WM rolls out in-cabin driver states in a compact latent space from frozen vision-language features, using traffic-conditioned dual streams and gated causal injection for long-horizon geometric and semantic forecasting.

  9. Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...

  10. Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Fine-tuning video generators on driving data can improve visual fidelity while degrading how accurately the model predicts the movement of cars and pedestrians.

  11. Non-invasive Assessment of Pancreatic Duct Hypertension Using Computational Flow Modeling

    physics.med-ph 2025-08 unverdicted novelty 5.0 of 10

    A computational model estimates pancreatic duct pressure non-invasively from MRCP geometry, with reported agreement against ERCP pressure measurements.

  12. 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model

    cs.CV 2025-09 conditional novelty 3.0 of 10

    A single-camera vision-language-model system scored 0.8747 on the CVPR 2024 E2E driving benchmark, the best camera-only result.

  13. Cosmos World Foundation Model Platform for Physical AI

    cs.CV 2025-01 unverdicted novelty 3.0 of 10

    The Cosmos platform supplies open-source pre-trained world models and supporting tools for building fine-tunable digital world simulations to train Physical AI.

Pith tools