REVIEW 8 cited by
WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Generating multi-camera street-view videos is critical for augmenting autonomous driving datasets, addressing the urgent demand for extensive and varied data. Due to the limitations in diversity and challenges in handling lighting conditions, traditional rendering-based methods are increasingly being supplanted by diffusion-based methods. However, a significant challenge in diffusion-based methods is ensuring that the generated sensor data preserve both intra-world consistency and inter-sensor coherence. To address these challenges, we combine an additional explicit world volume and propose the World Volume-aware Multi-camera Driving Scene Generator (WoVoGen). This system is specifically designed to leverage 4D world volume as a foundational element for video generation. Our model operates in two distinct phases: (i) envisioning the future 4D temporal world volume based on vehicle control sequences, and (ii) generating multi-camera videos, informed by this envisioned 4D temporal world volume and sensor interconnectivity. The incorporation of the 4D world volume empowers WoVoGen not only to generate high-quality street-view videos in response to vehicle control inputs but also to facilitate scene editing tasks.
Forward citations
Cited by 8 Pith papers
-
OmniNWM: Omniscient Driving Navigation World Models
OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.
-
Epona: Autoregressive Diffusion World Model for Autonomous Driving
An autoregressive diffusion world model jointly generates the next camera frame and a multi-step trajectory, enabling long videos and real-time planning for autonomous driving.
-
NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and Collaboration
NoiseController decomposes initial diffusion noise into scene-level foreground/background and shared/residual components, then collaborates them across views and frames, improving multi-view video consistency on nuScenes.
-
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
A 1B-parameter autoregressive world model generates over 40 seconds of controllable driving video with near-SOTA FVD, but key claims rest on inconsistent and cross-paper comparisons.
-
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
GEM generates controllable future RGB and depth ego-vision frames, conditioned on ego-trajectories, sparse object tokens, and human poses, across driving, egocentric, and drone domains.
-
Doe-1: Closed-Loop Autonomous Driving with Large World Model
Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.
-
Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method
UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...
-
2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model
A single-camera vision-language-model system scored 0.8747 on the CVPR 2024 E2E driving benchmark, the best camera-only result.
Discussion (0). Continue with ORCID to comment.