REVIEW 7 cited by
Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning world models can teach an agent how the world works in an unsupervised manner. Even though it can be viewed as a special case of sequence modeling, progress for scaling world models on robotic applications such as autonomous driving has been somewhat less rapid than scaling language models with Generative Pre-trained Transformers (GPT). We identify two reasons as major bottlenecks: dealing with complex and unstructured observation space, and having a scalable generative model. Consequently, we propose Copilot4D, a novel world modeling approach that first tokenizes sensor observations with VQVAE, then predicts the future via discrete diffusion. To efficiently decode and denoise tokens in parallel, we recast Masked Generative Image Transformer as discrete diffusion and enhance it with a few simple changes, resulting in notable improvement. When applied to learning world models on point cloud observations, Copilot4D reduces prior SOTA Chamfer distance by more than 65% for 1s prediction, and more than 50% for 3s prediction, across NuScenes, KITTI Odometry, and Argoverse2 datasets. Our results demonstrate that discrete diffusion on tokenized agent experience can unlock the power of GPT-like unsupervised learning for robotics.
Forward citations
Cited by 7 Pith papers
-
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
NeSy-Route supplies 10,821 optimally labeled remote-sensing route-planning tasks plus a three-level neuro-symbolic protocol that reveals major perception and planning deficits in current MLLMs.
-
AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models
AutoWorld learns a self-supervised LiDAR occupancy world model and conditions a diffusion-based motion generator on its forecasts, reporting the top Waymo Sim Agents realism score.
-
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
A joint video and LiDAR generation framework for driving scenes, conditioned on shared scene layouts, VLM captions, and BEV features, achieves SOTA generation and downstream perception gains on nuScenes.
-
Quo Vadis, World Modeling?
An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.
-
UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving
A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.
-
Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving
Diffusion-based generative models, using discrete categorical diffusion conditioned on BEV features, improve 3D occupancy prediction and downstream planning for autonomous driving.
-
Reinforcement Learning: From Algorithms To Foundation Models
A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.
Discussion (0). Continue with ORCID to comment.