{"work":{"id":"313484e6-a442-4522-8e19-d07e502844a8","openalex_id":null,"doi":null,"arxiv_id":"2309.17080","raw_key":null,"title":"GAIA-1: A Generative World Model for Autonomous Driving","authors":null,"authors_text":"Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall","year":2023,"venue":"cs.CV","abstract":"Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves.\n  To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle behavior and scene features. Our approach casts world modeling as an unsupervised sequence modeling problem by mapping the inputs to discrete tokens, and predicting the next token in the sequence. Emerging properties from our model include learning high-level structures and scene dynamics, contextual awareness, generalization, and understanding of geometry. The power of GAIA-1's learned representation that captures expectations of future events, combined with its ability to generate realistic samples, provides new possibilities for innovation in the field of autonomy, enabling enhanced and accelerated training of autonomous driving technology.","external_url":"https://arxiv.org/abs/2309.17080","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-10T11:47:02.949785+00:00","pith_arxiv_id":"2309.17080","created_at":"2026-05-10T05:25:54.495379+00:00","updated_at":"2026-07-10T11:47:02.949785+00:00","title_quality_ok":true,"display_title":"GAIA-1: A Generative World Model for Autonomous Driving","render_title":"GAIA-1: A Generative World Model for Autonomous Driving"},"hub":{"state":{"work_id":"313484e6-a442-4522-8e19-d07e502844a8","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":130,"external_cited_by_count":null,"distinct_field_count":7,"first_pith_cited_at":"2023-12-21T18:46:41+00:00","last_pith_cited_at":"2026-07-09T07:28:21+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T14:39:27.279446+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":25}],"polarity_counts":[{"context_polarity":"background","n":24},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"GAIA-1: A Generative World Model for Autonomous Driving","claims":[{"claim_text":"Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves.\n  To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle b","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"tently surpasses state-of-the-art VLA and world-model base- lines on both planning and future-generation benchmarks. Project page:https://vlaworld.github.io 1. Introduction Recently, two major paradigms have gained attention in end- to-end autonomous driving: Vision-Language-Action (VLA) models [3, 16, 20, 30, 33, 42, 76, 84, 86] and World Mod- els [22, 28, 41, 44, 59-61]. Unlike traditional end-to-end pipelines [12, 13, 15, 25, 29] that learn perception and con- trol only from driving data, VLA","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Multimodal Large Language Models (MLLMs) have rapidly advanced beyond basic perceptual tasks such as image recognition and captioning, and are now being deployed in physically grounded applications, including robotics [14, 94, 34] and autonomous driving [81, 69]. These deployments position MLLMs as a foundation toward world models that can understand and predict the dynamics of physical environments [28, 30]. However, despite this progress, current MLLMs still exhibit notable challenges in under","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Vista: A generalizable driving world model with high fidelity and versatile controllability.Advances in Neural Information Processing Systems, 37:91560-91596, 2024. [31] Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, and Jiwen Lu. Drive- dreamer: Towards real-world-drive world models for autonomous driving. InEuropean confer- ence on computer vision, pages 55-72. Springer, 2024. 11 [32] Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"In diffusion form, models denoise latent repre- sentations conditioned on past observations and controls. The key differentiator from occupancy-based models discussed in the next section is targeting view-level appearance and temporal coherence directly, which is useful for training and evaluating end-to-end autonomy stacks that operate on raw sensor inputs. GAIA-1 [138] demonstrated that large-scale generative models can produce controllable scenario generation with emergent understanding of sc","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Related Works World Models in Autonomous Driving.Given histor- ical observations, a world model aims to predict future states. It is gaining growing attention in autonomous driv- ing for its capability to generate high fidelity data [53] and enhance driving safety [17, 39]. In autonomous driv- ing, three mainstream world models have emerged: vision- based [12, 43, 44], occupancy-based [9, 41, 54], and LiDAR-based [13, 20, 45, 46, 52]. While former two types of world models are widely explored, t","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"lates environment dynamics, has regained attention. Video generation has become a leading paradigm, supported by advances in generative modeling, large-scale video datasets, and with wide applicability. In autonomous driving, tem- porally grounded video prediction provides rich context for understanding and decision-making. Several methods treat pure video generation as world modeling. GAIA [20] conditions generation on image, text, and action inputs. GAIA-2 [39] extends this to multi-view scene","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks GAIA-1: A Generative World Model for Autonomous Driving because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (25 contexts).","role_counts":[{"n":25,"context_role":"background"}]},"error":null,"updated_at":"2026-07-02T03:42:13.119046+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"57a6c204-5cae-49da-a6e1-f8faf0edb080","orcid":null,"display_name":"Anthony Hu"},{"id":"e94496b4-1824-4334-92d9-b6af8cad0e0f","orcid":null,"display_name":"Lloyd Russell"},{"id":"634a9703-c788-40f4-95a2-07cc1f0d1b43","orcid":null,"display_name":"Hudson Yeo"},{"id":"558bf95d-f0f5-4204-ab07-0b55c022a013","orcid":null,"display_name":"Zak Murez"},{"id":"b19d06a0-99a3-4ddc-ae49-c3ad0bf61f67","orcid":null,"display_name":"George Fedoseev"},{"id":"b1948448-c734-4989-bf15-f4fa4ce51b9c","orcid":null,"display_name":"Alex Kendall"}]},"error":null,"updated_at":"2026-07-02T03:42:13.115217+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:09:27.358410+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","work_id":"a9c28401-f16a-4933-89f0-788e2f94e52b","shared_citers":13},{"title":"Gaia-2: A controllable multi-view generative world model for autonomous driving","work_id":"1339e674-d09b-48b4-8e6f-efe55dcab22e","shared_citers":12},{"title":"World Models","work_id":"07227eee-8445-4c98-bce4-c6a6fd5ed907","shared_citers":12},{"title":"A survey of world models for autonomous driving","work_id":"5775e072-5965-490b-b444-42d28c541e6f","shared_citers":9},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":9},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":9},{"title":"Available: https://arxiv.org/abs/2311.13549","work_id":"e601ca37-4a47-4d11-aefc-857f5ca5501a","shared_citers":7},{"title":"Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation","work_id":"c1e2622d-92db-4a96-aa17-8ef832270ed1","shared_citers":7},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":6},{"title":"A comprehensive survey on world models for embodied AI","work_id":"d4ba9e8d-69e0-462a-b353-627172c1e415","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":6},{"title":"Dream to Control: Learning Behaviors by Latent Imagination","work_id":"5103f4be-344a-4139-8504-eaa59f5bac9d","shared_citers":6},{"title":"Enhancing end-to-end autonomous driving with latent world model","work_id":"f07d6be8-7c83-4e2c-9f71-013acac7ffda","shared_citers":6},{"title":"Mastering Diverse Domains through World Models","work_id":"6aeb260f-8c7c-4f9c-b98b-067cd7c59acd","shared_citers":6},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":6},{"title":"World Simulation with Video Foundation Models for Physical AI","work_id":"1dc393b8-98c3-43bd-8ab0-25d7c2a9705b","shared_citers":6},{"title":"Advancing open-source world models","work_id":"3446760e-6460-429e-82f6-6a5133ce4c8a","shared_citers":5},{"title":"arXiv preprint arXiv:2506.08052 (2025)","work_id":"02bd7e58-a437-4e26-912d-8f1a2e695fa7","shared_citers":5},{"title":"arXiv preprint arXiv:2506.13757 (2025)","work_id":"945172fb-0f3b-43be-88b4-ae61042517e9","shared_citers":5},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":5},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":5},{"title":"Learning interactive real-world simulators","work_id":"16f38691-7ab6-4e23-bba5-6b656579e579","shared_citers":5},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":5},{"title":"Rethinking the open-loop evaluation of end-to-end autonomous driving in nuscenes","work_id":"c287db5a-5852-407d-a032-8f9ed913f70a","shared_citers":5}],"time_series":[{"n":1,"year":2024},{"n":2,"year":2025},{"n":34,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:09:54.367123+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:09:43.915265+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"GAIA-1: A Generative World Model for Autonomous Driving","claims":[{"claim_text":"Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves.\n  To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle b","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"tently surpasses state-of-the-art VLA and world-model base- lines on both planning and future-generation benchmarks. Project page:https://vlaworld.github.io 1. Introduction Recently, two major paradigms have gained attention in end- to-end autonomous driving: Vision-Language-Action (VLA) models [3, 16, 20, 30, 33, 42, 76, 84, 86] and World Mod- els [22, 28, 41, 44, 59-61]. Unlike traditional end-to-end pipelines [12, 13, 15, 25, 29] that learn perception and con- trol only from driving data, VLA","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Multimodal Large Language Models (MLLMs) have rapidly advanced beyond basic perceptual tasks such as image recognition and captioning, and are now being deployed in physically grounded applications, including robotics [14, 94, 34] and autonomous driving [81, 69]. These deployments position MLLMs as a foundation toward world models that can understand and predict the dynamics of physical environments [28, 30]. However, despite this progress, current MLLMs still exhibit notable challenges in under","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Vista: A generalizable driving world model with high fidelity and versatile controllability.Advances in Neural Information Processing Systems, 37:91560-91596, 2024. [31] Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, and Jiwen Lu. Drive- dreamer: Towards real-world-drive world models for autonomous driving. InEuropean confer- ence on computer vision, pages 55-72. Springer, 2024. 11 [32] Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"In diffusion form, models denoise latent repre- sentations conditioned on past observations and controls. The key differentiator from occupancy-based models discussed in the next section is targeting view-level appearance and temporal coherence directly, which is useful for training and evaluating end-to-end autonomy stacks that operate on raw sensor inputs. GAIA-1 [138] demonstrated that large-scale generative models can produce controllable scenario generation with emergent understanding of sc","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Related Works World Models in Autonomous Driving.Given histor- ical observations, a world model aims to predict future states. It is gaining growing attention in autonomous driv- ing for its capability to generate high fidelity data [53] and enhance driving safety [17, 39]. In autonomous driv- ing, three mainstream world models have emerged: vision- based [12, 43, 44], occupancy-based [9, 41, 54], and LiDAR-based [13, 20, 45, 46, 52]. While former two types of world models are widely explored, t","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"lates environment dynamics, has regained attention. Video generation has become a leading paradigm, supported by advances in generative modeling, large-scale video datasets, and with wide applicability. In autonomous driving, tem- porally grounded video prediction provides rich context for understanding and decision-making. Several methods treat pure video generation as world modeling. GAIA [20] conditions generation on image, text, and action inputs. GAIA-2 [39] extends this to multi-view scene","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks GAIA-1: A Generative World Model for Autonomous Driving because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (25 contexts).","role_counts":[{"n":25,"context_role":"background"}]},"error":null,"updated_at":"2026-07-02T03:42:13.121447+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"GAIA-1: A Generative World Model for Autonomous Driving","claims":[{"claim_text":"Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves.\n  To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle b","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks GAIA-1: A Generative World Model for Autonomous Driving because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:09:39.663226+00:00"}},"summary":{"title":"GAIA-1: A Generative World Model for Autonomous Driving","claims":[{"claim_text":"Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves.\n  To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle b","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks GAIA-1: A Generative World Model for Autonomous Driving because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","work_id":"a9c28401-f16a-4933-89f0-788e2f94e52b","shared_citers":13},{"title":"Gaia-2: A controllable multi-view generative world model for autonomous driving","work_id":"1339e674-d09b-48b4-8e6f-efe55dcab22e","shared_citers":12},{"title":"World Models","work_id":"07227eee-8445-4c98-bce4-c6a6fd5ed907","shared_citers":12},{"title":"A survey of world models for autonomous driving","work_id":"5775e072-5965-490b-b444-42d28c541e6f","shared_citers":9},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":9},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":9},{"title":"Available: https://arxiv.org/abs/2311.13549","work_id":"e601ca37-4a47-4d11-aefc-857f5ca5501a","shared_citers":7},{"title":"Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation","work_id":"c1e2622d-92db-4a96-aa17-8ef832270ed1","shared_citers":7},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":6},{"title":"A comprehensive survey on world models for embodied AI","work_id":"d4ba9e8d-69e0-462a-b353-627172c1e415","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":6},{"title":"Dream to Control: Learning Behaviors by Latent Imagination","work_id":"5103f4be-344a-4139-8504-eaa59f5bac9d","shared_citers":6},{"title":"Enhancing end-to-end autonomous driving with latent world model","work_id":"f07d6be8-7c83-4e2c-9f71-013acac7ffda","shared_citers":6},{"title":"Mastering Diverse Domains through World Models","work_id":"6aeb260f-8c7c-4f9c-b98b-067cd7c59acd","shared_citers":6},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":6},{"title":"World Simulation with Video Foundation Models for Physical AI","work_id":"1dc393b8-98c3-43bd-8ab0-25d7c2a9705b","shared_citers":6},{"title":"Advancing open-source world models","work_id":"3446760e-6460-429e-82f6-6a5133ce4c8a","shared_citers":5},{"title":"arXiv preprint arXiv:2506.08052 (2025)","work_id":"02bd7e58-a437-4e26-912d-8f1a2e695fa7","shared_citers":5},{"title":"arXiv preprint arXiv:2506.13757 (2025)","work_id":"945172fb-0f3b-43be-88b4-ae61042517e9","shared_citers":5},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":5},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":5},{"title":"Learning interactive real-world simulators","work_id":"16f38691-7ab6-4e23-bba5-6b656579e579","shared_citers":5},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":5},{"title":"Rethinking the open-loop evaluation of end-to-end autonomous driving in nuscenes","work_id":"c287db5a-5852-407d-a032-8f9ed913f70a","shared_citers":5}],"time_series":[{"n":1,"year":2024},{"n":2,"year":2025},{"n":34,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"b1948448-c734-4989-bf15-f4fa4ce51b9c","orcid":null,"display_name":"Alex Kendall","source":"manual","import_confidence":0.72},{"id":"57a6c204-5cae-49da-a6e1-f8faf0edb080","orcid":null,"display_name":"Anthony Hu","source":"manual","import_confidence":0.72},{"id":"b19d06a0-99a3-4ddc-ae49-c3ad0bf61f67","orcid":null,"display_name":"George Fedoseev","source":"manual","import_confidence":0.72},{"id":"634a9703-c788-40f4-95a2-07cc1f0d1b43","orcid":null,"display_name":"Hudson Yeo","source":"manual","import_confidence":0.72},{"id":"e94496b4-1824-4334-92d9-b6af8cad0e0f","orcid":null,"display_name":"Lloyd Russell","source":"manual","import_confidence":0.72},{"id":"558bf95d-f0f5-4204-ab07-0b55c022a013","orcid":null,"display_name":"Zak Murez","source":"manual","import_confidence":0.72}]}}