Pith. sign in

Paper Citation Record · LEDGER

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

As of 21 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 32 inbound Pith citation observations for arXiv:2412.19505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19505 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:21:49.227847Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:54:07.752159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:57.095080Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 992cf9e4-efa8-4255-b75c-ddaa1a364e19 · outbound

This paper cites Model-Based Offline Planning.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Model-Based Offline Planning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.862449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.862449Z digest=sha256:7ea12b4adfeeb3968b6742fb83f10f2a6321b945bbb67c1d14ef15e2aa1cfd0c

Observation 5ef92d33-62fc-40cb-ae00-cc8a1633112b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.872608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.872608Z digest=sha256:f99d5d5fcf49e766c2fb20847e22127ef4927a9d286490447453eb6e0b866b25

Observation 5e2d042e-35ff-48ed-9fbe-5effc3a6053a · outbound

This paper cites Language Models are Few-Shot Learners.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.885106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.885106Z digest=sha256:8255279215b5bfef30dbad5681ad35800320061f00d5dd945ca26340248a8f00

Observation 7f6c6b40-c619-41ee-a230-40ab15309a8c · outbound

This paper cites nuscenes: A multi- modal dataset for autonomous driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT nuscenes: A multi- modal dataset for autonomous driving

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.294271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:48.892714Z digest=sha256:22dcb8b84685e33cf34b997be1105b3b8ba0d44ae1e37c3242b782096779b844

Observation 56125511-3d54-4d01-a35c-ee61dcd8f350 · outbound

This paper cites NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.904049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.904049Z digest=sha256:0f30031db29c57025f85b14e7c64c612d37fbba9046b860d5692d0458fc75d48

Observation f84c48d1-dd00-426a-8600-923881c8408d · outbound

This paper cites UMBRELLA: Uncertainty-Aware Model-Based Offline Reinforcement Learning Leveraging Planning.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT UMBRELLA: Uncertainty-Aware Model-Based Offline Reinforcement Learning Leveraging Planning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.920295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.920295Z digest=sha256:08d292607c4cc7866b1d6f5f4eb9ca9e4dd1a25a5cea1708d9ea06d3b36a66fe

Observation 5b727d2d-5693-4a1d-b599-5065ae919f25 · outbound

This paper cites Uncertainty-aware model-based offline reinforcement learning for automated driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Uncertainty-aware model-based offline reinforcement learning for automated driving

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.272580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:48.927406Z digest=sha256:fb91b788939a2ef2523720ad940e7cd795df4851ee5877df7ab49316125a1532

Observation 60b72b98-76cb-4212-9fe7-4542f11bd8b8 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Taming transformers for high-resolution image synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.253430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:48.934941Z digest=sha256:d10032968f95b05b6d71903cabba3c3efe74ad59a1b4c3bc404b3e2c4ecdb958

Observation 4347f416-1b71-4c11-ad51-97f1d229fca8 · outbound

This paper cites Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.940157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.940157Z digest=sha256:80885750d1d01203550cf9c9c33d5d3e94a41f7cf7dfc9d1fc037d8564169ecc

Observation 5b5eaa11-74a3-4209-a7a1-809a6a59c8fe · outbound

This paper cites World models for autonomous driving: An initial survey.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT World models for autonomous driving: An initial survey

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.235250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:48.950257Z digest=sha256:a8bdfba5023ef95a0184bc2d77dcaea925bf312fdff04cf6fcd897bd10549d63

Observation fb62c909-4854-4e03-9d28-c94ccdd7b717 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Dream to Control: Learning Behaviors by Latent Imagination

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.960535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.960535Z digest=sha256:12f0cb44eb971aa12077a0ea3f0ba3a195750d90a9f0be0f6bbf7f16dc72ed21

Observation 3b129139-b1e6-4793-8022-cd87baa51d96 · outbound

This paper cites Mastering Atari with Discrete World Models.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Mastering Atari with Discrete World Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.971614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.971614Z digest=sha256:34372f9a02e7091fbf0e19526c10e42e8b8600c8e4191bf9867f87675a3babfa

Observation 47cb3455-5fae-47df-87a5-c5cf99c30e47 · outbound

This paper cites Mastering Diverse Domains through World Models.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Mastering Diverse Domains through World Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.977956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.977956Z digest=sha256:7d02ab693664921bc70c208719fb8ce3051150e573490c415d2f8763bdf948ff

Observation 4f597b95-f6cd-4f37-a728-ac2544093686 · outbound

This paper cites Show me what and tell me how: Video synthesis via multimodal conditioning.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Show me what and tell me how: Video synthesis via multimodal conditioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.218513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:48.985048Z digest=sha256:31f336e0db0ad8521b6a28e15d93bb3491613d5ff10598dd48ef6b59c439b723

Observation af2682c2-60df-4882-a81e-e464c58594dc · outbound

This paper cites Model-Predictive Policy Learning with Uncertainty Regularization for Driving in Dense Traffic.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Model-Predictive Policy Learning with Uncertainty Regularization for Driving in Dense Traffic

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.990647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.990647Z digest=sha256:f3333b39d31096863966d1ebba1014b40b40453f45265060a4ec2fac60636664

Observation fdc8f408-ed22-440d-b429-40bccd627f93 · outbound

This paper cites Query-Key Normalization for Transformers.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Query-Key Normalization for Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:48.995839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:48.995839Z digest=sha256:1c3033d070f133f5d518a766bbd4d1dbb4fec5f37bf20a2837d30ac5a914c76e

Observation f2f0a016-8d43-46a0-a247-a30dfe0ec8c8 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT GAIA-1: A Generative World Model for Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.001732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.001732Z digest=sha256:728e2512de7a39fbbac9b632eb1c7755970e5338a1714980f1d4c7818ffe2857

Observation 9c164218-d3bf-464f-8bd1-e5af50137c5b · outbound

This paper cites Image-to-image translation with conditional adversarial net- works.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Image-to-image translation with conditional adversarial net- works

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.007117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.007117Z digest=sha256:9080a815ff2b9aa91355e5894254be437e15812f131af0c81e4d2994f2c8a4d6

Observation f3e2005e-9778-4319-a4de-7751335f7636 · outbound

This paper cites ADriver-I: A General World Model for Autonomous Driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT ADriver-I: A General World Model for Autonomous Driving

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.011945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.011945Z digest=sha256:9acdeedcb332fe695a71a9bf9d22abc357d8e6034d2a5fe4b5410cb3e1f43248

Observation 94589463-b9cc-4b15-a39d-b98587ee2f30 · outbound

This paper cites Drivegan: Towards a controllable high-quality neural simulation.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Drivegan: Towards a controllable high-quality neural simulation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.186676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.016523Z digest=sha256:14a0bfc8890a8384596833588361d20d5c35ba7e376c775ef6c56e118ec8a14c

Observation 0dd486fe-edca-41be-b031-b7417d8a22e4 · outbound

This paper cites Adam: A method for stochastic optimization.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Adam: A method for stochastic optimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.169142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.025369Z digest=sha256:f5eb1628c39e57cfa289cb2d61e65cb93bd82a7d44a9578b816b57d07665af65

Observation eb052aa0-1ab7-44bb-8395-385fc26a2b3f · outbound

This paper cites The open im- ages dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT The open im- ages dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.151917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.031890Z digest=sha256:feb679bcb38839370a53d4dea73db1fbf0d6f04dc3c51486db72151d8cd59ca0

Observation 0203f2ee-0557-4ca3-a464-a9e94489fbe7 · outbound

This paper cites Fast and accurate image super-resolution with deep laplacian pyramid networks.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Fast and accurate image super-resolution with deep laplacian pyramid networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.132811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.040380Z digest=sha256:8bb33c15fbae321aca4a4019a8cb338c9faab166abde13466e93dd01f7371637

Observation e70f41e2-101b-4fa9-8ade-c5c4feecafbc · outbound

This paper cites A path towards autonomous machine intelli- gence version 0.9.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT A path towards autonomous machine intelli- gence version 0.9

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.047118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.047118Z digest=sha256:57d7e77c5e5c31511b75d76fa78924b45118f3561eaa20401f9f07164631f7f3

Observation d7f4f25e-64b5-42e8-8667-a858c733d8fc · outbound

This paper cites Microsoft coco: Common objects in context.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Microsoft coco: Common objects in context

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.101344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.056740Z digest=sha256:c8440512183a9cacb1d23f9cd39f2e3c7e99ee42d76fc6287e3abfdbb142ee8e

Observation b872b970-e03f-48ea-89c4-be0692140a87 · outbound

This paper cites Decoupled Weight Decay Regularization.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.064106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.064106Z digest=sha256:6412e901771b1774cd87056b39c9f3763f08035c707f1f525897298500b33d41

Observation 97c13c3f-72e9-4ca0-aff5-7f3e74e8eb5c · outbound

This paper cites WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.071699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.071699Z digest=sha256:e1d616abed2aea884549e8675b03cbd105ea901e39799a7fb57370435c98387b

Observation f6acfe43-03fe-419f-8fc6-d8550ea53f1b · outbound

This paper cites Improving language understanding by genera- tive pre-training.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Improving language understanding by genera- tive pre-training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.077752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.077711Z digest=sha256:09a4fa91271232d5440ccffa7b2dbeb035dd58dcecef0fd51008ceca69f165d3

Observation 245b9cb9-d6be-4a48-821e-3647adc481eb · outbound

This paper cites Language models are unsuper- vised multitask learners.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Language models are unsuper- vised multitask learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.057463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.084622Z digest=sha256:dd446617309a9a8f468687f376bea416dd13e5b220f0139171040ab108d141dc

Observation 0d410798-ec2a-4296-b5dd-a849c578c43a · outbound

This paper cites Learning a Driving Simulator.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Learning a Driving Simulator

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.092154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.092154Z digest=sha256:eae9fa12d30dbc8e0f587a7d5f6dd11975404c951a9af53d87e5bc8b3c5e0dda

Observation 5813e40e-caf5-4355-9cf2-b440bdb9dbe5 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.097879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.097879Z digest=sha256:f0618d6c172717dfe17e4b3291360176d0d4a6fd8177c9f4d4cb9f35ca654def

Observation 2c9ee85b-d917-4ed7-b298-236e1382c618 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Roformer: Enhanced transformer with rotary position embedding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.105194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.105194Z digest=sha256:625a244610b48601bdb35da1293575714fc7bec760bd17dc349c50c71a9529a6

Observation af69369a-118b-42f5-a1ee-7ba031ac8e4c · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.111282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.111282Z digest=sha256:6d58e8ec5d2420fa305bf74f79c2bde3f297346f8120265f17798b9c937374f7

Observation 441621c2-8bf1-4738-869a-8cfc14e27081 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.117719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.117719Z digest=sha256:1e068e669b07df00366337f27bc0df0e1f125bc97c344da7b01c12d905bcb138

Observation 7392c75f-c6b8-447b-8ece-d7b0f381cfb1 · outbound

This paper cites A Good Image Generator Is What You Need for High-Resolution Video Synthesis.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT A Good Image Generator Is What You Need for High-Resolution Video Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.125257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.125257Z digest=sha256:47e37385664842a1cec8e6b7f18e3ad40196030ef5c94006aea19ac913d9ac76

Observation 3119c715-45ee-44a2-89e1-01de51f3409d · outbound

This paper cites Neural discrete representation learning.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Neural discrete representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:50.011147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.132585Z digest=sha256:1bc4d90b2585b2e098a7f06bbd92215db8ae19ed025a202f7be831377a5d334b

Observation 05547562-d343-4383-9d3b-f677ca994290 · outbound

This paper cites DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.139374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.139374Z digest=sha256:7264da33d6af8fa911d021470a9eb788d0e41fb41fbcbb68e7006152ad373476

Observation 5fd2c5a9-005a-47dd-8634-0aa9cfbc59d6 · outbound

This paper cites Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Au- tonomous Driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Au- tonomous Driving

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.995483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.147436Z digest=sha256:a4f571f20916c90ff0311ad4f133be9cd91e4c461a00809eb5ef9fb452ed01af

Observation 82deb782-5ead-4184-aaed-f3c1598552e7 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.153030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.153030Z digest=sha256:bb6ae6b6ddc466d60610297ccadadd37490c7a6c2123235fad383c1dcd7e5646

Observation fc217fea-bcd6-4692-8bcc-9955cd889d94 · outbound

This paper cites Daydreamer: World models for physical robot learning.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Daydreamer: World models for physical robot learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.964335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.159691Z digest=sha256:ed54bdc4aa78e33e877fbe518aef937cfca541969eee1c37c9fb8c23166db375

Observation fa749b3e-ccb2-47aa-8860-449b51f9b65a · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.164470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.164470Z digest=sha256:542b10d891201432f88c685ce4430ceaaa755271d62d23c374f8fcf4eb679fbd

Observation 22927ece-7dd2-4871-a19c-16689a281071 · outbound

This paper cites Generalized Predictive Model for Autonomous Driving.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Generalized Predictive Model for Autonomous Driving

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.938902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.171293Z digest=sha256:6c88cfe5982ef9a6d2c0bde2116eb2d900bd8770969331057f45dd6280fd6b3d

Observation 09a8872d-7e12-4e2e-961a-2a8860dfd52b · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Vector-quantized Image Modeling with Improved VQGAN

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.178281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.178281Z digest=sha256:233d5bf56d7012fd60f5d17054b4d42306779e74f14b3dea2a67fd9665de8d28

Observation ac5acc64-fbbf-492f-9ba1-5114455b2c36 · outbound

This paper cites Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:49.184250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:49.184250Z digest=sha256:ea83f4e5edb3d9ef93417d79b4d8b772ff5a45dbc07ff2bcb66172d874c6cafc

Observation 9aa6987b-fbd7-41ac-a37a-59132c2f0ad7 · outbound

This paper cites Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Learning to drive by watching youtube videos: Action-conditioned contrastive policy pretraining

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.923303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.189846Z digest=sha256:9be72a34728742962c5181454ceec1d5d9c85898d0f580b8d6fe31980a7d4e70

Observation 26c8db5e-159f-4ec0-9d42-a0b041bb72bc · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT The unreasonable effectiveness of deep features as a perceptual metric

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.905925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.194430Z digest=sha256:9888a613d1d216bc7b04201e07863783dd22b847a5fb4377ea885d52e2d886b2

Observation 3bfe43c7-265e-4d44-9d85-9eed6368f5af · outbound

This paper cites Movq: Modulating quantized vectors for high-fidelity image generation.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Movq: Modulating quantized vectors for high-fidelity image generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.888778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.199512Z digest=sha256:6d3cce68632b72c87862da157bbe9ef719600b4449099d13d463fce225a83aba

Observation 0162864f-c5ca-467e-96df-1293cdbce82c · outbound

This paper cites The authors believe that this work has small potential negative impacts.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT The authors believe that this work has small potential negative impacts

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.866768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.204381Z digest=sha256:72295c8fd5a2ffcbd82bb9f1107659a77cf709690488c0d4d422ed671e9071dc

Observation 7f7d23a9-99bb-4d9c-9e3c-c66df082f550 · outbound

This paper cites Due to limited GPU memory, each image is restricted to a resolution of 256 × 512.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT Due to limited GPU memory, each image is restricted to a resolution of 256 × 512

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.844782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.209863Z digest=sha256:0c301cc73d971dc156d5686ad18cdf88fe1482b8110c2e3538bbc2d7d804b78f

Observation d4d847be-07b0-4687-90bb-977aef0c73ec · outbound

This paper cites This dataset is automatically annotated using a state-of-the-art offline perception system.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT This dataset is automatically annotated using a state-of-the-art offline perception system

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.825507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.216091Z digest=sha256:64cf8b63eebfce2ab4c81d2965501a48d8994ec243b82021c8ff2a9139b60829

Observation a824a3a7-acd9-46c3-98db-f8f99313c4d5 · outbound

This paper cites The images are with size of 256 × 512 and tokenized into 512 tokens.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT The images are with size of 256 × 512 and tokenized into 512 tokens

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:49.806904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.222557Z digest=sha256:dd6ea0f6920af1e80d38fbc9256d5fdaed5619e7039b89ae41acd4024ee1d9eb

Observation 6f9985d5-488a-4f52-885a-7de8d1069b53 · outbound

This paper cites w/o AR” and “Ours.

DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT w/o AR” and “Ours

Reference 52

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T00:21:49.784312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T00:21:49.227847Z digest=sha256:31f7eefd31d9ec985f280621adcda76c9bda2f21538348aa9785ced878ae31fb

Pith citing papers

Observation 9acad089-1f21-497d-9a75-b60439efa4ba · inbound

ARCON: Advancing Auto-Regressive Continuation for Driving Videos cites this paper.

ARCON: Advancing Auto-Regressive Continuation for Driving Videos DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:11:52.549073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:11:52.549073Z digest=sha256:3a336ad4fe4a13acdf7a3925f0dbb459eb695a5a04c3a176d3605a985fb53f7f

Observation 7f7ba31e-8639-496b-8631-e77f823c5b90 · inbound

DriveGPT: Scaling Autoregressive Behavior Models for Driving cites this paper.

DriveGPT: Scaling Autoregressive Behavior Models for Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T12:18:48.098811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:18:48.098811Z digest=sha256:6c5483a3bef696d507f3072920780f8414bda750e59518fd4e786eea159460ad

Observation 4727ccc5-6884-4c1b-97e9-7ea6f12e52d5 · inbound

A Survey of World Models for Autonomous Driving cites this paper.

A Survey of World Models for Autonomous Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T18:31:52.474058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:31:52.474058Z digest=sha256:e2e82b280585b5b1904b02a36d48a958ddd86d95dabf11f39d77c868f657e95a

Observation 0db9d819-b5bb-41d5-8218-c317dea44932 · inbound

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment cites this paper.

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T17:50:59.797593Z digest=sha256:c8a05012f82346ce62023ab0aeb731347af19e0747c46032f0d898ea84fd09aa

Observation cd00c96b-81ea-4696-8e55-4298f1af5821 · inbound

ARTEMIS: Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving cites this paper.

ARTEMIS: Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:54:07.752159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:54:07.752159Z digest=sha256:9b418ad025cb61d43aa586a986105e80000f132e61c24cc859ecec7311c4c70d

Observation c3c39100-9998-42e8-96aa-a59ba53efb26 · inbound

PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth cites this paper.

PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:18:26.096340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:18:26.096340Z digest=sha256:430377d77b27e9d56a24d78906f0a87b0985e89c7cbd89953356b5584bbe3362

Observation b78d4b06-e9f4-4de9-acbc-71041aa5bb28 · inbound

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control cites this paper.

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:43.254703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:43.254703Z digest=sha256:7daaac2b4e79bee379fae6f0fd9b461ea4c9d1244865b58473f282820244dea3

Observation 82e4a2a7-0850-40ef-9200-49d35d6dd197 · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:12.935432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:12.935432Z digest=sha256:cc5fcfeacb96bdd981e93d9a2e35b21391d00a103dc7df52bdadfc9b207b70dd

Observation bb55e9a0-0669-4394-a68e-ee0fb00a8013 · inbound

Epona: Autoregressive Diffusion World Model for Autonomous Driving cites this paper.

Epona: Autoregressive Diffusion World Model for Autonomous Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:29:49.144575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:29:49.144575Z digest=sha256:ab595caeb3a356bdabfd39d65551c19853d18b19aff50839027e0d7736652698

Observation 65b2649b-0e03-4020-8598-f513f30aee3d · inbound

Depth Anything at Any Condition cites this paper.

Depth Anything at Any Condition DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:52:02.480908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:52:02.480908Z digest=sha256:91c3762a179fef9a796066561b1ab00e933e9d2683ff642395eeef2345119046

Observation b2409850-6265-41af-87aa-044ac4955e59 · inbound

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization cites this paper.

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:54.718781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:54.718781Z digest=sha256:391d59416c3a190bf9c5055034955dcceeefc78e996a5f2d9817d715e9ee0f8e

Observation c4fa05f2-f4f2-4f68-b7ba-c6ffff482753 · inbound

3D and 4D World Modeling: A Survey cites this paper.

3D and 4D World Modeling: A Survey DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T06:04:15.535672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:04:15.535672Z digest=sha256:de78258d2e5a0a120e1142a0960eaf9004b7494743871e6d889d068d24148ed5

Observation db6ed7d7-8fec-4f55-87cc-dcd438111853 · inbound

A Comprehensive Survey on World Models for Embodied AI cites this paper.

A Comprehensive Survey on World Models for Embodied AI DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:45.347912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:45.347912Z digest=sha256:f46283008f36e234358c2c21459af4f80ec7b01b5541087d9ee2e167a6c32453

Observation 4be31751-4d72-4bd0-8886-6c587a079584 · inbound

OmniNWM: Omniscient Driving Navigation World Models cites this paper.

OmniNWM: Omniscient Driving Navigation World Models DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:57:09.462192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:57:09.462192Z digest=sha256:afa72088c9fc9f6953f6694b2d370f9b0138b7ff08a0bbaed6208ace865b35ea

Observation d054fc5c-8512-4b6f-8410-dd03d7a4fa39 · inbound

Thinking Ahead: Foresight Intelligence in MLLMs and World Model cites this paper.

Thinking Ahead: Foresight Intelligence in MLLMs and World Model DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:41:57.783447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:41:57.783447Z digest=sha256:bdfa0c211f32bb26b272c1e4d8bc2d3f880f02e2bc21cdee8d3417b360a1c678

Observation 6e6cbc5b-8e79-4619-8c5a-1bd740c0b85f · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:ee3d3f2bd99bd5c517dbed5d15cde8e744cf030829617973080fa9f2594b9598

Observation fccf9409-c6f8-4df7-8646-f78bce711583 · inbound

Learning Vision-Language-Action World Models for Autonomous Driving cites this paper.

Learning Vision-Language-Action World Models for Autonomous Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:08:10.442655Z digest=sha256:aaca3820fc7872b6530c536b16abf012c664acc5799aac507a08f01982be10fa

Observation 88dd64db-3028-4efc-94c8-2045588a8501 · inbound

Artificial Intelligence for Modeling and Simulation of Mixed Automated and Human Traffic cites this paper.

Artificial Intelligence for Modeling and Simulation of Mixed Automated and Human Traffic DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:00:59.662003Z digest=sha256:675ed494c756d1fcfebf26029ab1d0d06ccd0ade9c3b7987abd11b581e7bbe24

Observation 1523f59a-96a8-4f5a-bc00-c1ea18730cce · inbound

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation cites this paper.

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T05:31:59.676725Z digest=sha256:8c0b180630de6f1eda0927706db01e5bbbb76485486a0cbbafcd8e2d484b2cbf

Observation 23b6a125-2a6d-4deb-a730-0d96c104108d · inbound

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving cites this paper.

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:48:36.717026Z digest=sha256:31de5b65f9b653c11d1cf39cb7891dc1faebafbd8e85eccc28eb6bb693c96db7

Observation 200f7029-45a7-4d53-903c-1e4a480634f1 · inbound

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving cites this paper.

CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T21:36:52.396245Z digest=sha256:64522bd1fff3f20cd03ca3608a1ce7a58db0a055d6bfaad3e0b4768fcada62db

Observation 63fe9d34-9637-40c3-81c9-f3479298f508 · inbound

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond cites this paper.

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T21:15:12.465970Z digest=sha256:c77d5ac56cc5660af1dd60040507e66ece708b1054e6e8b1ceb0f2418cc9ca56

Observation 724b52d1-eaac-4d39-9e19-98173b2f592c · inbound

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends cites this paper.

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 213

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:29:18.513507Z digest=sha256:b075d19d3e09671343ea7f2beff57c62277c23d3a0964e1cd28a48195f4d5cdb

Observation ef710abf-7490-4f43-9c2a-61978eb80f36 · inbound

Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning cites this paper.

Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-08-18T01:13:47.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T01:49:25.510681Z digest=sha256:2830102b95de95c6eefe4b848deacb0f104448dc12686651ad3325145321c964

Observation 68ca223f-0f8f-4af5-9abd-6e3415aa9a5d · inbound

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation cites this paper.

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T08:19:04.131379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:19:04.131379Z digest=sha256:afeb027b470872a79c66b4543e84f564ad3fa8fc0418f40975ed4427f0a62923

Observation 18800ef9-5faf-4b61-897d-cf4217d7db68 · inbound

OpenLongTail: Generative Scaling of Long-Tail Driving Data cites this paper.

OpenLongTail: Generative Scaling of Long-Tail Driving Data DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T01:26:27.220907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:26:27.220907Z digest=sha256:4258e9f1bce27fb13f0a9a1fb3c7f4276af4c57d04f2de9cdf10d1a71594e1d7

Observation 7d925026-4feb-4bd7-a003-b0f70351c1f7 · inbound

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation cites this paper.

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T02:55:46.103411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:55:46.103411Z digest=sha256:8a743e8ed65eec1e3f738a0d90d499dd9c429b8a4ea68884cf72ee4559647a33

Observation 8bd34162-89f3-46e2-a849-b45a45eed1cb · inbound

Orbis 2: A Hierarchical World Model for Driving cites this paper.

Orbis 2: A Hierarchical World Model for Driving DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:05:02.737118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:05:02.737118Z digest=sha256:c24bd8cd42063a5cc0eb776a32d96c1438b33060fc3c798be71ce48022a1d35c

Observation a95fcd14-a47f-445b-b5dc-410b0a3670d1 · inbound

Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features cites this paper.

Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:25:51.323654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:25:51.323654Z digest=sha256:d2615b82e747dd4eb6b7bd4f4686d0aa225e8718adb4cce1e2e098ccca7042ac

Observation 90a528d6-e066-47c5-a81c-0a10ad410c84 · inbound

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting cites this paper.

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T00:28:52.210464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:28:52.210464Z digest=sha256:8dd96183bef04c5b8d22b786d5f91f56cd30e1f4155b6f37b94226f1c4105960

Observation 848b8722-b994-4d37-80d0-8505559c48ed · inbound

How Can Driving World Models Do Counterfactual Prediction? cites this paper.

How Can Driving World Models Do Counterfactual Prediction? DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:58.427662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:40:58.427662Z digest=sha256:b8e663104d09757428702cf1433743600152a352897a8a7432d74c3a1b1deed5

Observation 671d015b-a27f-4d65-9d33-36c446c29423 · inbound

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives cites this paper.

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:23:02.078100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:23:02.078100Z digest=sha256:a977878c86b7edd9a89c9e1da93448047d11865beb5e0eb65632a15391dab446