Pith. sign in

REVIEW 2 major objections 5 minor 132 references

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read One video model generates four road futures and steers by them

desk verdict A solid, genuinely new interface for future-scene prediction in driving—one shared video expert generating four stream types and an action expert reading their latents—with planning results that survive scrutiny, though the structured-stream evaluation is partly circular against its own frozen teachers. read the letter →

arxiv 2608.03084 v1 pith:OCC73BNS submitted 2026-08-04 cs.CV

classification cs.CV
keywords videogenerationend-to-enddrivingfuturesceneunderstandingjointvideo-actionattentionflowmatchingtrajectoryplanningworldmodelmulti-streamprediction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

SUV asks whether a single pretrained video generator can serve as the shared predictor of everything a self-driving policy needs to know about the next four seconds: appearance (RGB), road semantics, relative geometry, and which object is which as it moves. The paper claims it can: one video expert, with no task-specific prediction heads, denoises all four streams natively, and a separate action expert that attends to their latent tokens during the same denoising pass produces the ego trajectory. The payoff would be a scaling path in which future-scene understanding grows by adding video targets rather than new decoders, heads, and losses. On NAVSIM-v2 with one front camera and no candidate-trajectory selection, the system scores 91.0 EPDMS on navtest and 36.9 on navhard, and controlled ablations link both structured future supervision and direct stream access to higher planning scores.

What carries the argument

The load-bearing machinery is the token-interaction mask inside joint video-action attention, together with the shared latent video space. The mask implements directed information flow: observation tokens read only the clean prefix, each future stream reads the prefix and itself, and the action expert reads the prefix, all four future streams, and its own trajectory tokens, and this mask is applied at every transformer block and denoising step. What makes it work is that all four streams share one latent video layout, one observation prefix, and one flow-matching schedule with a shifted timestep $\phi_\kappa(\rho)=\kappa\rho/(1+(\kappa-1)\rho)$, $\kappa=5$, so the pretrained video expert generates them natively; the action expert needs no decoded video, only intermediate latents.

What would settle it

Replace the future-stream latents fed to the action expert with shape-matched Gaussian noise at inference; if navhard EPDMS stays near 36.9, the claimed causal contribution of direct future-stream access is refuted.

Watch

Extended reading notes

Core claim

The central claim is that future scene understanding for end-to-end driving can be recast as video generation: the same video expert that predicts future RGB also predicts segmentation, relative depth, and instance tracks as three-channel video streams in a shared latent space, initialized from generative video pretraining and post-trained with flow matching. A Mixture-of-Transformers action expert shares the denoising loop; a token-interaction mask lets action queries attend to every future-stream token group while blocking cross-stream and action-to-future information flow. The paper's evidence is that this setting reaches 91.0 EPDMS on NAVSIM-v2 navtest and 36.9 on navhard with a single front camera and no trajectory-candidate selection, outperforming a broad set of recent methods, and that the multi-stream model matches or slightly exceeds RGB-only future quality. It further argues, via a 2x2 ablation, that adding segmentation/depth/track supervision raises EPDMS on navtest from 89.7 to 90.7 and on navhard from 30.5 to 32.8, and that giving the action expert direct access to future-stream latents raises it further to 91.0 and 36.9.

Load-bearing premise

Future-scene quality and the planning benefit attributed to structured futures are measured against the same frozen teacher models that generated the training targets, so systematic teacher errors would inflate both the reported stream quality and, potentially, the planning gain.

Editorial extensions

If this is right

  • Adding a new kind of future knowledge, such as drivable-area gradient or object velocities, reduces to rendering it as another video stream and adding a prompt, with no new visual head or decoder required.
  • Planning can happen entirely in latent space: the trajectory is denoised alongside the future streams, and the VAE decoder is needed only for evaluation or visualization.
  • The accuracy-latency trade-off is controllable through solver steps: one Euler step gives 89.8 EPDMS on navtest and 33.0 on navhard at 177 ms on an RTX 4090, while ten steps reach 36.9 on navhard at 1356 ms.
  • Both structured future supervision and direct future-stream access improve planning, with the larger gains on the long-tail navhard split.
  • Removing access to any single structured stream lowers navhard EPDMS, with track removal hurting Stage 2 most; no single stream alone explains the gain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • (Editorial inference) If the shared-expert design transfers beyond roads, a single generative backbone could become the common future-predictor for other embodied tasks, such as manipulation or navigation, where semantic, geometric, and instance futures are rendered as color-coded streams.
  • (Editorial inference) Because the structured metrics judge agreement with the frozen teachers that built the targets, an independent-ground-truth evaluation could reorder the comparison between native generation and generate-then-perceive.
  • (Editorial inference) The per-stream access ablations suggest a cheaper deployment recipe: generate only RGB and read its latents while dropping the other streams at inference, which the paper's own numbers predict would cost little on navtest but several points on navhard.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper proposes SUV, an end-to-end driving framework that casts future scene understanding as video generation. A single video expert initialized from Wan2.2-5B is post-trained to generate four future streams (RGB, semantic segmentation, relative depth, instance tracks) as native videos, without stream-specific visual prediction heads. A separate action expert attends to the latent tokens of all streams during joint denoising to produce a trajectory. The model is evaluated on NAVSIM-v2 (91.0 EPDMS on navtest, 36.9 on navhard) and WOD-E2E (RFS 7.94) with a single front camera and no candidate selection, along with ablations isolating the contributions of generative initialization, structured supervision, and future-stream access. The paper also provides six-run reproducibility checks and detailed supplementary documentation of target construction and training.

Significance. If the results hold, SUV offers a conceptually clean demonstration that one pretrained video generator can serve as a shared predictor for heterogeneous future scene signals and that the resulting future-stream latents can improve trajectory planning. The external planning benchmark results are competitive or state-of-the-art, the experimental design includes controlled 2x2 ablations and per-stream access ablations, and the paper ships code and reproducibility details. The principal weakness is that the future-scene evaluation is self-referential: segmentation, depth, and tracking references are produced by the same frozen SAM3 and DA3 models that generated the training targets, so the reported mIoU, δ1, AbsRel, and AssA@50 values measure teacher-agreement rather than independent scene understanding. This gap is explicitly acknowledged in the paper but remains unresolved, and it weakens the headline 'future scene understanding' contribution.

major comments (2)
  1. [Future-scene evaluation (main text) and Supplementary §4.1] The structured-stream metrics are computed against references generated by the same frozen SAM3 (segmentation, tracks) and DA3 (depth) models that produced the training targets, as the paper explicitly acknowledges: 'The segmentation, tracking, and depth metrics measure agreement with frozen teachers rather than accuracy against independent ground truth.' This circularity is load-bearing because the paper's headline contribution is future scene understanding, and the evidence for it is currently only teacher-agreement. Please either (a) add an independent evaluation subset (e.g., human annotations or a different off-the-shelf model for a sample of frames) and report the same metrics on that subset, or (b) explicitly reframe the future-scene claims as 'agreement with frozen teacher models' rather than 'accuracy' or 'understanding'. The Table 5 comparison between native generation and Generate-then-Perceive is useful for task-alignment but cannot detect teacher-specific artifacts because both arms use the same teacher-generated references.
  2. [Planning Benefits of Future Representations, Table 6] The 2x2 ablation shows that adding segmentation/relative-depth/instance-track supervision (S/G/I) improves EPDMS on both navtest and navhard. Since the structured targets come from the same teachers used in the future-scene evaluation, this improvement could partly reflect the planner exploiting teacher-specific artifacts rather than general scene structure. The planning scores themselves are external and therefore not invalidated, but the interpretation of the improvement as evidence for 'future scene understanding' inherits the circularity concern. Please address this by either adding a control with targets from a different teacher on a subset (e.g., a different depth or segmentation model) or explicitly discussing the risk and tempering the corresponding conclusion.
minor comments (5)
  1. [Introduction / Contributions] The novelty claim ('the first end-to-end driving framework ...') should be cross-checked against the most recent concurrent work on generative world models for driving (e.g., EponaV2, GeoSem-WAM, WAM4D). A sentence explaining how SUV specifically differs in the format of the predicted streams (native video vs. task-specific heads) would strengthen the claim.
  2. [Planning Benefits of Future Representations] The abbreviation 'S/G/I' is used in the text without definition; it is only defined in the Table 6 caption. Please define it at first use.
  3. [Supplementary §4.3] 'Rankr seeds Python, NumPy, PyTorch, and CUDA' appears to be a typo; it should read 'rank r seeds ...'.
  4. [Table 4] The sentence 'Wan2.2-5B initialization yields better point estimates for all five metrics' is correct, but the large gaps (e.g., mIoU 64.2 vs 49.7) would benefit from a note on whether these differences are consistent across repeated training runs, or at least a statement that the six-run reproducibility checks were only reported for planning scores.
  5. [Method, Eq. (3)] The notation A(g) is slightly confusing because it is a set of visible key groups. Consider defining it explicitly as a function from a token group to the set of allowed key groups to improve readability.

Circularity Check

2 steps flagged · score 6.0 of 10

Future-scene metrics are scored against the same frozen SAM3/DA3 teachers that generated the training targets, so the reported mIoU, AssA@50, δ1, and AbsRel measure teacher imitation rather than independent scene understanding; the paper explicitly acknowledges this caveat but the self-reference remains unresolved.

  1. self definitional [Experiments, Future-scene evaluation; Supplementary Material §4.1]
    "Frozen teachers convert each recorded future RGB clip into videos for semantic segmentation, relative depth, and instance tracks. ... These videos serve as training targets for both benchmarks and as evaluation references for the NAVSIM future-scene analysis. ... The segmentation, tracking, and depth metrics measure agreement with frozen teachers rather than accuracy against independent ground truth."

    The evaluation references for mIoU (segmentation), AssA@50 (tracks), and δ1/AbsRel (depth) are produced by the same frozen SAM3 and DA3 models that generated the training targets. The model is trained with MSE to reproduce those teacher labels, and the reported 'future scene understanding' scores then measure agreement with those same labels. Perfectly memorizing the training targets would, by construction, yield near-perfect teacher-agreement scores, so these numbers cannot validate the structured streams as independent scene understanding. Table 5 is also internal to this loop: Generate-then-Perceive applies the same teachers to generated RGB, and native generation is scored against the same teacher references, so the comparison only ranks which pipeline best reproduces the teachers.

  2. fitted input called prediction [Supplementary Material §1.2, Eq. (4)–(5); Supplementary Material §4.1]
    "At evaluation, each generated depth pixel is mapped to its nearest Turbo entry. ... For NAVSIM future-scene evaluation, Equation 4 fits one affine map from these normalized values to the DA3-reference scale over all eight frames. ... We report δ1 and AbsRel as affine-aligned, percentile-clipped measures of agreement with the DA3 teacher."

    The depth training target is DA3's clip-level percentile-normalized, Turbo-quantized output, and the depth evaluation reference is DA3 applied to the recorded RGB frames. Before computing δ1 and AbsRel, the prediction is affinely aligned to that same DA3 reference, removing global scale and shift errors—among the most basic geometric errors a depth predictor can make. The metric then rewards reproduction of the teacher's normalization and color mapping after optimal reparameterization.

full rationale

The paper's headline planning results are evaluated on external benchmarks (NAVSIM-v2 EPDMS and WOD-E2E RFS) and are not circular; the trajectory scores come from an independent protocol with no fitted parameters tied to the future-stream evaluation. The circularity is confined to the future-scene understanding contribution. By the paper's own description, frozen SAM3 and DA3 generate the segmentation, relative-depth, and instance-track training targets, and the same frozen teachers provide the evaluation references for mIoU, AssA@50, δ1, and AbsRel. The structured-stream metrics therefore measure how faithfully the video expert reproduces its teacher-derived supervision, not whether the predicted streams are accurate against independent ground truth. The paper acknowledges this plainly ('measure agreement with frozen teachers rather than accuracy against independent ground truth'), but an acknowledgement does not break the construction: the evaluation loop is closed by design. Because the central novelty is precisely that the generated structured streams constitute future scene understanding, the self-referential evaluation prevents the structured-stream results from independently validating that claim. The RGB quality metrics and the external planning numbers remain independent; however, the integrated claim that structured future supervision and future-stream access improve planning inherits some risk if the planner exploits teacher-specific artifacts. This is partial circularity, not total: score 6, with the non-circular external planning results as the mitigating factor.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central results depend on several hand-chosen design parameters (timestep mapping, scheduler shape, loss weights, decoding thresholds) and on the validity of pretrained video generators and frozen teacher models. No independent evidence for teacher accuracy is provided, and the planning benchmark is treated as ground truth without critical examination.

free parameters (6)
  • Shifted flow-time mapping kappa = 5
    Eq. 12 defines phi_kappa(rho) = kappa*rho/(1+(kappa-1)*rho) with kappa=5; hand-chosen to bias training timestep sampling and used for the Euler grid at inference. No sensitivity analysis is provided.
  • Scheduler weight shape exponent = 2 in q(lambda) = exp(-2(lambda-1/2)^2) - exp(-1/2)
    Eq. 14 in the supplement; the quadratic exponent and normalization are hand-chosen weighting for the flow-matching loss.
  • Action-vs-video loss weight = 1
    Eq. 5 and supplementary text: equal coefficients of 1 for the averaged four-stream video loss and the action loss; chosen without sensitivity analysis.
  • Depth normalization quantiles = q_0.01 and q_0.99
    Supplementary Section 4.1: relative depth is clipped to 1st and 99th percentiles per clip before quantization; choice affects depth encoding and evaluation.
  • Instance-track decoding thresholds = distance 30, min size 32, IoU/centroid cost weights
    Supplementary Section 4.1: decoding predicted instance videos uses a Euclidean distance threshold of 30 from black, 32-pixel minimum component size, and Hungarian assignment cost combining mask IoU and centroid displacement; these hand-chosen values directly affect AssA@50.
  • SAM3 target-construction thresholds = confidence 0.5, 32 detections per prompt, 128 detection cap
    Supplementary Section 4.1: frozen SAM3 target generation uses confidence threshold 0.5, up to 32 detections per prompt, and a 128-detection frame cap; these determine segmentation and tracking targets.
assumptions (5)
  • standard math Flow-matching training with the scored velocity estimator converges and shifted Euler inference is numerically stable.
    Invoked implicitly in Section 'Multi-stream flow matching' and 'Joint inference'; no convergence proof is given, which is standard for this empirical setting.
  • domain assumption Pretrained Wan2.2-5B provides a useful generative prior for jointly predicting all four future streams in driving scenes.
    The method initializes the video expert from Wan2.2-5B (Section 'Shared latent video space'). Table 4 tests this partially by comparing against random initialization, but the broader claim of transfer to structured streams is assumed.
  • domain assumption Frozen SAM3 and DA3 outputs provide reliable supervision and valid references for future semantics, relative depth, and instance tracks.
    Structured targets are built from these teachers (Section 'Structured target construction') and the same teachers define evaluation references (Section 'Future-scene evaluation'). The paper explicitly notes this is not independent ground truth.
  • domain assumption The masked attention scheme in Eq. 3, blocking cross-stream interaction and action-to-future feedback, is a suitable inductive bias.
    The paper motivates the mask by preserving pretrained attention structure, but no ablation on alternative masks is provided, so the optimality of this specific information-flow restriction is assumed.
  • domain assumption NAVSIM-v2 EPDMS with human-penalty filtering is an appropriate proxy for autonomous driving performance.
    The benchmark protocol is adopted from prior work (Dauner et al., Cao et al.) and used for the central planning claims; its validity as a safety metric is not examined.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SUV: Future Scene Understanding as Video Generation for End-to-End Driving." pith.science (2026). https://pith.science/paper/OCC73BNS

@misc{pith2026260803084,
  author       = {Pith},
  title        = {Pith review of: SUV: Future Scene Understanding as Video Generation for End-to-End Driving},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OCC73BNS}},
  note         = {Machine review of arXiv:2608.03084}
}
read the original abstract

End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.

Figures

Figures reproduced from arXiv: 2608.03084 by the authors.

Figure 1
Figure 1. Future scene understanding for end-to-end driving. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of SUV. Left: the shared video expert generates four future streams, while the action expert denoises the trajectory using their latent tokens. Right: the token-interaction mask preserves observation-to-stream pathways, blocks cross-stream interaction and action-to-future feedback, and lets action queries read every token group. The label Obs. denotes observation tokens encoded from the input RGB video. Pro… view at source ↗
Figure 3
Figure 3. Qualitative multi-stream prediction and planning in a lane-change scenario. Left: Past video and planning. Right: The [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 1
Figure 1. Figure 1: Qualitative zero-shot planning on two in-house [PITH_FULL_IMAGE:figures/full_fig_p012_1.png]
Figure 2
Figure 2. Figure 2: High-curvature turns on NAVSIM-v2 navtest. Each case shows the current observation, bird’s-eye-view trajectories, [PITH_FULL_IMAGE:figures/full_fig_p013_2.png]
Figure 3
Figure 3. Figure 3: Lower progress in dense urban traffic. The [PITH_FULL_IMAGE:figures/full_fig_p014_3.png]
Figure 5
Figure 5. Figure 5: Horizon-wise RGB metrics for multi-stream (Uni [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 4
Figure 4. Figure 4: Horizon-wise teacher-agreement metrics for native [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

132 extracted references · 55 canonical work pages

  1. [1]

    Luiten, Jonathon and Osep, Aljosa and Dendorfer, Patrick and Torr, Philip H. S. and Geiger, Andreas and Leal-Taix. 2021 , doi=

  2. [2]

    2026 , eprint=

    Image Generators are Generalist Vision Learners , author=. 2026 , eprint=

  3. [3]

    2026 , eprint=

    Video Generation Models are General-Purpose Vision Learners , author=. 2026 , eprint=

  4. [4]

    2026 , eprint=

    Vision as Unified Multimodal Generation , author=. 2026 , eprint=

  5. [5]

    Carion, Nicolas and Gustafson, Laura and Hu, Yuan-Ting and Debnath, Shoubhik and Hu, Ronghang and Suris Coll-Vinent, Didac and Ryali, Chaitanya and Alwala, Kalyan Vasudev and Khedr, Haitham and Huang, Andrew and Lei, Jie and Ma, Tengyu and Guo, Baishan and Kalla, Arpit and Marks, Markus and Greer, Joseph and Wang, Meng and Sun, Peize and Rädle, Roman and ...

  6. [6]

    2026 , url=

    Depth Anything 3: Recovering the Visual Space from Any Views , author=. 2026 , url=

  7. [7]

    2025 , issn=

    Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models , author=. 2025 , issn=

  8. [8]

    Jiang, Bo and Chen, Shaoyu and Xu, Qing and Liao, Bencheng and Chen, Jiajie and Zhou, Helong and Zhang, Qian and Liu, Wenyu and Huang, Chang and Wang, Xinggang , booktitle=ICCV, year=

Show all 132 references
  1. [9]

    Sun, Wenchao and Lin, Xuewu and Shi, Yining and Zhang, Chuang and Wu, Haoran and Zheng, Sifa , booktitle=ICRA, year=

  2. [10]

    Planning-

    Hu, Yihan and Yang, Jiazhi and Chen, Li and Li, Keyu and Sima, Chonghao and Zhu, Xizhou and Chai, Siqi and Du, Senyao and Lin, Tianwei and Wang, Wenhai and Lu, Lewei and Jia, Xiaosong and Liu, Qiang and Dai, Jifeng and Qiao, Yu and Li, Hongyang , booktitle=CVPR, year=. Planning-

  3. [11]

    Tian, Xiaoyu and Gu, Junru and Li, Bailin and Liu, Yicheng and Wang, Yang and Zhao, Zhiyong and Zhan, Kun and Jia, Peng and Lang, XianPeng and Zhao, Hang , booktitle=CORL, year=

  4. [12]

    Hwang, Jyh-Jing and Xu, Runsheng and Lin, Hubert and Hung, Wei-Chih and Ji, Jingwei and Choi, Kristy and Huang, Di and He, Tong and Covington, Paul and Sapp, Benjamin and Zhou, Yin and Guo, James and Anguelov, Dragomir and Tan, Mingxing , journal=TMLR, year=

  5. [13]

    , booktitle=CVPR, year=

    Wang, Shihao and Yu, Zhiding and Jiang, Xiaohui and Lan, Shiyi and Shi, Min and Chang, Nadine and Kautz, Jan and Li, Ying and Alvarez, Jose M. , booktitle=CVPR, year=

  6. [14]

    Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in

    Zhai, Jiang-Tian and Feng, Ze and Du, Jinhao and Mao, Yongqiang and Liu, Jiang-Jiang and Tan, Zichang and Zhang, Yifu and Ye, Xiaoqing and Wang, Jingdong , year=. Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in. 2305.10430 , archivePrefix=

  7. [15]

    2024 , doi=

    Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? , author=. 2024 , doi=

  8. [17]

    Dauner, Daniel and Hallgarten, Marcel and Li, Tianyu and Weng, Xinshuo and Huang, Zhiyu and Yang, Zetong and Li, Hongyang and Gilitschenski, Igor and Ivanovic, Boris and Pavone, Marco and Geiger, Andreas and Chitta, Kashyap , booktitle=NeurIPS, year=

  9. [18]

    2025 , url=

    Pseudo-Simulation for Autonomous Driving , author=. 2025 , url=

  10. [19]

    Xu, Runsheng and Lin, Hubert and Jeon, Wonseok and Feng, Hao and Zou, Yuliang and Sun, Liting and Gorman, John and Tolstaya, Kate and Tang, Sarah and White, Brandyn and Sapp, Ben and Tan, Mingxing and Hwang, Jyh-Jing and Anguelov, Dragomir , booktitle=CVPR, year=

  11. [20]

    2026 , eprint=

    World Model for Robot Learning: A Comprehensive Survey , author=. 2026 , eprint=

  12. [21]

    2022 , doi=

    Model-Based Imitation Learning for Urban Driving , author=. 2022 , doi=

  13. [23]

    Wang, Xiaofeng and Zhu, Zheng and Huang, Guan and Chen, Xinze and Zhu, Jiagang and Lu, Jiwen , booktitle=ECCV, year=

  14. [24]

    2024 , doi=

    Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving , author=. 2024 , doi=

  15. [25]

    2024 , doi=

    Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability , author=. 2024 , doi=

  16. [26]

    Chen, Yuntao and Wang, Yuqi and Zhang, Zhaoxiang , booktitle=ICCV, year=

  17. [27]

    2024 , eprint=

    Doe-1: Closed-Loop Autonomous Driving with Large World Model , author=. 2024 , eprint=

  18. [28]

    Wei, Julong and Yuan, Shanshuai and Li, Pengfei and Quan, Xinyi and Tai, Lei and Zhao, Jieru and Gan, Zhongxue and Ding, Wenchao , booktitle=ICRA, year=

  19. [29]

    Zheng, Wenzhao and Chen, Weiliang and Huang, Yuanhui and Zhang, Borui and Duan, Yueqi and Lu, Jiwen , booktitle=ECCV, year=

  20. [30]

    2025 , doi=

    Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving , author=. 2025 , doi=

  21. [31]

    2024 , doi=

    Visual Point Cloud Forecasting enables Scalable Autonomous Driving , author=. 2024 , doi=

  22. [32]

    Min, Chen and Zhao, Dawei and Xiao, Liang and Zhao, Jian and Xu, Xinli and Zhu, Zheng and Jin, Lei and Li, Jianshu and Guo, Yulan and Xing, Junliang and Jing, Liping and Nie, Yiming and Dai, Bin , booktitle=CVPR, year=

  23. [33]

    Enhancing End-to-End Autonomous Driving with Latent World Model , author=

  24. [34]

    Li, Yingyan and Shang, Shuyao and Liu, Weisong and Zhan, Bing and Wang, Haochen and Wang, Yuqi and Chen, Yuntao and Wang, Xiaoman and An, Yasong and Tang, Chufeng and Hou, Lu and Fan, Lue and Zhang, Zhaoxiang , booktitle=ICLR, year=

  25. [35]

    doi:10.1109/LRA.2026.3678836 , url=

    Wozniak, Maciej and Liu, Lianhang and Cai, Yixi and Jensfelt, Patric , journal=RAL, year=. doi:10.1109/LRA.2026.3678836 , url=

  26. [39]

    Liang, Dingkang and Zhang, Dingyuan and Zhou, Xin and Tu, Sifan and Feng, Tianrui and Li, Xiaofan and Zhang, Yumeng and Du, Mingyang and Tan, Xiao and Bai, Xiang , booktitle=ICRA, year=

  27. [41]

    Li, Bohan and Ma, Zhuang and Du, Dalong and Peng, Baorui and Liang, Zhujin and Liu, Zhenqiang and Guo, Xianda and Zhu, Zheng and Ma, Chao and Jin, Yueming and Jin, Xin and Zhao, Hao and Zeng, Wenjun , booktitle=ECCV, year=

  28. [42]

    2024 , doi=

    Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation , author=. 2024 , doi=

  29. [43]

    What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? , author=

  30. [44]

    Zhao, Canyu and Sun, Yanlong and Liu, Mingyu and Zheng, Huanyi and Zhu, Muzhi and Zhao, Zhiyue and Chen, Hao and He, Tong and Shen, Chunhua , booktitle=NeurIPS, year=

  31. [45]

    2603.25892 , archivePrefix=

    Wang, Letian and Zanfir, Andrei and Bazavan, Eduard Gabriel and Andriluka, Misha and Sminchisescu, Cristian , year=. 2603.25892 , archivePrefix=

  32. [46]

    Huang, Wenhui and Zhang, Songyan and Huang, Qihang and Wang, Zhidong and Mao, Zhiqi and Chua, Collister and Chen, Zhan and Chen, Long and Lv, Chen , booktitle=ICML, year=

  33. [48]

    Bi, Hongzhe and Tan, Hengkai and Xie, Shenghao and Wang, Zeyuan and Huang, Shuhe and Liu, Haitian and Zhao, Ruowen and Feng, Yao and Xiang, Chendong and Rong, Yinze and Zhao, Hongyan and Liu, Hanyu and Su, Zhizhong and Ma, Lei and Su, Hang and Zhu, Jun , booktitle=CVPR, year=

  34. [51]

    Li, Jingyu and Zhang, Bozhou and Jin, Xin and Deng, Jiankang and Zhu, Xiatian and Zhang, Li , booktitle=ICRA, year=

  35. [53]

    Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution , author=

  36. [55]

    Dang, Chenxu and Ang, Sining and Li, Yongkang and Tian, Haochen and Wang, Jie and Li, Guang and Ye, Hangjun and Ma, Jie and Chen, Long and Wang, Yan , booktitle=ECCV, year=

  37. [56]

    2605.15120 , archivePrefix=

    Ang, Sining and Yang, Yuguang and Chen, Canyu and Wang, Yan , year=. 2605.15120 , archivePrefix=

  38. [57]

    Liu, Lin and Jia, Caiyan and Yu, Guanyi and Song, Ziying and Li, Junqiao and Jia, Feiyang and Wu, Peiliang and Hao, Xiaoshuai and Luo, Yadan , booktitle=CVPR, year=

  39. [58]

    2025 , doi=

    Learning 4D Embodied World Models , author=. 2025 , doi=

  40. [59]

    Xiong, Zhexiao and Ye, Xin and Yaman, Burhan and Cheng, Sheng and Lu, Yiren and Luo, Jingru and Jacobs, Nathan and Ren, Liu , booktitle=ECCV, year=

  41. [61]

    Sheng, Zihao and Ye, Xin and Luo, Jingru and Chen, Sikai and Ren, Liu , booktitle=ECCV, year=

  42. [62]

    Jia, Feiyang and Liu, Lin and Song, Ziying and Jia, Caiyan and Ye, Hangjun and Hao, Xiaoshuai and Chen, Long , booktitle=ICML, year=

  43. [63]

    Li, Jingyu and Wu, Junjie and Hu, Dongnan and Huang, Xiangkai and Sun, Bin and Hao, Zhihui and Lang, Xianpeng and Zhu, Xiatian and Zhang, Li , booktitle=CVPR, year=

  44. [66]

    and Wu, Zuxuan , booktitle=AAAI, year=

    Yao, Wenhao and Li, Zhenxin and Lan, Shiyi and Wang, Zi and Sun, Xinglong and Alvarez, Jose M. and Wu, Zuxuan , booktitle=AAAI, year=

  45. [68]

    Li, Yongkang and Xiong, Kaixin and Guo, Xiangyu and Li, Fang and Yan, Sixu and Xu, Gangwei and Zhou, Lijun and Chen, Long and Sun, Haiyang and Wang, Bing and Ma, Kun and Chen, Guang and Ye, Hangjun and Liu, Wenyu and Wang, Xinggang , booktitle=ICLR, year=

  46. [69]

    Chitta, Kashyap and Prakash, Aditya and Jaeger, Bernhard and Yu, Zehao and Renz, Katrin and Geiger, Andreas , journal=PAMI, year=

  47. [70]

    Liao, Bencheng and Chen, Shaoyu and Yin, Haoran and Jiang, Bo and Wang, Cheng and Yan, Sixu and Zhang, Xinbang and Li, Xiangyu and Zhang, Ying and Zhang, Qian and Wang, Xinggang , booktitle=CVPR, year=

  48. [71]

    Wang, Jie and Li, Guang and Huang, Zhijian and Dang, Chenxu and Ye, Hangjun and Han, Yahong and Chen, Long , booktitle=CVPR, year=

  49. [72]

    Zhang, Jinqing and Fu, Zehua and Xu, Zelin and Dai, Wenying and Liu, Qingjie and Wang, Yunhong , booktitle=ICLR, year=

  50. [73]

    Wang, Junli and Zheng, Yinan and Liu, Xueyi and Xing, Zebin and Li, Pengfei and Ma, Kun and Ye, Hangjun and Chen, Guang and Li, Guang and Chen, Long and Xia, Zhongpu and Zhang, Qichao , booktitle=CVPR, year=

  51. [74]

    Zhou, Zewei and Cai, Tianhui and Zhao, Seth and Zhang, Yun and Huang, Zhiyu and Zhou, Bolei and Ma, Jiaqi , booktitle=NeurIPS, year=

  52. [77]

    Ye, Yuqi and Zhang, Zijian and Lin, Junhong and Sun, Shangkun and Peng, Changhao and Gao, Wei , booktitle=ICLR, year=

  53. [78]

    Zheng, Yupeng and Yang, Pengxuan and Xing, Zebin and Zhang, Qichao and Zheng, Yuhang and Gao, Yinfeng and Li, Pengfei and Zhang, Teng and Xia, Zhongpu and Jia, Peng and Lang, Xianpeng and Zhao, Dongbin , booktitle=ICCV, year=

  54. [79]

    Zhang, Kaiwen and Tang, Zhenyu and Hu, Xiaotao and Pan, Xingang and Guo, Xiaoyang and Liu, Yuan and Huang, Jingwei and Yuan, Li and Zhang, Qian and Long, Xiao-Xiao and Cao, Xun and Yin, Wei , booktitle=ICCV, year=

  55. [80]

    From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction , author=

  56. [81]

    Xia, Tianze and Li, Yongkang and Zhou, Lijun and Yao, Jingfeng and Xiong, Kaixin and Sun, Haiyang and Wang, Bing and Ma, Kun and Chen, Guang and Ye, Hangjun and Liu, Wenyu and Wang, Xinggang , booktitle=CVPR, year=

  57. [82]

    Feng, Lan and Gao, Yang and Zablocki, Eloi and Li, Quanyi and Li, Wuyang and Liu, Sichao and Cord, Matthieu and Alahi, Alexandre , booktitle=ICLR, year=

  58. [83]

    2605.12624 , archivePrefix=

    Huang, Yuzhou and Zhu, Benjin and Lu, Hengtong and Huang, Victor Shea-Jay and Zhang, Haiming and Chen, Wei and Dai, Jifeng and Xie, Yan and Li, Hongsheng , year=. 2605.12624 , archivePrefix=

  59. [86]

    Luiten, Jonathon and Osep, Aljosa and Dendorfer, Patrick and Torr, Philip H. S. and Geiger, Andreas and Leal-Taix. IJCV , year=

  60. [87]

    Bi, H.; Tan, H.; Xie, S.; Wang, Z.; Huang, S.; Liu, H.; Zhao, R.; Feng, Y.; Xiang, C.; Rong, Y.; Zhao, H.; Liu, H.; Su, Z.; Ma, L.; Su, H.; and Zhu, J. 2026. Motus : A Unified Latent Action World Model. In CVPR

  61. [88]

    Cao, W.; Hallgarten, M.; Li, T.; Dauner, D.; Gu, X.; Wang, C.; Miron, Y.; Aiello, M.; Li, H.; Gilitschenski, I.; Ivanovic, B.; Pavone, M.; Geiger, A.; and Chitta, K. 2025. Pseudo-Simulation for Autonomous Driving. In CoRL

  62. [89]

    Carion, N.; Gustafson, L.; Hu, Y.-T.; Debnath, S.; Hu, R.; Suris Coll-Vinent, D.; Ryali, C.; Alwala, K. V.; Khedr, H.; Huang, A.; Lei, J.; Ma, T.; Guo, B.; Kalla, A.; Marks, M.; Greer, J.; Wang, M.; Sun, P.; Rädle, R.; Afouras, T.; Mavroudi, E.; Xu, K.; Wu, T.-H.; Zhou, Y.; Mo...

  63. [90]

    Chen, Y.; Wang, Y.; and Zhang, Z. 2025. DrivingGPT : Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers. In ICCV

  64. [91]

    Chitta, K.; Prakash, A.; Jaeger, B.; Yu, Z.; Renz, K.; and Geiger, A. 2023. TransFuser : Imitation With Transformer-Based Sensor Fusion for Autonomous Driving. IEEE TPAMI

  65. [92]

    Dang, C.; Ang, S.; Li, Y.; Tian, H.; Wang, J.; Li, G.; Ye, H.; Ma, J.; Chen, L.; and Wang, Y. 2026. DriveFine : Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving. In ECCV

  66. [93]

    Dauner, D.; Hallgarten, M.; Li, T.; Weng, X.; Huang, Z.; Yang, Z.; Li, H.; Gilitschenski, I.; Ivanovic, B.; Pavone, M.; Geiger, A.; and Chitta, K. 2024. NAVSIM : Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking. In NeurIPS

  67. [94]

    T.; Genova, K.; Kannen, N.; Ben, S.; Li, Y.; Guo, M.; Yogin, S.; Gu, Y.; Chen, H.; Wang, O.; Xie, S.; Zhou, H.; He, K.; Funkhouser, T.; Alayrac, J.-B.; and Soricut, R

    Gabeur, V.; Long, S.; Peng, S.; Voigtlaender, P.; Sun, S.; Bao, Y.; Truong, K.; Wang, Z.; Zhou, W.; Barron, J. T.; Genova, K.; Kannen, N.; Ben, S.; Li, Y.; Guo, M.; Yogin, S.; Gu, Y.; Chen, H.; Wang, O.; Xie, S.; Zhou, H.; He, K.; Funkhouser, T.; Alayrac, J.-B.; and Soricut, R...

  68. [95]

    Gao, S.; Yang, J.; Chen, L.; Chitta, K.; Qiu, Y.; Geiger, A.; Zhang, J.; and Li, H. 2024. Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability. In NeurIPS

  69. [96]

    Guo, J.; Li, Q.; Li, P.; Chen, Z.; Sun, N.; Su, Y.; Wang, H.; Zhang, Y.; Li, X.; and Liu, H. 2026. Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising. arXiv:2604.26694

  70. [97]

    Han, X.; Li, J.; Deng, K.; Chen, Z.; Shi, X.; Wang, S.; Li, B.; Wang, L.; Xie, S.; You, X.; Quan, J.; Cai, Z.; Diao, H.; Liu, Z.; Yang, L.; Lin, D.; and Wang, Q. 2026. Vision as Unified Multimodal Generation. arXiv:2607.06560

  71. [98]

    Hong, Y.; Zhou, X.; Li, Y.; Zhou, X.; Liu, L.; Luo, Y.; Xu, S.; Yang, L.; and Song, Z. 2026. DriveFuture : Future-Aware Latent World Models for Autonomous Driving. arXiv:2605.09701

  72. [99]

    Hou, B.; Li, G.; Jia, J.; An, T.; Guo, X.; Leng, S.; Geng, H.; Ze, Y.; Harada, T.; Torr, P.; Mees, O.; Pollefeys, M.; Liu, Z.; Wu, J.; Abbeel, P.; Malik, J.; Du, Y.; and Yang, J. 2026. World Model for Robot Learning: A Comprehensive Survey. arXiv:2605.00080

  73. [100]

    Hu, A.; Russell, L.; Yeo, H.; Murez, Z.; Fedoseev, G.; Kendall, A.; Shotton, J.; and Corrado, G. 2023 a . GAIA-1 : A Generative World Model for Autonomous Driving. arXiv:2309.17080

  74. [101]

    Hu, Y.; Yang, J.; Chen, L.; Li, K.; Sima, C.; Zhu, X.; Chai, S.; Du, S.; Lin, T.; Wang, W.; Lu, L.; Jia, X.; Liu, Q.; Dai, J.; Qiao, Y.; and Li, H. 2023 b . Planning- O riented Autonomous Driving. In CVPR

  75. [102]

    Huang, W.; Zhang, S.; Huang, Q.; Wang, Z.; Mao, Z.; Chua, C.; Chen, Z.; Chen, L.; and Lv, C. 2026. AutoMoT : A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving. In ICML

  76. [103]

    Hwang, J.-J.; Xu, R.; Lin, H.; Hung, W.-C.; Ji, J.; Choi, K.; Huang, D.; He, T.; Covington, P.; Sapp, B.; Zhou, Y.; Guo, J.; Anguelov, D.; and Tan, M. 2025. EMMA : End-to-End Multimodal Model for Autonomous Driving. TMLR

  77. [104]

    Jia, F.; Liu, L.; Song, Z.; Jia, C.; Ye, H.; Hao, X.; and Chen, L. 2026. DriveWorld-VLA : Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving. In ICML

  78. [105]

    Jiang, A.; Gao, Y.; Wang, Y.; Sun, Z.; Wang, S.; Heng, Y.; Sun, H.; Tang, S.; Zhu, L.; Chai, J.; Wang, J.; Gu, Z.; Jiang, H.; and Sun, L. 2025. IRL-VLA : Training an Vision-Language-Action Policy via Reward World Model. arXiv:2508.06571

  79. [106]

    Jiang, B.; Chen, S.; Liao, B.; Zhang, X.; Yin, W.; Zhang, Q.; Huang, C.; Liu, W.; and Wang, X. 2024. Senna : Bridging Large Vision-Language Models and End-to-End Autonomous Driving. arXiv:2410.22313

  80. [107]

    Jiang, B.; Chen, S.; Xu, Q.; Liao, B.; Chen, J.; Zhou, H.; Zhang, Q.; Liu, W.; Huang, C.; and Wang, X. 2023. VAD : Vectorized Scene Representation for Efficient Autonomous Driving. In ICCV

  81. [108]

    Li, J.; Liu, Z.; Hu, D.; Wu, J.; Ma, Z.; Wu, W.; Han, C.; Hao, Z.; Liu, Z.; Zhan, K.; Deng, J.; Zhu, X.; and Zhang, L. 2026 a . Metis : A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation. arXiv:2606.15869

  82. [109]

    Li, J.; Wu, J.; Hu, D.; Huang, X.; Sun, B.; Hao, Z.; Lang, X.; Zhu, X.; and Zhang, L. 2026 b . SGDrive : Scene-to-Goal Hierarchical World Cognition for Autonomous Driving. In CVPR

  83. [110]

    Li, J.; Zhang, B.; Jin, X.; Deng, J.; Zhu, X.; and Zhang, L. 2026 c . ImagiDrive : A Unified Imagination-and-Planning Framework for Autonomous Driving. In ICRA

  84. [111]

    Li, K.; Li, Z.; Lan, S.; Xie, Y.; Zhang, Z.; Liu, J.; Wu, Z.; Yu, Z.; and Alvarez, J. M. 2025 a . Hydra-MDP++ : Advancing End-to-End Driving via Expert-Guided Hydra-Distillation. arXiv:2503.12820

  85. [112]

    Li, Y.; Fan, L.; He, J.; Wang, Y.; Chen, Y.; Zhang, Z.; and Tan, T. 2025 b . Enhancing End-to-End Autonomous Driving with Latent World Model. In ICLR

  86. [113]

    Li, Y.; Shang, S.; Liu, W.; Zhan, B.; Wang, H.; Wang, Y.; Chen, Y.; Wang, X.; An, Y.; Tang, C.; Hou, L.; Fan, L.; and Zhang, Z. 2026 d . DriveVLA-W0 : World Models Amplify Data Scaling Law in Autonomous Driving. In ICLR

  87. [114]

    Li, Y.; Wei, X.; Cao, J.; Wang, H.; Chi, X.; Bai, C.; Sun, Q.; Li, J.; Zhang, X.; Jia, P.; Tang, J.; Han, S.; and Zhang, S. 2026 e . WAM4D : Fast 4D World Action Model via Spatial Register Tokens. arXiv:2606.14048

  88. [115]

    Li, Y.; Xiong, K.; Guo, X.; Li, F.; Yan, S.; Xu, G.; Zhou, L.; Chen, L.; Sun, H.; Wang, B.; Ma, K.; Chen, G.; Ye, H.; Liu, W.; and Wang, X. 2026 f . ReCogDrive : A Reinforced Cognitive Framework for End-to-End Autonomous Driving. In ICLR

  89. [116]

    Liang, D.; Zhang, D.; Zhou, X.; Tu, S.; Feng, T.; Li, X.; Zhang, Y.; Du, M.; Tan, X.; and Bai, X. 2026. UniFuture : A 4D Driving World Model for Future Generation and Perception. In ICRA

  90. [117]

    Liang, W.; Yu, L.; Luo, L.; Iyer, S.; Dong, N.; Zhou, C.; Ghosh, G.; Lewis, M.; Yih, W.-t.; Zettlemoyer, L.; and Lin, X. V. 2025. Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR

  91. [118]

    Liao, B.; Chen, S.; Yin, H.; Jiang, B.; Wang, C.; Yan, S.; Zhang, X.; Li, X.; Zhang, Y.; Zhang, Q.; and Wang, X. 2025. DiffusionDrive : Truncated Diffusion Model for End-to-End Autonomous Driving. In CVPR

  92. [119]

    H.; Chen, D

    Lin, H.; Chen, S.; Liew, J. H.; Chen, D. Y.; Li, Z.; Zhao, Y.; Peng, S.; Guo, H.; Zhou, X.; Shi, G.; Feng, J.; and Kang, B. 2026. Depth Anything 3: Recovering the Visual Space from Any Views. In ICLR

  93. [120]

    Liu, L.; Jia, C.; Yu, G.; Song, Z.; Li, J.; Jia, F.; Wu, P.; Hao, X.; and Luo, Y. 2026. GuideFlow : Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving. In CVPR

  94. [121]

    Luiten, J.; Osep, A.; Dendorfer, P.; Torr, P. H. S.; Geiger, A.; Leal-Taix \'e , L.; and Leibe, B. 2021. HOTA : A Higher Order Metric for Evaluating Multi-Object Tracking. IJCV

  95. [122]

    Ma, F.; Peng, D.; Yue, W.; Cao, J.; Wang, B.; Zhang, Q.; and Ma, J. 2026. GeoSem-WAM : Geometry- and Semantic-Aware World Action Models. arXiv:2606.03188

  96. [123]

    NVIDIA ; et al. 2025. Cosmos World Foundation Model Platform for Physical AI . arXiv:2501.03575

  97. [124]

    Rowe, L.; de Schaetzen, R.; Girgis, R.; Pal, C.; and Paull, L. 2025. Poutine : Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving. arXiv:2506.11234

  98. [125]

    L.; Zhan, W.; and Li, H

    Shao, H.; Wang, L.; Zhou, Y.; Hu, Y.; Zong, Z.; Waslander, S. L.; Zhan, W.; and Li, H. 2026. LMGenDrive : Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving. arXiv:2604.08719

  99. [126]

    Sheng, Z.; Ye, X.; Luo, J.; Chen, S.; and Ren, L. 2026. ExploreVLA : Dense World Modeling and Exploration for End-to-End Autonomous Driving. In ECCV

  100. [127]

    Sun, W.; Lin, X.; Chen, K.; Pei, Z.; Li, X.; Shi, Y.; and Zheng, S. 2026. SparseDriveV2 : Scoring is All You Need for End-to-End Autonomous Driving. arXiv:2603.29163

  101. [128]

    Sun, W.; Lin, X.; Shi, Y.; Zhang, C.; Wu, H.; and Zheng, S. 2025. SparseDrive : End-to-End Autonomous Driving via Sparse Scene Representation. In ICRA

  102. [129]

    Team Wan ; et al. 2025. Wan : Open and Advanced Large-Scale Video Generative Models. arXiv:2503.20314

  103. [130]

    Tian, X.; Gu, J.; Li, B.; Liu, Y.; Wang, Y.; Zhao, Z.; Zhan, K.; Jia, P.; Lang, X.; and Zhao, H. 2025. DriveVLM : The Convergence of Autonomous Driving and Large Vision-Language Models. In CoRL

  104. [131]

    Wang, D.; Song, Y.; He, Z.; Chen, K.; Pan, X.; Deng, L.; and Gu, W. 2025 a . HMVLM : Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios. arXiv:2506.05883

  105. [132]

    Wang, L.; Yang, Z.; Bai, C.; Zhang, G.; Liu, X.; Zheng, X.; Long, X.-X.; Lu, C.-T.; and Lu, C. 2026 a . Drive-JEPA : Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving. arXiv:2601.22032

  106. [133]

    G.; Zanfir, A.; and Sminchisescu, C

    Wang, L.; Zhang, C.; Kabra, R.; Uijlings, J.; Waslander, S.; Zisserman, A.; Carreira, J.; He, K.; Andriluka, M.; Bazavan, E. G.; Zanfir, A.; and Sminchisescu, C. 2026 b . Video Generation Models are General-Purpose Vision Learners. In ECCV

  107. [134]

    Wang, S.; Yu, Z.; Jiang, X.; Lan, S.; Shi, M.; Chang, N.; Kautz, J.; Li, Y.; and Alvarez, J. M. 2025 b . OmniDrive : A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning. In CVPR

  108. [135]

    Wang, X.; Zhu, Z.; Huang, G.; Chen, X.; Zhu, J.; and Lu, J. 2024 a . DriveDreamer : Towards Real-world-driven World Models for Autonomous Driving. In ECCV

  109. [136]

    Wang, Y.; He, J.; Fan, L.; Li, H.; Chen, Y.; and Zhang, Z. 2024 b . Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving. In CVPR

  110. [137]

    Xia, T.; Li, Y.; Zhou, L.; Yao, J.; Xiong, K.; Sun, H.; Wang, B.; Ma, K.; Chen, G.; Ye, H.; Liu, W.; and Wang, X. 2026. DriveLaW : Unifying Planning and Video Generation in a Latent Driving World. In CVPR

  111. [138]

    Xu, J.; Zhong, Z.; Shu, Z.; Jia, M.; Li, M.; Bian, J.-W.; Zhang, Q.; Zhang, K.; Xie, J.; Yang, J.; and Yin, W. 2026 a . EponaV2 : Driving World Model with Comprehensive Future Reasoning. arXiv:2605.14696

  112. [139]

    Xu, R.; Lin, H.; Jeon, W.; Feng, H.; Zou, Y.; Sun, L.; Gorman, J.; Tolstaya, K.; Tang, S.; White, B.; Sapp, B.; Tan, M.; Hwang, J.-J.; and Anguelov, D. 2026 b . WOD-E2E : Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios. In CVPR

  113. [140]

    M.; and Wu, Z

    Yao, W.; Li, Z.; Lan, S.; Wang, Z.; Sun, X.; Alvarez, J. M.; and Wu, Z. 2026. DriveSuprim : Towards Precise Trajectory Selection for End-to-End Planning. In AAAI

  114. [141]

    Ye, Y.; Zhang, Z.; Lin, J.; Sun, S.; Peng, C.; and Gao, W. 2026. AutoDrive-P ^3 : Unified Chain of Perception--Prediction--Planning Thought via Reinforcement Fine-Tuning. In ICLR

  115. [142]

    Yuan, T.; Dong, Z.; Liu, Y.; and Zhao, H. 2026. Fast-WAM : Do World Action Models Need Test-time Future Imagination? arXiv:2603.16666

  116. [143]

    Zhang, B.; Song, N.; Li, J.; Zhu, X.; Deng, J.; and Zhang, L. 2025. Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution. In NeurIPS

  117. [144]

    Zhang, K.; Wang, J.; Gao, S.; Wu, C.; Cao, Y.; Han, S.; Ivanovic, B.; Liu, L.; Pavone, M.; Han, S.; Zhou, D.; and Xie, E. 2026. Fast-dDrive : Efficient Block-Diffusion VLM for Autonomous Driving. arXiv:2605.23163

  118. [145]

    Zhao, Z.; Fu, T.; Wang, Y.; Wang, L.; and Lu, H. 2025. From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction. In NeurIPS

  119. [146]

    Zhen, H.; Sun, Q.; Zhang, H.; Li, J.; Zhou, S.; Du, Y.; and Gan, C. 2025. Learning 4D Embodied World Models. In ICCV

  120. [147]

    Zheng, W.; Xia, Z.; Huang, Y.; Zuo, S.; Zhou, J.; and Lu, J. 2024. Doe-1: Closed-Loop Autonomous Driving with Large World Model. arXiv:2412.09627

  121. [148]

    Zheng, Y.; Yang, P.; Xing, Z.; Zhang, Q.; Zheng, Y.; Gao, Y.; Li, P.; Zhang, T.; Xia, Z.; Jia, P.; Lang, X.; and Zhao, D. 2025. World4Drive : End-to-End Autonomous Driving via Intention-Aware Physical Latent World Model. In ICCV

  122. [149]

    Zhou, Y.; Wang, X.; Shao, H.; Wang, L.; Zhao, G.; Shao, J.; Zhu, J.; Yu, T.; Zhu, Z.; Huang, G.; and Waslander, S. L. 2026. DriveDreamer-Policy : A Geometry-Grounded World-Action Model for Unified Generation and Planning. arXiv:2604.01765

  123. [150]

    Zhou, Z.; Cai, T.; Zhao, S.; Zhang, Y.; Huang, Z.; Zhou, B.; and Ma, J. 2025. AutoVLA : A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning. In NeurIPS

  124. [151]

    Zou, J.; Chen, S.; Liao, B.; Zheng, Z.; Song, Y.; Zhang, L.; Zhang, Q.; Liu, W.; and Wang, X. 2025. DiffusionDriveV2 : Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving. arXiv:2512.07745

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.