REVIEW 2 major objections 5 minor 132 references
SUV: Future Scene Understanding as Video Generation for End-to-End Driving
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read One video model generates four road futures and steers by them
desk verdict A solid, genuinely new interface for future-scene prediction in driving—one shared video expert generating four stream types and an action expert reading their latents—with planning results that survive scrutiny, though the structured-stream evaluation is partly circular against its own frozen teachers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the token-interaction mask inside joint video-action attention, together with the shared latent video space. The mask implements directed information flow: observation tokens read only the clean prefix, each future stream reads the prefix and itself, and the action expert reads the prefix, all four future streams, and its own trajectory tokens, and this mask is applied at every transformer block and denoising step. What makes it work is that all four streams share one latent video layout, one observation prefix, and one flow-matching schedule with a shifted timestep $\phi_\kappa(\rho)=\kappa\rho/(1+(\kappa-1)\rho)$, $\kappa=5$, so the pretrained video expert generates them natively; the action expert needs no decoded video, only intermediate latents.
What would settle it
Replace the future-stream latents fed to the action expert with shape-matched Gaussian noise at inference; if navhard EPDMS stays near 36.9, the claimed causal contribution of direct future-stream access is refuted.
Extended reading notes
Core claim
The central claim is that future scene understanding for end-to-end driving can be recast as video generation: the same video expert that predicts future RGB also predicts segmentation, relative depth, and instance tracks as three-channel video streams in a shared latent space, initialized from generative video pretraining and post-trained with flow matching. A Mixture-of-Transformers action expert shares the denoising loop; a token-interaction mask lets action queries attend to every future-stream token group while blocking cross-stream and action-to-future information flow. The paper's evidence is that this setting reaches 91.0 EPDMS on NAVSIM-v2 navtest and 36.9 on navhard with a single front camera and no trajectory-candidate selection, outperforming a broad set of recent methods, and that the multi-stream model matches or slightly exceeds RGB-only future quality. It further argues, via a 2x2 ablation, that adding segmentation/depth/track supervision raises EPDMS on navtest from 89.7 to 90.7 and on navhard from 30.5 to 32.8, and that giving the action expert direct access to future-stream latents raises it further to 91.0 and 36.9.
Load-bearing premise
Future-scene quality and the planning benefit attributed to structured futures are measured against the same frozen teacher models that generated the training targets, so systematic teacher errors would inflate both the reported stream quality and, potentially, the planning gain.
Editorial extensions
If this is right
- Adding a new kind of future knowledge, such as drivable-area gradient or object velocities, reduces to rendering it as another video stream and adding a prompt, with no new visual head or decoder required.
- Planning can happen entirely in latent space: the trajectory is denoised alongside the future streams, and the VAE decoder is needed only for evaluation or visualization.
- The accuracy-latency trade-off is controllable through solver steps: one Euler step gives 89.8 EPDMS on navtest and 33.0 on navhard at 177 ms on an RTX 4090, while ten steps reach 36.9 on navhard at 1356 ms.
- Both structured future supervision and direct future-stream access improve planning, with the larger gains on the long-tail navhard split.
- Removing access to any single structured stream lowers navhard EPDMS, with track removal hurting Stage 2 most; no single stream alone explains the gain.
Reading between the lines
- (Editorial inference) If the shared-expert design transfers beyond roads, a single generative backbone could become the common future-predictor for other embodied tasks, such as manipulation or navigation, where semantic, geometric, and instance futures are rendered as color-coded streams.
- (Editorial inference) Because the structured metrics judge agreement with the frozen teachers that built the targets, an independent-ground-truth evaluation could reorder the comparison between native generation and generate-then-perceive.
- (Editorial inference) The per-stream access ablations suggest a cheaper deployment recipe: generate only RGB and read its latents while dropping the other streams at inference, which the paper's own numbers predict would cost little on navtest but several points on navhard.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes SUV, an end-to-end driving framework that casts future scene understanding as video generation. A single video expert initialized from Wan2.2-5B is post-trained to generate four future streams (RGB, semantic segmentation, relative depth, instance tracks) as native videos, without stream-specific visual prediction heads. A separate action expert attends to the latent tokens of all streams during joint denoising to produce a trajectory. The model is evaluated on NAVSIM-v2 (91.0 EPDMS on navtest, 36.9 on navhard) and WOD-E2E (RFS 7.94) with a single front camera and no candidate selection, along with ablations isolating the contributions of generative initialization, structured supervision, and future-stream access. The paper also provides six-run reproducibility checks and detailed supplementary documentation of target construction and training.
Significance. If the results hold, SUV offers a conceptually clean demonstration that one pretrained video generator can serve as a shared predictor for heterogeneous future scene signals and that the resulting future-stream latents can improve trajectory planning. The external planning benchmark results are competitive or state-of-the-art, the experimental design includes controlled 2x2 ablations and per-stream access ablations, and the paper ships code and reproducibility details. The principal weakness is that the future-scene evaluation is self-referential: segmentation, depth, and tracking references are produced by the same frozen SAM3 and DA3 models that generated the training targets, so the reported mIoU, δ1, AbsRel, and AssA@50 values measure teacher-agreement rather than independent scene understanding. This gap is explicitly acknowledged in the paper but remains unresolved, and it weakens the headline 'future scene understanding' contribution.
major comments (2)
- [Future-scene evaluation (main text) and Supplementary §4.1] The structured-stream metrics are computed against references generated by the same frozen SAM3 (segmentation, tracks) and DA3 (depth) models that produced the training targets, as the paper explicitly acknowledges: 'The segmentation, tracking, and depth metrics measure agreement with frozen teachers rather than accuracy against independent ground truth.' This circularity is load-bearing because the paper's headline contribution is future scene understanding, and the evidence for it is currently only teacher-agreement. Please either (a) add an independent evaluation subset (e.g., human annotations or a different off-the-shelf model for a sample of frames) and report the same metrics on that subset, or (b) explicitly reframe the future-scene claims as 'agreement with frozen teacher models' rather than 'accuracy' or 'understanding'. The Table 5 comparison between native generation and Generate-then-Perceive is useful for task-alignment but cannot detect teacher-specific artifacts because both arms use the same teacher-generated references.
- [Planning Benefits of Future Representations, Table 6] The 2x2 ablation shows that adding segmentation/relative-depth/instance-track supervision (S/G/I) improves EPDMS on both navtest and navhard. Since the structured targets come from the same teachers used in the future-scene evaluation, this improvement could partly reflect the planner exploiting teacher-specific artifacts rather than general scene structure. The planning scores themselves are external and therefore not invalidated, but the interpretation of the improvement as evidence for 'future scene understanding' inherits the circularity concern. Please address this by either adding a control with targets from a different teacher on a subset (e.g., a different depth or segmentation model) or explicitly discussing the risk and tempering the corresponding conclusion.
minor comments (5)
- [Introduction / Contributions] The novelty claim ('the first end-to-end driving framework ...') should be cross-checked against the most recent concurrent work on generative world models for driving (e.g., EponaV2, GeoSem-WAM, WAM4D). A sentence explaining how SUV specifically differs in the format of the predicted streams (native video vs. task-specific heads) would strengthen the claim.
- [Planning Benefits of Future Representations] The abbreviation 'S/G/I' is used in the text without definition; it is only defined in the Table 6 caption. Please define it at first use.
- [Supplementary §4.3] 'Rankr seeds Python, NumPy, PyTorch, and CUDA' appears to be a typo; it should read 'rank r seeds ...'.
- [Table 4] The sentence 'Wan2.2-5B initialization yields better point estimates for all five metrics' is correct, but the large gaps (e.g., mIoU 64.2 vs 49.7) would benefit from a note on whether these differences are consistent across repeated training runs, or at least a statement that the six-run reproducibility checks were only reported for planning scores.
- [Method, Eq. (3)] The notation A(g) is slightly confusing because it is a set of visible key groups. Consider defining it explicitly as a function from a token group to the set of allowed key groups to improve readability.
Circularity Check
Future-scene metrics are scored against the same frozen SAM3/DA3 teachers that generated the training targets, so the reported mIoU, AssA@50, δ1, and AbsRel measure teacher imitation rather than independent scene understanding; the paper explicitly acknowledges this caveat but the self-reference remains unresolved.
-
self definitional
[Experiments, Future-scene evaluation; Supplementary Material §4.1]
"Frozen teachers convert each recorded future RGB clip into videos for semantic segmentation, relative depth, and instance tracks. ... These videos serve as training targets for both benchmarks and as evaluation references for the NAVSIM future-scene analysis. ... The segmentation, tracking, and depth metrics measure agreement with frozen teachers rather than accuracy against independent ground truth."
The evaluation references for mIoU (segmentation), AssA@50 (tracks), and δ1/AbsRel (depth) are produced by the same frozen SAM3 and DA3 models that generated the training targets. The model is trained with MSE to reproduce those teacher labels, and the reported 'future scene understanding' scores then measure agreement with those same labels. Perfectly memorizing the training targets would, by construction, yield near-perfect teacher-agreement scores, so these numbers cannot validate the structured streams as independent scene understanding. Table 5 is also internal to this loop: Generate-then-Perceive applies the same teachers to generated RGB, and native generation is scored against the same teacher references, so the comparison only ranks which pipeline best reproduces the teachers.
-
fitted input called prediction
[Supplementary Material §1.2, Eq. (4)–(5); Supplementary Material §4.1]
"At evaluation, each generated depth pixel is mapped to its nearest Turbo entry. ... For NAVSIM future-scene evaluation, Equation 4 fits one affine map from these normalized values to the DA3-reference scale over all eight frames. ... We report δ1 and AbsRel as affine-aligned, percentile-clipped measures of agreement with the DA3 teacher."
The depth training target is DA3's clip-level percentile-normalized, Turbo-quantized output, and the depth evaluation reference is DA3 applied to the recorded RGB frames. Before computing δ1 and AbsRel, the prediction is affinely aligned to that same DA3 reference, removing global scale and shift errors—among the most basic geometric errors a depth predictor can make. The metric then rewards reproduction of the teacher's normalization and color mapping after optimal reparameterization.
full rationale
The paper's headline planning results are evaluated on external benchmarks (NAVSIM-v2 EPDMS and WOD-E2E RFS) and are not circular; the trajectory scores come from an independent protocol with no fitted parameters tied to the future-stream evaluation. The circularity is confined to the future-scene understanding contribution. By the paper's own description, frozen SAM3 and DA3 generate the segmentation, relative-depth, and instance-track training targets, and the same frozen teachers provide the evaluation references for mIoU, AssA@50, δ1, and AbsRel. The structured-stream metrics therefore measure how faithfully the video expert reproduces its teacher-derived supervision, not whether the predicted streams are accurate against independent ground truth. The paper acknowledges this plainly ('measure agreement with frozen teachers rather than accuracy against independent ground truth'), but an acknowledgement does not break the construction: the evaluation loop is closed by design. Because the central novelty is precisely that the generated structured streams constitute future scene understanding, the self-referential evaluation prevents the structured-stream results from independently validating that claim. The RGB quality metrics and the external planning numbers remain independent; however, the integrated claim that structured future supervision and future-stream access improve planning inherits some risk if the planner exploits teacher-specific artifacts. This is partial circularity, not total: score 6, with the non-circular external planning results as the mitigating factor.
Assumptions & free parameters
free parameters (6)
- Shifted flow-time mapping kappa =
5
- Scheduler weight shape exponent =
2 in q(lambda) = exp(-2(lambda-1/2)^2) - exp(-1/2)
- Action-vs-video loss weight =
1
- Depth normalization quantiles =
q_0.01 and q_0.99
- Instance-track decoding thresholds =
distance 30, min size 32, IoU/centroid cost weights
- SAM3 target-construction thresholds =
confidence 0.5, 32 detections per prompt, 128 detection cap
assumptions (5)
- standard math Flow-matching training with the scored velocity estimator converges and shifted Euler inference is numerically stable.
- domain assumption Pretrained Wan2.2-5B provides a useful generative prior for jointly predicting all four future streams in driving scenes.
- domain assumption Frozen SAM3 and DA3 outputs provide reliable supervision and valid references for future semantics, relative depth, and instance tracks.
- domain assumption The masked attention scheme in Eq. 3, blocking cross-stream interaction and action-to-future feedback, is a suitable inductive bias.
- domain assumption NAVSIM-v2 EPDMS with human-penalty filtering is an appropriate proxy for autonomous driving performance.
Cite this review
Pith. "Pith review of SUV: Future Scene Understanding as Video Generation for End-to-End Driving." pith.science (2026). https://pith.science/paper/OCC73BNS
@misc{pith2026260803084,
author = {Pith},
title = {Pith review of: SUV: Future Scene Understanding as Video Generation for End-to-End Driving},
year = {2026},
howpublished = {\url{https://pith.science/paper/OCC73BNS}},
note = {Machine review of arXiv:2608.03084}
}
read the original abstract
End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide a shared predictor? We introduce SUV, a unified end-to-end driving framework that casts future Scene Understanding as Video generation using a pretrained video foundation model. SUV models future appearance, semantics, relative depth, and instance-level dynamics as video streams with a shared video expert, without stream-specific visual prediction heads. Through joint video-action attention, the action expert attends to the latent representations of all future streams and generates the ego trajectory. Experiments show that SUV directly predicts all four future streams, while controlled ablations show that structured future supervision and direct future-stream access yield higher trajectory planning scores. With only a single front camera and no candidate-trajectory selection, SUV outperforms a broad set of recent state-of-the-art methods on both NAVSIM-v2 splits, achieving 91.0 EPDMS on navtest and 36.9 on navhard. On the long-tail WOD-E2E benchmark, SUV achieves a competitive RFS of 7.94.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Luiten, Jonathon and Osep, Aljosa and Dendorfer, Patrick and Torr, Philip H. S. and Geiger, Andreas and Leal-Taix. 2021 , doi=
2021
-
[2]
2026 , eprint=
Image Generators are Generalist Vision Learners , author=. 2026 , eprint=
2026
-
[3]
2026 , eprint=
Video Generation Models are General-Purpose Vision Learners , author=. 2026 , eprint=
2026
-
[4]
2026 , eprint=
Vision as Unified Multimodal Generation , author=. 2026 , eprint=
2026
-
[5]
Carion, Nicolas and Gustafson, Laura and Hu, Yuan-Ting and Debnath, Shoubhik and Hu, Ronghang and Suris Coll-Vinent, Didac and Ryali, Chaitanya and Alwala, Kalyan Vasudev and Khedr, Haitham and Huang, Andrew and Lei, Jie and Ma, Tengyu and Guo, Baishan and Kalla, Arpit and Marks, Markus and Greer, Joseph and Wang, Meng and Sun, Peize and Rädle, Roman and ...
-
[6]
2026 , url=
Depth Anything 3: Recovering the Visual Space from Any Views , author=. 2026 , url=
2026
-
[7]
2025 , issn=
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models , author=. 2025 , issn=
2025
-
[8]
Jiang, Bo and Chen, Shaoyu and Xu, Qing and Liao, Bencheng and Chen, Jiajie and Zhou, Helong and Zhang, Qian and Liu, Wenyu and Huang, Chang and Wang, Xinggang , booktitle=ICCV, year=
Show all 132 references
-
[9]
Sun, Wenchao and Lin, Xuewu and Shi, Yining and Zhang, Chuang and Wu, Haoran and Zheng, Sifa , booktitle=ICRA, year=
-
[10]
Planning-
Hu, Yihan and Yang, Jiazhi and Chen, Li and Li, Keyu and Sima, Chonghao and Zhu, Xizhou and Chai, Siqi and Du, Senyao and Lin, Tianwei and Wang, Wenhai and Lu, Lewei and Jia, Xiaosong and Liu, Qiang and Dai, Jifeng and Qiao, Yu and Li, Hongyang , booktitle=CVPR, year=. Planning-
-
[11]
Tian, Xiaoyu and Gu, Junru and Li, Bailin and Liu, Yicheng and Wang, Yang and Zhao, Zhiyong and Zhan, Kun and Jia, Peng and Lang, XianPeng and Zhao, Hang , booktitle=CORL, year=
-
[12]
Hwang, Jyh-Jing and Xu, Runsheng and Lin, Hubert and Hung, Wei-Chih and Ji, Jingwei and Choi, Kristy and Huang, Di and He, Tong and Covington, Paul and Sapp, Benjamin and Zhou, Yin and Guo, James and Anguelov, Dragomir and Tan, Mingxing , journal=TMLR, year=
-
[13]
, booktitle=CVPR, year=
Wang, Shihao and Yu, Zhiding and Jiang, Xiaohui and Lan, Shiyi and Shi, Min and Chang, Nadine and Kautz, Jan and Li, Ying and Alvarez, Jose M. , booktitle=CVPR, year=
-
[14]
Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in
Zhai, Jiang-Tian and Feng, Ze and Du, Jinhao and Mao, Yongqiang and Liu, Jiang-Jiang and Tan, Zichang and Zhang, Yifu and Ye, Xiaoqing and Wang, Jingdong , year=. Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in. 2305.10430 , archivePrefix=
-
[15]
2024 , doi=
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? , author=. 2024 , doi=
2024
-
[17]
Dauner, Daniel and Hallgarten, Marcel and Li, Tianyu and Weng, Xinshuo and Huang, Zhiyu and Yang, Zetong and Li, Hongyang and Gilitschenski, Igor and Ivanovic, Boris and Pavone, Marco and Geiger, Andreas and Chitta, Kashyap , booktitle=NeurIPS, year=
-
[18]
2025 , url=
Pseudo-Simulation for Autonomous Driving , author=. 2025 , url=
2025
-
[19]
Xu, Runsheng and Lin, Hubert and Jeon, Wonseok and Feng, Hao and Zou, Yuliang and Sun, Liting and Gorman, John and Tolstaya, Kate and Tang, Sarah and White, Brandyn and Sapp, Ben and Tan, Mingxing and Hwang, Jyh-Jing and Anguelov, Dragomir , booktitle=CVPR, year=
-
[20]
2026 , eprint=
World Model for Robot Learning: A Comprehensive Survey , author=. 2026 , eprint=
2026
-
[21]
2022 , doi=
Model-Based Imitation Learning for Urban Driving , author=. 2022 , doi=
2022
-
[23]
Wang, Xiaofeng and Zhu, Zheng and Huang, Guan and Chen, Xinze and Zhu, Jiagang and Lu, Jiwen , booktitle=ECCV, year=
-
[24]
2024 , doi=
Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving , author=. 2024 , doi=
2024
-
[25]
2024 , doi=
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability , author=. 2024 , doi=
2024
-
[26]
Chen, Yuntao and Wang, Yuqi and Zhang, Zhaoxiang , booktitle=ICCV, year=
-
[27]
2024 , eprint=
Doe-1: Closed-Loop Autonomous Driving with Large World Model , author=. 2024 , eprint=
2024
-
[28]
Wei, Julong and Yuan, Shanshuai and Li, Pengfei and Quan, Xinyi and Tai, Lei and Zhao, Jieru and Gan, Zhongxue and Ding, Wenchao , booktitle=ICRA, year=
-
[29]
Zheng, Wenzhao and Chen, Weiliang and Huang, Yuanhui and Zhang, Borui and Duan, Yueqi and Lu, Jiwen , booktitle=ECCV, year=
-
[30]
2025 , doi=
Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving , author=. 2025 , doi=
2025
-
[31]
2024 , doi=
Visual Point Cloud Forecasting enables Scalable Autonomous Driving , author=. 2024 , doi=
2024
-
[32]
Min, Chen and Zhao, Dawei and Xiao, Liang and Zhao, Jian and Xu, Xinli and Zhu, Zheng and Jin, Lei and Li, Jianshu and Guo, Yulan and Xing, Junliang and Jing, Liping and Nie, Yiming and Dai, Bin , booktitle=CVPR, year=
-
[33]
Enhancing End-to-End Autonomous Driving with Latent World Model , author=
-
[34]
Li, Yingyan and Shang, Shuyao and Liu, Weisong and Zhan, Bing and Wang, Haochen and Wang, Yuqi and Chen, Yuntao and Wang, Xiaoman and An, Yasong and Tang, Chufeng and Hou, Lu and Fan, Lue and Zhang, Zhaoxiang , booktitle=ICLR, year=
-
[35]
doi:10.1109/LRA.2026.3678836 , url=
Wozniak, Maciej and Liu, Lianhang and Cai, Yixi and Jensfelt, Patric , journal=RAL, year=. doi:10.1109/LRA.2026.3678836 , url=
2026
-
[39]
Liang, Dingkang and Zhang, Dingyuan and Zhou, Xin and Tu, Sifan and Feng, Tianrui and Li, Xiaofan and Zhang, Yumeng and Du, Mingyang and Tan, Xiao and Bai, Xiang , booktitle=ICRA, year=
-
[41]
Li, Bohan and Ma, Zhuang and Du, Dalong and Peng, Baorui and Liang, Zhujin and Liu, Zhenqiang and Guo, Xianda and Zhu, Zheng and Ma, Chao and Jin, Yueming and Jin, Xin and Zhao, Hao and Zeng, Wenjun , booktitle=ECCV, year=
-
[42]
2024 , doi=
Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation , author=. 2024 , doi=
2024
-
[43]
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? , author=
-
[44]
Zhao, Canyu and Sun, Yanlong and Liu, Mingyu and Zheng, Huanyi and Zhu, Muzhi and Zhao, Zhiyue and Chen, Hao and He, Tong and Shen, Chunhua , booktitle=NeurIPS, year=
-
[45]
2603.25892 , archivePrefix=
Wang, Letian and Zanfir, Andrei and Bazavan, Eduard Gabriel and Andriluka, Misha and Sminchisescu, Cristian , year=. 2603.25892 , archivePrefix=
-
[46]
Huang, Wenhui and Zhang, Songyan and Huang, Qihang and Wang, Zhidong and Mao, Zhiqi and Chua, Collister and Chen, Zhan and Chen, Long and Lv, Chen , booktitle=ICML, year=
-
[48]
Bi, Hongzhe and Tan, Hengkai and Xie, Shenghao and Wang, Zeyuan and Huang, Shuhe and Liu, Haitian and Zhao, Ruowen and Feng, Yao and Xiang, Chendong and Rong, Yinze and Zhao, Hongyan and Liu, Hanyu and Su, Zhizhong and Ma, Lei and Su, Hang and Zhu, Jun , booktitle=CVPR, year=
-
[51]
Li, Jingyu and Zhang, Bozhou and Jin, Xin and Deng, Jiankang and Zhu, Xiatian and Zhang, Li , booktitle=ICRA, year=
-
[53]
Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution , author=
-
[55]
Dang, Chenxu and Ang, Sining and Li, Yongkang and Tian, Haochen and Wang, Jie and Li, Guang and Ye, Hangjun and Ma, Jie and Chen, Long and Wang, Yan , booktitle=ECCV, year=
-
[56]
2605.15120 , archivePrefix=
Ang, Sining and Yang, Yuguang and Chen, Canyu and Wang, Yan , year=. 2605.15120 , archivePrefix=
-
[57]
Liu, Lin and Jia, Caiyan and Yu, Guanyi and Song, Ziying and Li, Junqiao and Jia, Feiyang and Wu, Peiliang and Hao, Xiaoshuai and Luo, Yadan , booktitle=CVPR, year=
-
[58]
2025 , doi=
Learning 4D Embodied World Models , author=. 2025 , doi=
2025
-
[59]
Xiong, Zhexiao and Ye, Xin and Yaman, Burhan and Cheng, Sheng and Lu, Yiren and Luo, Jingru and Jacobs, Nathan and Ren, Liu , booktitle=ECCV, year=
-
[61]
Sheng, Zihao and Ye, Xin and Luo, Jingru and Chen, Sikai and Ren, Liu , booktitle=ECCV, year=
-
[62]
Jia, Feiyang and Liu, Lin and Song, Ziying and Jia, Caiyan and Ye, Hangjun and Hao, Xiaoshuai and Chen, Long , booktitle=ICML, year=
-
[63]
Li, Jingyu and Wu, Junjie and Hu, Dongnan and Huang, Xiangkai and Sun, Bin and Hao, Zhihui and Lang, Xianpeng and Zhu, Xiatian and Zhang, Li , booktitle=CVPR, year=
-
[66]
and Wu, Zuxuan , booktitle=AAAI, year=
Yao, Wenhao and Li, Zhenxin and Lan, Shiyi and Wang, Zi and Sun, Xinglong and Alvarez, Jose M. and Wu, Zuxuan , booktitle=AAAI, year=
-
[68]
Li, Yongkang and Xiong, Kaixin and Guo, Xiangyu and Li, Fang and Yan, Sixu and Xu, Gangwei and Zhou, Lijun and Chen, Long and Sun, Haiyang and Wang, Bing and Ma, Kun and Chen, Guang and Ye, Hangjun and Liu, Wenyu and Wang, Xinggang , booktitle=ICLR, year=
-
[69]
Chitta, Kashyap and Prakash, Aditya and Jaeger, Bernhard and Yu, Zehao and Renz, Katrin and Geiger, Andreas , journal=PAMI, year=
-
[70]
Liao, Bencheng and Chen, Shaoyu and Yin, Haoran and Jiang, Bo and Wang, Cheng and Yan, Sixu and Zhang, Xinbang and Li, Xiangyu and Zhang, Ying and Zhang, Qian and Wang, Xinggang , booktitle=CVPR, year=
-
[71]
Wang, Jie and Li, Guang and Huang, Zhijian and Dang, Chenxu and Ye, Hangjun and Han, Yahong and Chen, Long , booktitle=CVPR, year=
-
[72]
Zhang, Jinqing and Fu, Zehua and Xu, Zelin and Dai, Wenying and Liu, Qingjie and Wang, Yunhong , booktitle=ICLR, year=
-
[73]
Wang, Junli and Zheng, Yinan and Liu, Xueyi and Xing, Zebin and Li, Pengfei and Ma, Kun and Ye, Hangjun and Chen, Guang and Li, Guang and Chen, Long and Xia, Zhongpu and Zhang, Qichao , booktitle=CVPR, year=
-
[74]
Zhou, Zewei and Cai, Tianhui and Zhao, Seth and Zhang, Yun and Huang, Zhiyu and Zhou, Bolei and Ma, Jiaqi , booktitle=NeurIPS, year=
-
[77]
Ye, Yuqi and Zhang, Zijian and Lin, Junhong and Sun, Shangkun and Peng, Changhao and Gao, Wei , booktitle=ICLR, year=
-
[78]
Zheng, Yupeng and Yang, Pengxuan and Xing, Zebin and Zhang, Qichao and Zheng, Yuhang and Gao, Yinfeng and Li, Pengfei and Zhang, Teng and Xia, Zhongpu and Jia, Peng and Lang, Xianpeng and Zhao, Dongbin , booktitle=ICCV, year=
-
[79]
Zhang, Kaiwen and Tang, Zhenyu and Hu, Xiaotao and Pan, Xingang and Guo, Xiaoyang and Liu, Yuan and Huang, Jingwei and Yuan, Li and Zhang, Qian and Long, Xiao-Xiao and Cao, Xun and Yin, Wei , booktitle=ICCV, year=
-
[80]
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction , author=
-
[81]
Xia, Tianze and Li, Yongkang and Zhou, Lijun and Yao, Jingfeng and Xiong, Kaixin and Sun, Haiyang and Wang, Bing and Ma, Kun and Chen, Guang and Ye, Hangjun and Liu, Wenyu and Wang, Xinggang , booktitle=CVPR, year=
-
[82]
Feng, Lan and Gao, Yang and Zablocki, Eloi and Li, Quanyi and Li, Wuyang and Liu, Sichao and Cord, Matthieu and Alahi, Alexandre , booktitle=ICLR, year=
-
[83]
2605.12624 , archivePrefix=
Huang, Yuzhou and Zhu, Benjin and Lu, Hengtong and Huang, Victor Shea-Jay and Zhang, Haiming and Chen, Wei and Dai, Jifeng and Xie, Yan and Li, Hongsheng , year=. 2605.12624 , archivePrefix=
-
[86]
Luiten, Jonathon and Osep, Aljosa and Dendorfer, Patrick and Torr, Philip H. S. and Geiger, Andreas and Leal-Taix. IJCV , year=
-
[87]
Bi, H.; Tan, H.; Xie, S.; Wang, Z.; Huang, S.; Liu, H.; Zhao, R.; Feng, Y.; Xiang, C.; Rong, Y.; Zhao, H.; Liu, H.; Su, Z.; Ma, L.; Su, H.; and Zhu, J. 2026. Motus : A Unified Latent Action World Model. In CVPR
2026
-
[88]
Cao, W.; Hallgarten, M.; Li, T.; Dauner, D.; Gu, X.; Wang, C.; Miron, Y.; Aiello, M.; Li, H.; Gilitschenski, I.; Ivanovic, B.; Pavone, M.; Geiger, A.; and Chitta, K. 2025. Pseudo-Simulation for Autonomous Driving. In CoRL
2025
-
[89]
Carion, N.; Gustafson, L.; Hu, Y.-T.; Debnath, S.; Hu, R.; Suris Coll-Vinent, D.; Ryali, C.; Alwala, K. V.; Khedr, H.; Huang, A.; Lei, J.; Ma, T.; Guo, B.; Kalla, A.; Marks, M.; Greer, J.; Wang, M.; Sun, P.; Rädle, R.; Afouras, T.; Mavroudi, E.; Xu, K.; Wu, T.-H.; Zhou, Y.; Mo...
2026
-
[90]
Chen, Y.; Wang, Y.; and Zhang, Z. 2025. DrivingGPT : Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers. In ICCV
2025
-
[91]
Chitta, K.; Prakash, A.; Jaeger, B.; Yu, Z.; Renz, K.; and Geiger, A. 2023. TransFuser : Imitation With Transformer-Based Sensor Fusion for Autonomous Driving. IEEE TPAMI
2023
-
[92]
Dang, C.; Ang, S.; Li, Y.; Tian, H.; Wang, J.; Li, G.; Ye, H.; Ma, J.; Chen, L.; and Wang, Y. 2026. DriveFine : Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving. In ECCV
2026
-
[93]
Dauner, D.; Hallgarten, M.; Li, T.; Weng, X.; Huang, Z.; Yang, Z.; Li, H.; Gilitschenski, I.; Ivanovic, B.; Pavone, M.; Geiger, A.; and Chitta, K. 2024. NAVSIM : Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking. In NeurIPS
2024
-
[94]
T.; Genova, K.; Kannen, N.; Ben, S.; Li, Y.; Guo, M.; Yogin, S.; Gu, Y.; Chen, H.; Wang, O.; Xie, S.; Zhou, H.; He, K.; Funkhouser, T.; Alayrac, J.-B.; and Soricut, R
Gabeur, V.; Long, S.; Peng, S.; Voigtlaender, P.; Sun, S.; Bao, Y.; Truong, K.; Wang, Z.; Zhou, W.; Barron, J. T.; Genova, K.; Kannen, N.; Ben, S.; Li, Y.; Guo, M.; Yogin, S.; Gu, Y.; Chen, H.; Wang, O.; Xie, S.; Zhou, H.; He, K.; Funkhouser, T.; Alayrac, J.-B.; and Soricut, R...
2026 arXiv
-
[95]
Gao, S.; Yang, J.; Chen, L.; Chitta, K.; Qiu, Y.; Geiger, A.; Zhang, J.; and Li, H. 2024. Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability. In NeurIPS
2024
-
[96]
Guo, J.; Li, Q.; Li, P.; Chen, Z.; Sun, N.; Su, Y.; Wang, H.; Zhang, Y.; Li, X.; and Liu, H. 2026. Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising. arXiv:2604.26694
2026 arXiv
-
[97]
Han, X.; Li, J.; Deng, K.; Chen, Z.; Shi, X.; Wang, S.; Li, B.; Wang, L.; Xie, S.; You, X.; Quan, J.; Cai, Z.; Diao, H.; Liu, Z.; Yang, L.; Lin, D.; and Wang, Q. 2026. Vision as Unified Multimodal Generation. arXiv:2607.06560
2026 arXiv
-
[98]
Hong, Y.; Zhou, X.; Li, Y.; Zhou, X.; Liu, L.; Luo, Y.; Xu, S.; Yang, L.; and Song, Z. 2026. DriveFuture : Future-Aware Latent World Models for Autonomous Driving. arXiv:2605.09701
2026 arXiv
-
[99]
Hou, B.; Li, G.; Jia, J.; An, T.; Guo, X.; Leng, S.; Geng, H.; Ze, Y.; Harada, T.; Torr, P.; Mees, O.; Pollefeys, M.; Liu, Z.; Wu, J.; Abbeel, P.; Malik, J.; Du, Y.; and Yang, J. 2026. World Model for Robot Learning: A Comprehensive Survey. arXiv:2605.00080
2026 arXiv
-
[100]
Hu, A.; Russell, L.; Yeo, H.; Murez, Z.; Fedoseev, G.; Kendall, A.; Shotton, J.; and Corrado, G. 2023 a . GAIA-1 : A Generative World Model for Autonomous Driving. arXiv:2309.17080
2023 arXiv
-
[101]
Hu, Y.; Yang, J.; Chen, L.; Li, K.; Sima, C.; Zhu, X.; Chai, S.; Du, S.; Lin, T.; Wang, W.; Lu, L.; Jia, X.; Liu, Q.; Dai, J.; Qiao, Y.; and Li, H. 2023 b . Planning- O riented Autonomous Driving. In CVPR
2023
-
[102]
Huang, W.; Zhang, S.; Huang, Q.; Wang, Z.; Mao, Z.; Chua, C.; Chen, Z.; Chen, L.; and Lv, C. 2026. AutoMoT : A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving. In ICML
2026
-
[103]
Hwang, J.-J.; Xu, R.; Lin, H.; Hung, W.-C.; Ji, J.; Choi, K.; Huang, D.; He, T.; Covington, P.; Sapp, B.; Zhou, Y.; Guo, J.; Anguelov, D.; and Tan, M. 2025. EMMA : End-to-End Multimodal Model for Autonomous Driving. TMLR
2025
-
[104]
Jia, F.; Liu, L.; Song, Z.; Jia, C.; Ye, H.; Hao, X.; and Chen, L. 2026. DriveWorld-VLA : Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving. In ICML
2026
-
[105]
Jiang, A.; Gao, Y.; Wang, Y.; Sun, Z.; Wang, S.; Heng, Y.; Sun, H.; Tang, S.; Zhu, L.; Chai, J.; Wang, J.; Gu, Z.; Jiang, H.; and Sun, L. 2025. IRL-VLA : Training an Vision-Language-Action Policy via Reward World Model. arXiv:2508.06571
2025 arXiv
-
[106]
Jiang, B.; Chen, S.; Liao, B.; Zhang, X.; Yin, W.; Zhang, Q.; Huang, C.; Liu, W.; and Wang, X. 2024. Senna : Bridging Large Vision-Language Models and End-to-End Autonomous Driving. arXiv:2410.22313
2024 arXiv
-
[107]
Jiang, B.; Chen, S.; Xu, Q.; Liao, B.; Chen, J.; Zhou, H.; Zhang, Q.; Liu, W.; Huang, C.; and Wang, X. 2023. VAD : Vectorized Scene Representation for Efficient Autonomous Driving. In ICCV
2023
-
[108]
Li, J.; Liu, Z.; Hu, D.; Wu, J.; Ma, Z.; Wu, W.; Han, C.; Hao, Z.; Liu, Z.; Zhan, K.; Deng, J.; Zhu, X.; and Zhang, L. 2026 a . Metis : A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation. arXiv:2606.15869
2026
-
[109]
Li, J.; Wu, J.; Hu, D.; Huang, X.; Sun, B.; Hao, Z.; Lang, X.; Zhu, X.; and Zhang, L. 2026 b . SGDrive : Scene-to-Goal Hierarchical World Cognition for Autonomous Driving. In CVPR
2026
-
[110]
Li, J.; Zhang, B.; Jin, X.; Deng, J.; Zhu, X.; and Zhang, L. 2026 c . ImagiDrive : A Unified Imagination-and-Planning Framework for Autonomous Driving. In ICRA
2026
-
[111]
Li, K.; Li, Z.; Lan, S.; Xie, Y.; Zhang, Z.; Liu, J.; Wu, Z.; Yu, Z.; and Alvarez, J. M. 2025 a . Hydra-MDP++ : Advancing End-to-End Driving via Expert-Guided Hydra-Distillation. arXiv:2503.12820
2025 arXiv
-
[112]
Li, Y.; Fan, L.; He, J.; Wang, Y.; Chen, Y.; Zhang, Z.; and Tan, T. 2025 b . Enhancing End-to-End Autonomous Driving with Latent World Model. In ICLR
2025
-
[113]
Li, Y.; Shang, S.; Liu, W.; Zhan, B.; Wang, H.; Wang, Y.; Chen, Y.; Wang, X.; An, Y.; Tang, C.; Hou, L.; Fan, L.; and Zhang, Z. 2026 d . DriveVLA-W0 : World Models Amplify Data Scaling Law in Autonomous Driving. In ICLR
2026
-
[114]
Li, Y.; Wei, X.; Cao, J.; Wang, H.; Chi, X.; Bai, C.; Sun, Q.; Li, J.; Zhang, X.; Jia, P.; Tang, J.; Han, S.; and Zhang, S. 2026 e . WAM4D : Fast 4D World Action Model via Spatial Register Tokens. arXiv:2606.14048
2026 arXiv
-
[115]
Li, Y.; Xiong, K.; Guo, X.; Li, F.; Yan, S.; Xu, G.; Zhou, L.; Chen, L.; Sun, H.; Wang, B.; Ma, K.; Chen, G.; Ye, H.; Liu, W.; and Wang, X. 2026 f . ReCogDrive : A Reinforced Cognitive Framework for End-to-End Autonomous Driving. In ICLR
2026
-
[116]
Liang, D.; Zhang, D.; Zhou, X.; Tu, S.; Feng, T.; Li, X.; Zhang, Y.; Du, M.; Tan, X.; and Bai, X. 2026. UniFuture : A 4D Driving World Model for Future Generation and Perception. In ICRA
2026
-
[117]
Liang, W.; Yu, L.; Luo, L.; Iyer, S.; Dong, N.; Zhou, C.; Ghosh, G.; Lewis, M.; Yih, W.-t.; Zettlemoyer, L.; and Lin, X. V. 2025. Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models. TMLR
2025
-
[118]
Liao, B.; Chen, S.; Yin, H.; Jiang, B.; Wang, C.; Yan, S.; Zhang, X.; Li, X.; Zhang, Y.; Zhang, Q.; and Wang, X. 2025. DiffusionDrive : Truncated Diffusion Model for End-to-End Autonomous Driving. In CVPR
2025
-
[119]
H.; Chen, D
Lin, H.; Chen, S.; Liew, J. H.; Chen, D. Y.; Li, Z.; Zhao, Y.; Peng, S.; Guo, H.; Zhou, X.; Shi, G.; Feng, J.; and Kang, B. 2026. Depth Anything 3: Recovering the Visual Space from Any Views. In ICLR
2026
-
[120]
Liu, L.; Jia, C.; Yu, G.; Song, Z.; Li, J.; Jia, F.; Wu, P.; Hao, X.; and Luo, Y. 2026. GuideFlow : Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving. In CVPR
2026
-
[121]
Luiten, J.; Osep, A.; Dendorfer, P.; Torr, P. H. S.; Geiger, A.; Leal-Taix \'e , L.; and Leibe, B. 2021. HOTA : A Higher Order Metric for Evaluating Multi-Object Tracking. IJCV
2021
-
[122]
Ma, F.; Peng, D.; Yue, W.; Cao, J.; Wang, B.; Zhang, Q.; and Ma, J. 2026. GeoSem-WAM : Geometry- and Semantic-Aware World Action Models. arXiv:2606.03188
2026 arXiv
-
[123]
NVIDIA ; et al. 2025. Cosmos World Foundation Model Platform for Physical AI . arXiv:2501.03575
2025 arXiv
-
[124]
Rowe, L.; de Schaetzen, R.; Girgis, R.; Pal, C.; and Paull, L. 2025. Poutine : Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving. arXiv:2506.11234
2025
-
[125]
L.; Zhan, W.; and Li, H
Shao, H.; Wang, L.; Zhou, Y.; Hu, Y.; Zong, Z.; Waslander, S. L.; Zhan, W.; and Li, H. 2026. LMGenDrive : Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving. arXiv:2604.08719
2026 arXiv
-
[126]
Sheng, Z.; Ye, X.; Luo, J.; Chen, S.; and Ren, L. 2026. ExploreVLA : Dense World Modeling and Exploration for End-to-End Autonomous Driving. In ECCV
2026
-
[127]
Sun, W.; Lin, X.; Chen, K.; Pei, Z.; Li, X.; Shi, Y.; and Zheng, S. 2026. SparseDriveV2 : Scoring is All You Need for End-to-End Autonomous Driving. arXiv:2603.29163
2026
-
[128]
Sun, W.; Lin, X.; Shi, Y.; Zhang, C.; Wu, H.; and Zheng, S. 2025. SparseDrive : End-to-End Autonomous Driving via Sparse Scene Representation. In ICRA
2025
-
[129]
Team Wan ; et al. 2025. Wan : Open and Advanced Large-Scale Video Generative Models. arXiv:2503.20314
2025 arXiv
-
[130]
Tian, X.; Gu, J.; Li, B.; Liu, Y.; Wang, Y.; Zhao, Z.; Zhan, K.; Jia, P.; Lang, X.; and Zhao, H. 2025. DriveVLM : The Convergence of Autonomous Driving and Large Vision-Language Models. In CoRL
2025
-
[131]
Wang, D.; Song, Y.; He, Z.; Chen, K.; Pan, X.; Deng, L.; and Gu, W. 2025 a . HMVLM : Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios. arXiv:2506.05883
2025 arXiv
-
[132]
Wang, L.; Yang, Z.; Bai, C.; Zhang, G.; Liu, X.; Zheng, X.; Long, X.-X.; Lu, C.-T.; and Lu, C. 2026 a . Drive-JEPA : Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving. arXiv:2601.22032
2026 arXiv
-
[133]
G.; Zanfir, A.; and Sminchisescu, C
Wang, L.; Zhang, C.; Kabra, R.; Uijlings, J.; Waslander, S.; Zisserman, A.; Carreira, J.; He, K.; Andriluka, M.; Bazavan, E. G.; Zanfir, A.; and Sminchisescu, C. 2026 b . Video Generation Models are General-Purpose Vision Learners. In ECCV
2026
-
[134]
Wang, S.; Yu, Z.; Jiang, X.; Lan, S.; Shi, M.; Chang, N.; Kautz, J.; Li, Y.; and Alvarez, J. M. 2025 b . OmniDrive : A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning. In CVPR
2025
-
[135]
Wang, X.; Zhu, Z.; Huang, G.; Chen, X.; Zhu, J.; and Lu, J. 2024 a . DriveDreamer : Towards Real-world-driven World Models for Autonomous Driving. In ECCV
2024
-
[136]
Wang, Y.; He, J.; Fan, L.; Li, H.; Chen, Y.; and Zhang, Z. 2024 b . Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving. In CVPR
2024
-
[137]
Xia, T.; Li, Y.; Zhou, L.; Yao, J.; Xiong, K.; Sun, H.; Wang, B.; Ma, K.; Chen, G.; Ye, H.; Liu, W.; and Wang, X. 2026. DriveLaW : Unifying Planning and Video Generation in a Latent Driving World. In CVPR
2026
-
[138]
Xu, J.; Zhong, Z.; Shu, Z.; Jia, M.; Li, M.; Bian, J.-W.; Zhang, Q.; Zhang, K.; Xie, J.; Yang, J.; and Yin, W. 2026 a . EponaV2 : Driving World Model with Comprehensive Future Reasoning. arXiv:2605.14696
2026 arXiv
-
[139]
Xu, R.; Lin, H.; Jeon, W.; Feng, H.; Zou, Y.; Sun, L.; Gorman, J.; Tolstaya, K.; Tang, S.; White, B.; Sapp, B.; Tan, M.; Hwang, J.-J.; and Anguelov, D. 2026 b . WOD-E2E : Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios. In CVPR
2026
-
[140]
M.; and Wu, Z
Yao, W.; Li, Z.; Lan, S.; Wang, Z.; Sun, X.; Alvarez, J. M.; and Wu, Z. 2026. DriveSuprim : Towards Precise Trajectory Selection for End-to-End Planning. In AAAI
2026
-
[141]
Ye, Y.; Zhang, Z.; Lin, J.; Sun, S.; Peng, C.; and Gao, W. 2026. AutoDrive-P ^3 : Unified Chain of Perception--Prediction--Planning Thought via Reinforcement Fine-Tuning. In ICLR
2026
-
[142]
Yuan, T.; Dong, Z.; Liu, Y.; and Zhao, H. 2026. Fast-WAM : Do World Action Models Need Test-time Future Imagination? arXiv:2603.16666
2026 arXiv
-
[143]
Zhang, B.; Song, N.; Li, J.; Zhu, X.; Deng, J.; and Zhang, L. 2025. Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution. In NeurIPS
2025
-
[144]
Zhang, K.; Wang, J.; Gao, S.; Wu, C.; Cao, Y.; Han, S.; Ivanovic, B.; Liu, L.; Pavone, M.; Han, S.; Zhou, D.; and Xie, E. 2026. Fast-dDrive : Efficient Block-Diffusion VLM for Autonomous Driving. arXiv:2605.23163
2026 arXiv
-
[145]
Zhao, Z.; Fu, T.; Wang, Y.; Wang, L.; and Lu, H. 2025. From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction. In NeurIPS
2025
-
[146]
Zhen, H.; Sun, Q.; Zhang, H.; Li, J.; Zhou, S.; Du, Y.; and Gan, C. 2025. Learning 4D Embodied World Models. In ICCV
2025
-
[147]
Zheng, W.; Xia, Z.; Huang, Y.; Zuo, S.; Zhou, J.; and Lu, J. 2024. Doe-1: Closed-Loop Autonomous Driving with Large World Model. arXiv:2412.09627
2024 arXiv
-
[148]
Zheng, Y.; Yang, P.; Xing, Z.; Zhang, Q.; Zheng, Y.; Gao, Y.; Li, P.; Zhang, T.; Xia, Z.; Jia, P.; Lang, X.; and Zhao, D. 2025. World4Drive : End-to-End Autonomous Driving via Intention-Aware Physical Latent World Model. In ICCV
2025
-
[149]
Zhou, Y.; Wang, X.; Shao, H.; Wang, L.; Zhao, G.; Shao, J.; Zhu, J.; Yu, T.; Zhu, Z.; Huang, G.; and Waslander, S. L. 2026. DriveDreamer-Policy : A Geometry-Grounded World-Action Model for Unified Generation and Planning. arXiv:2604.01765
2026
-
[150]
Zhou, Z.; Cai, T.; Zhao, S.; Zhang, Y.; Huang, Z.; Zhou, B.; and Ma, J. 2025. AutoVLA : A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning. In NeurIPS
2025
-
[151]
Zou, J.; Chen, S.; Liao, B.; Zheng, Z.; Song, Y.; Zhang, L.; Zhang, Q.; Liu, W.; and Wang, X. 2025. DiffusionDriveV2 : Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving. arXiv:2512.07745
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.