REVIEW 3 major objections 5 minor 1 cited by
A single-camera planner that routes frozen perception priors per scene reaches the top open-loop driving scores without test-time search.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 10:33 UTC pith:TQRZKL6X
load-bearing objection PerceptDrive is a believable new NAVSIM SOTA with unusually careful internal validation, but part of its gain is explicitly alignment to the NAVSIM rule-based scorer, and the closed-loop transfer question is left open. the 3 major comments →
PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the bottleneck between frozen perception priors and continuous trajectory generation is not the priors themselves but the interface: query compression tends to discard source-specific cues, and static fusion cannot adapt expert contributions per scene. PerceptDrive addresses both with two coupled mechanisms. Per-branch prior retention makes each compressed readout reconstruct its own frozen prior, preventing branch collapse and preserving geometry, semantics, and dynamics through the query bottleneck. Metric-distilled routing uses privileged rule-based sub-metric scores, available only during training, to supervise soft gates that recombine expert conditions at the
What carries the argument
The carrying mechanism is an expert-routed world-action model built around three expert conditioning branches — geometric, semantic, and dynamic — each with its own query bank. Two trainable pieces do the work: a per-branch retention probe that reconstructs the frozen prior's token states from each branch's compressed readout (with stop-gradient targets), and a lightweight router that maps a shared scene vector to soft gates on the simplex, mixing the three expert conditions before trajectory generation. A flow-matching actor is conditioned on the gated mixture plus a self-predicted action-free future latent from a frozen video encoder. During training, privileged rule-based sub-metric score
Load-bearing premise
The load-bearing premise is that the benchmark's rule-based sub-metrics for safety, comfort, compliance, and progress are a valid proxy for real driving quality, because the router and quality regressor are trained to maximize exactly those metrics; if those metrics do not reflect closed-loop or differently weighted objectives, the reported advantage could shrink or disappear.
What would settle it
Run the planner in interactive closed-loop simulation, or re-score it with a differently weighted objective that is not used during training (e.g., a learned safety critic or naturalistic driving costs); if the routing and retention gains shrink or reverse, the improvements are largely an artifact of aligning to the benchmark's specific non-reactive scorer.
If this is right
- A direct single-trajectory planner can match or exceed systems that rely on candidate scoring, reranking, or multiple cameras on open-loop benchmarks.
- Evaluator knowledge can be amortized: distilling privileged sub-metrics into the router during training keeps inference a single feed-forward pass.
- Frozen, heterogeneous perception priors can be combined at the conditioning level rather than through static fusion, with per-branch retention preventing the compression bottleneck from erasing source-specific information.
- Per-branch prior retention induces specialized reads — branch readouts become more distinguishable and reconstruct their own priors better — so additional frozen priors can in principle be added with dedicated branches.
- The ablations show the gains are not solely from learnable weights or trajectory averaging, so routing before generation is the component that carries the improvement.
Where Pith is reading between the lines
- Editorial extension: Because the router is supervised by the same sub-metrics that define the benchmark, part of the reported margin likely reflects alignment to that specific scorer; retraining the gate on a differently weighted objective is a direct way to test how much of the gain is transferable.
- Editorial extension: The gate values could be read as an interpretable per-scene signal of which prior matters most (curves toward geometry, straight roads toward dynamics), suggesting they might be useful for downstream monitoring or explanation, something the paper does not claim.
- Editorial extension: The same retention-plus-routing recipe could apply to other embodied tasks that condition on frozen foundation models — for example manipulation or navigation — but the paper only demonstrates it for driving.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. PerceptDrive proposes a perception-prior world-action model for end-to-end driving. A frozen, driving-adapted VLM provides geometric/semantic/dynamic expert tokens; a frozen V-JEPA encoder provides dense observation latents. A trainable world-action model compresses these streams with query banks, anchors each expert branch to its prior via a retention loss, predicts an action-free future latent, and fuses the expert conditions through a scene-conditioned router. Training additionally uses privileged NAVSIM sub-metrics to supervise a quality regressor and to distill soft-gate targets from one-step branch drafts. At inference the system produces a single trajectory from one front camera without candidate scoring or test-time search. The paper reports 90.4 PDMS on NAVSIM v1 navtest, 90.2 EPDMS on NAVSIM v2 navtest, and 34.5 EPDMS on navhard, with ablations across provider design, routing/fusion, and retention choices.
Significance. If the empirical results hold, this is a strong benchmark contribution: state-of-the-art NAVSIM v1/v2 numbers with single-trajectory, single-camera inference. The paper's support is unusually careful: official evaluator scores, three training seeds, paired bootstrap separation from ablations, surrogate validation of the routing supervision (Spearman 0.91 for k-NN interpolation, 86.2% draft-rollout agreement), hyperparameter sensitivity analysis, and a transparent Discussion of limitations. The conceptual mechanism — prior-retention-anchored expert branches with metric-distilled scene-conditioned routing — is also well motivated. The main caveat is that the routing and quality supervision are trained on exactly the rule-based sub-metrics that define the evaluation score, so part of the gain is evaluator alignment; the paper acknowledges this but does not quantify it, and transfer to closed-loop or differently weighted objectives is untested.
major comments (3)
- [§3.2, Appendix B] Data-provenance check is missing and potentially serious. The provider is full fine-tuned on 1,398,858 samples including NuScenes-QA and DriveQA, which are built on the nuScenes dataset, the same dataset from which NAVSIM navtest/navhard scenes are drawn. The paper does not state that navtest/navhard scenes or their frames were excluded from Stage 1a QA fine-tuning or from the multi-teacher distillation corpus. Since the VLM stream contributes a large share of the final EPDMS (81.9 without the stream vs. 90.2 full, Table A1), even partial overlap could materially inflate the reported SOTA. Please provide a scene-level overlap analysis and, if any test scenes are present in the provider training corpus, rerun with exclusion filtering.
- [§4.3 Discussion, Appendix C Eq. (A7)] The broader claim that PerceptDrive 'improves direct planning' is not cleanly separable from 'optimizes the NAVSIM scorer.' The routing targets are α* = softmax(q/T_r), where q is the mean of privileged sub-metric scores (NC, DAC, DDC, TLC, EP, TTC, LK, EC) evaluated by the same rule-based planner family that defines EPDMS; the auxiliary metric branch uses the same pool. The ablations show that metric-distilled routing gives a real score gain (Table A2), but every internal validation is inside the same scorer family, so the gain could reflect alignment to the benchmark's particular weighting rather than a general improvement in driving quality. The Discussion concedes that transfer to differently weighted objectives and closed-loop evaluation are untested. This is acceptable for a benchmark-specific claim, but the abstract and conclusion should say 'under NAVSIM metrics' rather than impl
- [§3.2, Eq. (5)] Reproduction is incomplete because the four provider-loss weights λ_geo, λ_sem, λ_dyn, λ_rep are never given. Appendix C reports routing/retention weights (λ_f, λ_ret, λ_r) and temperature, but the provider construction is a central component and its weighting is unspecified. Please report the values or the search range used.
minor comments (5)
- [Abstract and §5] The phrase 'improves direct planning' should be qualified as 'improves NAVSIM PDMS/EPDMS' unless transfer evidence is added. This is partly addressed in the Discussion, but the abstract/conclusion currently overstate the scope.
- [Figure 4(b)] The axis label 'uniform 7-point range below Full' is unclear. Please state explicitly what the rings correspond to and how the bootstrap bands were computed.
- [Table A3] The navhard table formatting splits PerceptDrive sub-scores across two stage rows without a per-stage EPDMS column, making the final 34.5 EPDMS attribution ambiguous. Add per-stage EPDMS and mark which stage produces the final score.
- [§3.2] The provider is described as 'frozen' after adaptation, but it was first full fine-tuned and LoRA-adapted. Please clarify in the main text that 'frozen' refers to the deployed inference-time state, to avoid confusion about what is being frozen.
- [§4.1, Table 1] The paper notes that sensor coverage, pretraining, proposal count, and inference budget differ across baselines, but this is only mentioned in passing. A small table or footnote listing these properties for each compared method would improve fairness assessment.
Circularity Check
No circularity: NAVSIM SOTA is an optimized benchmark result; acknowledged evaluator alignment is a generalization caveat, not a self-referential derivation.
full rationale
PerceptDrive's benchmark numbers are obtained by training on navtrain with losses that include metric distillation (Eqs. A4, A7), then evaluating with the official NAVSIM evaluator on navtest/navhard; this is optimizing the stated benchmark objective, not deriving the benchmark from itself. The router target gates alpha* = softmax(q/T_r) use the same PDM sub-metric family as EPDMS, and the paper explicitly states 'part of the gain reflects evaluator alignment', but the reported result is a trained model's external-evaluator score, not a fitted parameter renamed as a prediction. The supervision pools are navtrain-only, and official evaluation is held out. The author self-citations [7,30,56] appear only in related-work surveys and are not used to justify the method or exclude alternatives. No uniqueness theorem, ansatz-by-citation, or definitional equivalence occurs. The transfer-to-closed-loop caveat is a correctness/generalization risk, not circularity.
Axiom & Free-Parameter Ledger
free parameters (6)
- Provider distillation weights λ_geo, λ_sem, λ_dyn, λ_rep =
not stated in provided text
- WAM loss weights λ_f, λ_ret, λ_r =
0.2, 0.1, 0.1
- Routing temperature T_r =
0.5
- Warmup steps T_w =
3000
- Offline pool perturbation scales and count =
lateral σ=0.5 m, longitudinal σ=1.0 m, heading σ=2°, 16 perturbations per scene
- k-NN neighbors for score interpolation =
5
axioms (6)
- domain assumption NAVSIM's non-reactive pseudo-simulation and PDMS/EPDMS rule-based sub-metrics are valid proxies for real driving safety, comfort, and planning quality.
- domain assumption Frozen VGGT, V-JEPA 2, and Wan 2.1 teacher representations encode geometric, semantic, and dynamic knowledge that transfers to driving planning through the adapted VLM.
- domain assumption The frozen V-JEPA 2-L encoder's future-frame latents are a valid self-supervised target for predicting the next 2 s.
- domain assumption Query-bank readouts plus lightweight retention probes can recover enough of the detached expert-slot targets to anchor branch specialization.
- standard math Flow matching with 25 Euler steps yields a good trajectory sample from the learned velocity field.
- domain assumption Training on NAVSIM navtrain generalizes to navtest and navhard under the NAVSIM protocol.
invented entities (2)
-
[GEO], [SEM], [DYN] expert-slot tokens in the adapted VLM
no independent evidence
-
Action-free future latent v̂_free
no independent evidence
read the original abstract
Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as the prior-to-plan transfer problem and introduce PerceptDrive, a perception prior world-action modeling framework with adaptive expert routing. PerceptDrive feeds teacher-distilled priors from a frozen, driving-adapted provider and dense observation latents from a frozen self-supervised video encoder into a trainable expert-routed world-action model. Expert-specific query branches process these signals, while a prior-retention objective anchors each branch to its prior. A router predicts soft gates from a shared scene representation and combines the expert conditions before trajectory generation. During training, privileged rule-based sub-metric estimates for branch-specific trajectory drafts provide soft-gate distillation targets. The predicted action-free future latent conditions a flow-matching actor. At inference, privileged components are absent; with one front-facing camera, PerceptDrive generates one trajectory per planning step without test-time scoring, reranking, or search. Experiments show that PerceptDrive achieves state-of-the-art performance with 90.4 PDMS on NAVSIM v1 and 90.2 EPDMS on NAVSIM v2, outperforming existing methods. Ablations confirm complementary gains from prior retention and scene-conditioned routing, alongside differential reliance on the three priors. These results demonstrate that preserving and adaptively routing perception priors improves direct planning without test-time candidate selection.
Forward citations
Cited by 1 Pith paper
-
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
Action-conditioned world-model verification with conformal first-intervention control and latency-aware suffix repair raises RoboCasa365 success 8.5 points over invocation-matched periodic replanning.
Reference graph
Works this paper leans on
-
[1]
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning, 2025
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, et al. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning, 2025
2025
-
[2]
Pseudo-Simulation for Autonomous Driving
Wei Cao, Marcel Hallgarten, Tianyu Li, Daniel Dauner, Xunjiang Gu, Caojun Wang, Yakov Miron, Marco Aiello, Hongyang Li, Igor Gilitschenski, Boris Ivanovic, Marco Pavone, Andreas Geiger, and Kashyap Chitta. Pseudo-Simulation for Autonomous Driving. InConference on Robot Learning (CoRL), 2025
2025
-
[3]
DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers, 2024
Yuntao Chen, Yuqi Wang, and Zhaoxiang Zhang. DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers, 2024
2024
-
[4]
TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger. TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
2023
-
[5]
NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
Daniel Dauner, Marcel Hallgarten, Tianyu Li, Xinshuo Weng, Zhiyu Huang, Zetong Yang, Hongyang Li, Igor Gilitschenski, Boris Ivanovic, Marco Pavone, Andreas Geiger, and Kashyap Chitta. NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking. InAdvances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[6]
FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder, 2026
Zeyu Dong, Yimin Zhu, Yu Wu, and Yu Sun. FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder, 2026
2026
-
[7]
FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution, 2026
Jingjing Fan, Yushan Liu, Shoujie Li, Botao Ren, Siyuan Li, Xiao-Ping Zhang, Wenbo Ding, and Zhidong Deng. FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution, 2026
2026
-
[8]
ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution, 2026
Chuyao Fu, Shengzhe Gan, Zhuoli Ouyang, Yuhan Rui, Xiaowei Chi, et al. ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution, 2026
2026
-
[9]
Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation, 2026
Xingtai Gui, Meijie Zhang, Tianyi Yan, Wencheng Han, Jiahao Gong, Feiyang Tan, Cheng-zhong Xu, and Jianbing Shen. Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation, 2026
2026
-
[10]
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving, 2025
Jianhua Han, Meng Tian, Jiangtong Zhu, Fan He, Huixin Zhang, Sitong Guo, Dechang Zhu, Hao Tang, Pei Xu, Yuze Guo, Minzhe Niu, Haojie Zhu, Qichao Dong, Xuechao Yan, Siyuan Dong, Lu Hou, Qingqiu Huang, Xiaosong Jia, and Hang Xu. Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving, 2025
2025
-
[11]
DriveFuture: Future-Aware Latent World Models for Autonomous Driving, 2026
Yufeng Hong, Xiaotian Zhou, Yingyan Li, Xiangpo Zhou, Lin Liu, Yadan Luo, Shaoqing Xu, Lei Yang, and Ziying Song. DriveFuture: Future-Aware Latent World Models for Autonomous Driving, 2026
2026
-
[12]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models, 2021
2021
-
[13]
Planning-oriented Autonomous Driving
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, Lewei Lu, Xiaosong Jia, Qiang Liu, Jifeng Dai, Yu Qiao, and Hongyang Li. Planning-oriented Autonomous Driving. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[14]
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving, 2026
Minqing Huang, Yujiao Xiang, Zihan Liang, Jiajie Huang, Jingqi Wang, Zhi Xu, Feiyang Tan, Hangning Zhou, Mu Yang, and Gong Che. CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving, 2026
2026
-
[15]
Jacobs, Michael I
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. Adaptive Mixtures of Local Experts.Neural Computation, 3(1):79–87, 1991
1991
-
[16]
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving, 2026
Feiyang Jia, Lin Liu, Ziying Song, Caiyan Jia, Hangjun Ye, Xiaoshuai Hao, and Long Chen. DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving, 2026
2026
-
[17]
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning, 2024
Bo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao, Qian Zhang, Wenyu Liu, and Xinggang Wang. VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning, 2024
2024
-
[18]
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving, 2024
Bo Jiang, Shaoyu Chen, Bencheng Liao, Xingyu Zhang, Wei Yin, Qian Zhang, Chang Huang, Wenyu Liu, and Xinggang Wang. Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving, 2024. 11
2024
-
[19]
Enhancing End-to-End Autonomous Driving with Latent World Model
Yingyan Li, Lue Fan, Jiawei He, Yuqi Wang, Yuntao Chen, Zhaoxiang Zhang, and Tieniu Tan. Enhancing End-to-End Autonomous Driving with Latent World Model. InThe Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/forum?id=fd2u60ryG0
2025
-
[20]
DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving, 2025
Yingyan Li, Shuyao Shang, Weisong Liu, Bing Zhan, Haochen Wang, Yuqi Wang, Yuntao Chen, Xiaoman Wang, Yasong An, Chufeng Tang, Lu Hou, Lue Fan, and Zhaoxiang Zhang. DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving, 2025
2025
-
[21]
End-to-End Driving with Online Trajectory Evaluation via BEV World Model, 2025
Yingyan Li, Yuqi Wang, Yang Liu, Jiawei He, Lue Fan, and Zhaoxiang Zhang. End-to-End Driving with Online Trajectory Evaluation via BEV World Model, 2025
2025
-
[22]
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving, 2025
Yongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li, Sixu Yan, Gangwei Xu, Lijun Zhou, Long Chen, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, and Xinggang Wang. ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving, 2025. URLhttps://arxiv.org/abs/2506.08052
Pith/arXiv arXiv 2025
-
[23]
Zhenxin Li, Kailin Li, Shihao Wang, Shiyi Lan, Zhiding Yu, Yishen Ji, Zhiqi Li, Ziyue Zhu, Jan Kautz, Zuxuan Wu, Yu-Gang Jiang, and Jose M. Alvarez. Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation, 2024
2024
-
[24]
Zhenxin Li, Wenhao Yao, Zi Wang, Xinglong Sun, Joshua Chen, Nadine Chang, Maying Shen, Zuxuan Wu, Shiyi Lan, and Jose M. Alvarez. Generalized Trajectory Scoring for End-to-end Multimodal Planning, 2025
2025
-
[25]
Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving, 2026
Weitong Lian, Zecong Tang, Haoran Li, Tianjian Gao, Yifei Wang, Zixu Wang, Lingyi Meng, Tengju Ru, Zhejun Cui, Yichen Zhu, Hangshuo Cao, Qi Kang, Tianxing Chen, Kaixuan Wang, and Yu Zhang. Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving, 2026
2026
-
[26]
Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving
Bencheng Liao, Shaoyu Chen, Haoran Yin, Bo Jiang, Cheng Wang, Sixu Yan, Xinbang Zhang, Xiangyu Li, Ying Zhang, Qian Zhang, and Xinggang Wang. Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
2025
-
[27]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow Matching for Generative Modeling. InInternational Conference on Learning Representations (ICLR), 2023
2023
-
[28]
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving, 2026
Qiqi Liu, Huan Xu, Jingyu Li, Bin Sun, Zhihui Hao, Dangen She, Xiatian Zhu, and Li Zhang. Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving, 2026
2026
-
[29]
Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. InInternational Conference on Learning Representations (ICLR), 2023
2023
-
[30]
OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation, 2026
Yushan Liu, Peibo Sun, Shoujie Li, Yifan Xie, Lingfeng Zhang, Xintao Chao, Shiyuan Dong, Fang Chen, Xiao-Ping Zhang, and Wenbo Ding. OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation, 2026
2026
-
[31]
LingoQA: Visual Question Answering for Autonomous Driving
Ana-Maria Marcu, Long Chen, Jan Huenermann, Alice Karnsund, Benoit Hanotte, Prajwal Chidananda, Saurabh Nair, Vijay Badrinarayanan, Alex Kendall, Jamie Shotton, Elahe Arani, and Oleg Sinavski. LingoQA: Visual Question Answering for Autonomous Driving. InEuropean Conference on Computer Vision (ECCV), 2024
2024
-
[32]
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving, 2023
Ming Nie, Renyuan Peng, Chunwei Wang, Xinyue Cai, Jianhua Han, Hang Xu, and Li Zhang. Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving, 2023
2023
-
[33]
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario, 2023
Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao, and Yu-Gang Jiang. NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario, 2023
2023
-
[34]
PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models, 2026
Xiaoyun Qiu, Jingtao He, Yijie Chen, Yusong Huang, Haotian Wang, et al. PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models, 2026
2026
-
[35]
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. InInternational Conference on Learning Representations (ICLR), 2017
2017
-
[36]
DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving, 2026
Chen Shi, Jinrui Xu, Shaoshuai Shi, Kehua Sheng, Bo Zhang, and Li Jiang. DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving, 2026. 12
2026
-
[37]
DriveLM: Driving with Graph Visual Question Answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens Beisswenger, Ping Luo, Andreas Geiger, and Hongyang Li. DriveLM: Driving with Graph Visual Question Answering. InEuropean Conference on Computer Vision (ECCV), 2024
2024
-
[38]
GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving, 2026
Ziying Song, Caiyan Jia, Lin Liu, Lei Yang, Shengkai Zhang, Feiyang Jia, Fengda Zhao, Peiliang Wu, Shaoqing Xu, Chen Lv, and Yadan Luo. GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving, 2026
2026
-
[39]
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving, 2025
Bin Sun, Yaoguang Cao, Yan Wang, Rui Wang, Jiachen Shang, Xiejie Feng, Jiayi Lu, Jia Shi, Shichun Yang, Xiaoyu Yan, and Ziying Song. MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving, 2025
2025
-
[40]
Wan: Open and Advanced Large-Scale Video Generative Models, 2025
Wan Team, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, et al. Wan: Open and Advanced Large-Scale Video Generative Models, 2025
2025
-
[41]
VGGT: Visual Geometry Grounded Transformer, 2025
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. VGGT: Visual Geometry Grounded Transformer, 2025
2025
-
[42]
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving, 2026
Linbo Wang, Yupeng Zheng, Qiang Chen, Shiwei Li, Yichen Zhang, Zebin Xing, Qichao Zhang, Xiang Li, Deheng Qian, Pengxuan Yang, Yihang Dong, Ce Hao, Xiaoqing Ye, Junyu Han, Yifeng Pan, and Dongbin Zhao. Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving, 2026
2026
-
[43]
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving, 2026
Linhan Wang, Zichong Yang, Chen Bai, Guoxiang Zhang, Xiaotong Liu, Xiaoyin Zheng, Xiao-Xiao Long, Chang- Tien Lu, and Cheng Lu. Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving, 2026
2026
-
[44]
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving, 2026
Yiru Wang, Zichong Gu, Yu Gao, Anqing Jiang, Zhigang Sun, Shuo Wang, Yuwen Heng, and Hao Sun. HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving, 2026
2026
-
[45]
DriveQA: Passing the Driving Knowledge Test, 2025
Maolin Wei, Wanzhou Liu, and Eshed Ohn-Bar. DriveQA: Passing the Driving Knowledge Test, 2025
2025
-
[46]
A Unified Candidate Set with Scene-Adaptive Refinement via Diffusion for End-to-End Autonomous Driving, 2026
Zhengfei Wu, Shuaixi Pan, Shuohan Chen, et al. A Unified Candidate Set with Scene-Adaptive Refinement via Diffusion for End-to-End Autonomous Driving, 2026
2026
-
[47]
DriveLaW: Unifying Planning and Video Generation in a Latent Driving World
Tianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao, Kaixin Xiong, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, and Xinggang Wang. DriveLaW: Unifying Planning and Video Generation in a Latent Driving World. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),pages39701–39712, 2026. URL https://ope...
2026
-
[48]
Meyer, Siva Karthik Mustikovela, Siddhartha Srinivasa, Eric M
Yi Xu, Yuxin Hu, Zaiwei Zhang, Gregory P. Meyer, Siva Karthik Mustikovela, Siddhartha Srinivasa, Eric M. Wolff, and Xin Huang. VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision. In Joseph Lim, Shuran Song, and Hae-Won Park, editors,Proceedings of The 9th Conference on Robot Learning, volume 305 ofProceedings of Machine Learni...
2025
-
[49]
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
Zhenjie Yang, Yilin Chai, Xiaosong Jia, Qifeng Li, Yuqian Shao, Xuekai Zhu, Haisheng Su, and Junchi Yan. DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
2026
-
[50]
S2-VLA: Decoupling Semantic and Spatial Streams in Vision- Language-Action Models for Autonomous Driving, 2026
Jianguo Yu, Rukang Wang, Duanfeng Chu, et al. S2-VLA: Decoupling Semantic and Spatial Streams in Vision- Language-Action Models for Autonomous Driving, 2026
2026
-
[51]
Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution, 2025
Bozhou Zhang, Nan Song, Jingyu Li, Xiatian Zhu, Jiankang Deng, and Li Zhang. Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution, 2025
2025
-
[52]
IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving, 2026
Chenghao Zhang, Timin Li, and Dongmei Li. IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving, 2026
2026
-
[53]
Epona: Autoregressive Diffusion World Model for Autonomous Driving, 2025
Kaiwen Zhang, Zhenyu Tang, Xiaotao Hu, Xingang Pan, Xiaoyang Guo, Yuan Liu, Jingwei Huang, Li Yuan, Qian Zhang, Xiao-Xiao Long, Xun Cao, and Wei Yin. Epona: Autoregressive Diffusion World Model for Autonomous Driving, 2025
2025
-
[54]
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction, 2025
Zhida Zhao, Talas Fu, Yifan Wang, Lijun Wang, and Huchuan Lu. From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction, 2025. 13
2025
-
[55]
World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model, 2025
Yupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang, Yuhang Zheng, Yinfeng Gao, Pengfei Li, Teng Zhang, Zhongpu Xia, Peng Jia, and Dongbin Zhao. World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model, 2025
2025
-
[56]
ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving, 2026
Xuchang Zhong, He Zheng, Chenxu Zhao, Tianxiong Lv, Hangqi Fan, Bohua Wang, Yushan Liu, Zhihao Liao, Leigang Luo, Congyang Zhao, and Yang Cai. ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving, 2026
2026
-
[57]
Waslander
Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang, Guosheng Zhao, Jiangnan Shao, Jiagang Zhu, Tingdong Yu, Zheng Zhu, Guan Huang, and Steven L. Waslander. DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning, 2026
2026
-
[58]
Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma
Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma. AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning, 2025
2025
-
[59]
w/o prior distillation
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, et al. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, 2025. 14 Appendix A Evaluation Protocols Data splits.The trainable WAM is optimized on NAVSIM navtrain and evaluated on navtest or navhard according to the pr...
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.