REVIEW 3 major objections 5 minor 65 references
FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read FlashDrive reduces a 10B-parameter driving VLA's per-frame inference latency from 717 ms to 151 ms (4.7×) while keeping trajectory error nearly unchanged.
desk verdict Real measured speedups with a plausible near-lossless claim, but missing error bars and a thin empirical basis for the adaptive step policy keep it just short of fully convincing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the four-stage decomposition of VLA inference and the matching shortcut for each stage: (1) streaming KV-cache reuse, which reduces effective sequence length by 75% and requires a streaming attention mask plus pre-RoPE key storage; (2) DFlash, a diffusion-based non-autoregressive drafter trained on driving-domain chain-of-causation tokens, which produces candidate blocks in one forward pass; (3) adaptive-step flow matching, which exploits the U-shaped velocity profile to cache four of eight denoising steps; and (4) CUDA Graphs with fused QKV and gate-up projections, which removes CPU launch overhead. These shortcut identities—streaming mask, pre-RoPE caching, the diffusion drafter, and velocity caching—are what carry the 4.7× compounding gain.
What would settle it
Measure the flow-matching velocity profile on a large, diverse set of clips and compare minADE6 with the four middle denoising steps cached versus recomputed: if the U-shape does not hold on scenes with sudden braking or sharp turns, or if the minADE6 gap exceeds the reported roughly 0.04 m, the adaptive-step claim would be falsified. A second check would run a continuous rollout much longer than the training windows and test whether minADE1 stops improving as streaming KV-cache approximation accumulates.
Extended reading notes
Core claim
FlashDrive's central claim is that all four stages of a reasoning VLA for driving—visual encoding, language prefill, autoregressive decoding of chain-of-causation tokens, and flow-matching trajectory denoising—contain a distinct form of redundancy that can be removed almost without accuracy loss. Concretely, only the newest frame of a sliding window needs to be encoded; the previous frames' KV entries can be reused if keys are stored pre-RoPE and rotated on the fly at shifted positions. A two-layer diffusion drafter generates candidate reasoning blocks in parallel, and with an average accepted length of 5.6 tokens it cuts decode latency 4.7× over the unoptimized baseline. The flow-matching velocity field is U-shaped across denoising steps, sharp at the endpoints and flat in the middle, so four of eight intermediate velocity evaluations can be cached and reused. Together with CUDA Graph compilation, kernel fusion, and W4A8 quantization, these tricks reduce end-to-end latency from 716.9 ms to 151.4 ms on an RTX PRO 6000 while minADE6@6.4s degrades only from 0.767 m to 0.844 m and minADE1 improves from 1.705 m to 1.573 m.
Load-bearing premise
The near-lossless adaptive-step flow-matching result rests on the velocity profile measured from only 10 clips with 20 window inputs each; if that U-shape is not representative of deployment scenes, skipping the same four middle denoising steps for every input could produce trajectory error well beyond the reported 0.04 m minADE6 degradation.
Editorial extensions
If this is right
- A single RTX PRO 6000 can run a 10B-parameter reasoning VLA at 6.6 Hz instead of 1.4 Hz, bringing end-to-end driving within the replanning rates used in urban driving.
- The 4.7× speedup carries to other GPUs: 4.0× on Jetson Thor with one trajectory sample, and 9.6×–10.6× on edge through workstation GPUs with six trajectory samples.
- W4A8 quantization is the appropriate compression regime for VLA inference because 8-bit activations accelerate the compute-bound prefill, whereas weight-only W4A16 quantization would leave prefill untouched.
- Streaming fine-tuning of only the action expert, not the VLM, recovers the trajectory accuracy lost to KV-cache approximation and appears to act as a regularizer that improves minADE1.
- In closed-loop AlpaSim evaluation, collision and off-road rates drop (0.19→0.15 and 0.41→0.32) while relative progress is unchanged, indicating that the latency reduction does not degrade safety.
Reading between the lines
- Because the four speedups are roughly multiplicative, the same profile-each-stage-and-exploit-its-redundancy recipe should transfer to other cascade pipelines, such as long-horizon video-language agents, with the largest remaining gain concentrated where autoregressive decode dominates.
- The fixed skip schedule in adaptive-step flow matching is an implementation choice rather than a learned policy; a scene-conditioned or confidence-based selection of which denoising steps to cache could preserve more accuracy in edge cases, but that variant is not tested in the paper.
- If a future VLA uses free-form rather than templated reasoning tokens, the low entropy that makes DFlash effective would weaken, so the decode speedup would need to be re-derived for that distribution rather than assumed.
- The streaming KV-cache approximation is only shown to be recoverable within the rollout lengths used in training; longer continuous drives may reveal where the action-expert regularizer saturates.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FlashDrive, an algorithm-system co-design framework that accelerates vision-language-action (VLA) inference for autonomous driving by attacking four pipeline stages simultaneously. On Alpamayo 1.5-10B, the authors combine streaming KV-cache reuse (§3.1), DFlash diffusion-based speculative decoding (§3.2), adaptive-step flow matching (§3.3), CUDA Graph compilation and kernel fusion (§3.5), and W4A8 quantization (§3.4). They report a 4.7x end-to-end latency reduction from 716.9 ms to 151.4 ms on an RTX PRO 6000, consistent speedups on four other GPUs, and approximately unchanged open-loop minADE6 (0.767 to 0.844 m) with improved minADE1 (1.705 to 1.573 m). Closed-loop results in AlpaSim show improved collision and off-road rates but regressions in Dtraj and Wrong Lane. The paper also reports a 2.5x per-step rollout speedup in closed-loop simulation. The central claim is that the speedup is achieved with negligible accuracy loss, making real-time VLA driving feasible.
Significance. If the central claim holds, FlashDrive is a valuable systems contribution. The latency measurements are direct and the per-stage ablations in Table 1 add consistently to the end-to-end total, which supports the claim that the four bottlenecks compound. The code and pre-trained checkpoints are promised, and the evaluation covers multiple hardware classes including edge devices. The main caveat is that the 'near-lossless' accuracy claim currently rests on single-run point estimates without confidence intervals, and the adaptive-step flow matching component is justified by velocity profiles from only 10 clips. These issues are fixable with additional analysis and do not undermine the speedup measurements themselves.
major comments (3)
- [§3.3, Fig. 5, Table 1] The adaptive-step flow-matching policy skips the same four middle denoising steps for every input, but the velocity-profile evidence (Fig. 5) is from only 10 clips with 20 window inputs per clip. The paper reports no variation of the U-shaped profile across clips, and Table 1 reports only aggregate minADE6, which is a minimum over six trajectories and can therefore hide large per-sample errors. The closed-loop regressions in Table 3 (Dtraj 20.0→22.4 m, Wrong Lane 0.45→0.51) are consistent with a subset of scenes degrading more than the open-loop mean suggests. Without per-clip accuracy breakdowns or a robustness analysis of the fixed skip, the 'near-lossless' claim for this component is not established.
- [§4.2, Table 1] All efficiency and accuracy numbers in Table 1 are single-run point estimates. The accuracy differences central to the paper (minADE6 +0.077 m, minADE1 −0.132 m) are small enough that they could lie within run-to-run or seed-to-seed variation, yet no error bars, multiple seeds, or confidence intervals are reported. The paper should provide mean±std over at least three seeds for the final comparison (and ideally for the ablations), with bootstrapped confidence intervals over the 12k evaluation windows, before 'essentially unchanged' can be substantiated.
- [§4.4, Table 3] The closed-loop safety interpretation is only partially supported by the metrics. While Collision and Off Road improve, Wrong Lane increases (0.45→0.51) and Dtraj increases (20.0→22.4 m). The text argues that Wrong Lane is noisy near intersections, but no episode-level statistics or significance tests are provided. To support the conclusion that 'acceleration does not compromise driving safety,' the paper should report the distribution of event rates across episodes and, if possible, a statistical test for the differences.
minor comments (5)
- [Table 1] The token throughput (tokens/s) is embedded in the Decode column; consider a separate subcolumn or annotation to avoid ambiguity.
- [§A.1, Table A1a] The row label 'w/ f.t. AE' appears misaligned; check the formatting.
- [§2] The related work discusses KV caching for streaming text generation (Xiao et al., 2024), but no comparison with recent token-caching methods for VLA (e.g., VLA-Cache, Xu et al., 2025a) is given; a brief quantitative or qualitative comparison would clarify the novelty of the streaming scheme.
- [Table 3] The closed-loop latency in Table 3 includes simulator rendering and trajectory optimization, so it is not directly comparable to the Table 1 inference latency; please state this explicitly in the table caption.
- [§3.2] The draft block size of 8 yields an average accepted length of 5.6 tokens, but the acceptance rate and verification cost are not reported; reporting these would help assess the DFlash contribution.
Circularity Check
No significant circularity: FlashDrive's speedup and near-lossless accuracy claims are direct measurements, not derivations from fitted inputs.
full rationale
The paper's central claims are empirical measurements: end-to-end latency drops from 716.9 ms to 151.4 ms, minADE6@6.4s changes from 0.767 m to 0.844 m, and minADE1@6.4s improves from 1.705 m to 1.573 m on the externally defined Alpamayo 1.5-10B model and NVIDIA AV dataset. No claimed result is equivalent to its inputs by construction. The adaptive-step flow-matching skip pattern is chosen after inspecting the U-shaped velocity profile on 10 clips / 200 windows (Sec. 3.3, Fig. 5), but the near-lossless accuracy figure is a separate measurement on 100 clips / 12k windows (Sec. 4.1, Table 1), not an algebraic consequence of the profile; at most this is small-sample model selection, a statistical limitation rather than circularity. DFlash (Chen et al., 2026) and ParoQuant (Liang et al., 2026) are prior works with overlapping authors, but they are not cited as evidence for the speedup: the paper trains its own two-layer drafter and reports a measured accepted length of 5.6 tokens and 58.2 ms decode, and it applies ParoQuant and measures the resulting 24.6 ms reduction. These citations are therefore not load-bearing for the central result. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via citation, and no known result is renamed as a derivation. The weakest point, the representativeness of the 10-clip velocity profile, is a generalization risk, not a circular step.
Assumptions & free parameters
free parameters (3)
- draft block size =
8
- cached flow-matching steps =
4 of 8
- fused hidden-state count for drafter =
8
assumptions (3)
- domain assumption Driving-domain reasoning tokens are short, template-like, low-entropy, and intra-block correlated.
- domain assumption The flow-matching velocity field changes sharply at the endpoints and is nearly constant in the middle.
- domain assumption Streaming KV-cache approximation error mainly affects the action expert and can be compensated by fine-tuning only the action expert.
Cite this review
Pith. "Pith review of FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving." pith.science (2026). https://pith.science/paper/D3MLEUGM
@misc{pith2026260812932,
author = {Pith},
title = {Pith review of: FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving},
year = {2026},
howpublished = {\url{https://pith.science/paper/D3MLEUGM}},
note = {Machine review of arXiv:2608.12932}
}
read the original abstract
Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.
Figures
Reference graph
Works this paper leans on
-
[1]
Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xi...
work page 2025
-
[3]
DFlash: Block Diffusion for Flash Speculative Decoding
Jian Chen, Yesheng Liang, and Zhijian Liu. DFlash: Block Diffusion for Flash Speculative Decoding . In International Conference on Machine Learning (ICML), 2026
work page 2026
-
[4]
Erfei Cui, Wenhai Wang, Zhiqi Li, Jiangwei Xie, Haoming Zou, Hanming Deng, Gen Luo, Lewei Lu, Xizhou Zhu, and Jifeng Dai. DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving . Visual Intelligence, 2025
work page 2025
-
[6]
Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Chong Zhang, Dingyuan Zhang, Hongwei Xie, Bing Wang, and Xiang Bai. ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation . In IEEE/CVF International Conference on Computer Vision (ICCV), 2025
work page 2025
-
[7]
EMMA: End-to-End Multimodal Model for Autonomous Driving
Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, Yin Zhou, James Guo, Dragomir Anguelov, and Mingxing Tan. EMMA: End-to-End Multimodal Model for Autonomous Driving . Transactions on Machine Learning Research (TMLR), 2025
work page 2025
-
[8]
OpenVLA: An Open-Source Vision-Language-Action Model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: An Open-Source Vision-Language-Action Model . In Conference on Robot Learning ...
work page 2024
-
[9]
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success . In Robotics: Science and Systems (RSS), 2025
work page 2025
-
[10]
Isabel Leal, Krzysztof Choromanski, Deepali Jain, Kumar Avinava Dubey, Jake Varley, Michael S. Ryoo, Yao Lu, Frederick Liu, Vikas Sindhwani, Quan Ho Vuong, Tam \'a s Sarl \'o s, Kenneth Oslund, Karol Hausman, and Kanishka Rao. SARA-RT: Scaling Up Robotics Transformers with Self-Adaptive Robust Attention . In IEEE International Conference on Robotics and A...
work page 2024
Show all 65 references
-
[11]
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
Yesheng Liang, Haisheng Chen, Zihan Zhang, Song Han, and Zhijian Liu. ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference . In International Conference on Learning Representations (ICLR), 2026
2026
-
[12]
AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration . In Conference on Machine Learning and Systems (MLSys), 2024
2024
-
[13]
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An, Xiaoqi Li, Kaichen Zhou, Senqiao Yang, Renrui Zhang, Yandong Guo, and Shanghang Zhang. RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation . In Advances in Neural Information Processing Sys...
2024
-
[14]
PhysicalAI Autonomous Vehicles Dataset , 2025
NVIDIA . PhysicalAI Autonomous Vehicles Dataset , 2025. URL https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles
2025
-
[15]
Alpasim: A modular, lightweight, and data-driven research simulator for autonomous driving, October 2025
NVIDIA, Yulong Cao, Riccardo de Lutio, Sanja Fidler, Guillermo Garcia Cobo, Zan Gojcic, Maximilian Igl, Boris Ivanovic, Peter Karkus, Janick Martinez Esturo, Marco Pavone, Aaron Smith, Ellie Tanimura, Michal Tyszkiewicz, Michael Watson, Qi Wu, and Le Zhang. Alpasim: A modular,...
2025
-
[16]
VLP: Vision Language Planning for Autonomous Driving
Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik, Alessandro G Allievi, Senem Velipasalar, and Liu Ren. VLP: Vision Language Planning for Autonomous Driving . In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[17]
FAST: Efficient Action Tokenization for Vision-Language-Action Models
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. FAST: Efficient Action Tokenization for Vision-Language-Action Models . In Robotics: Science and Systems (RSS), 2025
2025
-
[18]
Waslander, Yu Liu, and Hongsheng Li
Hao Shao, Yuxuan Hu, Letian Wang, Steven L. Waslander, Yu Liu, and Hongsheng Li. LMDrive: Closed-Loop End-to-End Driving with Large Language Models . In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[19]
DriveLM: Driving with Graph Visual Question Answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens Bei wenger, Ping Luo, Andreas Geiger, and Hongyang Li. DriveLM: Driving with Graph Visual Question Answering . In European Conference on Computer Vision (ECCV), 2024
2024
-
[20]
GeRM: A Generalist Robotic Model with Mixture-of-Experts for Quadruped Robot
Wenxuan Song, Han Zhao, Pengxiang Ding, Can Cui, Shangke Lyu, Yaning Fan, and Donglin Wang. GeRM: A Generalist Robotic Model with Mixture-of-Experts for Quadruped Robot . In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024
2024
-
[21]
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
Wenxuan Song, Jiayi Chen, Pengxiang Ding, Han Zhao, Wei Zhao, Zhide Zhong, Zongyuan Ge, Zhijun Li, Donglin Wang, Jun Ma, Lujia Wang, and Haoang Li. PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding . arXiv preprint arXiv:25...
2025
-
[22]
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models
Xudong Tan, Yaoxin Yang, Peng Ye, Jialin Zheng, Bizhe Bai, Xinyi Wang, Jia Hao, and Tao Chen. Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models . arXiv preprint arXiv:2505.21200, 2025
2025 arXiv
-
[23]
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, Xianpeng Lang, and Hang Zhao. DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models . In Conference on Robot Learning (CoRL), 2024
2024
-
[24]
BitVLA: 1-Bit Vision-Language-Action Models for Robotics Manipulation
Hongyu Wang, Chuyan Xiong, Ruiping Wang, and Xilin Chen. BitVLA: 1-Bit Vision-Language-Action Models for Robotics Manipulation . arXiv preprint arXiv:2506.07530, 2025 a
2025
-
[25]
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
Yan Wang, Wenjie Luo, Junjie Bai, Yulong Cao, Tong Che, Ke Chen, Yuxiao Chen, Jenna Diamond, Yifan Ding, Wenhao Ding, Liang Feng, Greg Heinrich, Jack Huang, Peter Karkus, Boyi Li, Pinyi Li, Tsung-Yi Lin, Dongran Liu, Ming-Yu Liu, Langechuan Liu, Zhijian Liu, Jason Lu, Yunxiang...
2025 arXiv
-
[26]
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, Yaxin Peng, Feifei Feng, and Jian Tang. TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation . IEEE Robotics and Automation Letters, 2025
2025
-
[27]
Efficient Streaming Language Models with Attention Sinks
Guangxuan Xiao, Yuandong Tang, Juntao Zuo, Junxian Guo, Shang Yang, Haotian Tang, Jinlong Fu, and Song Han. Efficient Streaming Language Models with Attention Sinks . In International Conference on Learning Representations (ICLR), 2024
2024
-
[28]
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
Siyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang, and Chang Xu. VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching . In Advances in Neural Information Processing Systems (NeurIPS), 2025 a
2025
-
[29]
KV-Efficient VLA: A Method to Speed Up Vision Language Models with RNN-Gated Chunked KV Cache
Wanshun Xu, Long Zhuang, and Lianlei Shan. KV-Efficient VLA: A Method to Speed Up Vision Language Models with RNN-Gated Chunked KV Cache . arXiv preprint arXiv:2509.21354, 2025 b
2025
-
[30]
Wong, Zhenguo Li, and Hengshuang Zhao
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee K. Wong, Zhenguo Li, and Hengshuang Zhao. DriveGPT4: Interpretable End-to-End Autonomous Driving via Large Language Model . IEEE Robotics and Automation Letters, 2024
2024
-
[31]
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
Yang Yue, Yulin Wang, Bingyi Kang, Yizeng Han, Shenzhi Wang, Shiji Song, Jiashi Feng, and Gao Huang. DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution . In Advances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[32]
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
Jianke Zhang, Yanjiang Guo, Xiaoyu Chen, Yen-Jen Wang, Yucheng Hu, Chengming Shi, and Jianyu Chen. HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers . In Conference on Robot Learning (CoRL), 2024
2024
-
[33]
MoLe-VLA: Dynamic Layer-Skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
Rongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng, Xiaowei Chi, Gaole Dai, Li Du, Yuan Du, and Shanghang Zhang. MoLe-VLA: Dynamic Layer-Skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation . In AAAI Conference on Artificial Intelligenc...
2026
-
[34]
Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma
Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma. AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning . In Advances in Neural Information Processing Systems (Neur...
2025
-
[35]
Xiao, Guangxuan and Tang, Yuandong and Zuo, Juntao and Guo, Junxian and Yang, Shang and Tang, Haotian and Fu, Jinlong and Han, Song , booktitle =
-
[36]
Chen, Jian and Liang, Yesheng and Liu, Zhijian , booktitle =
-
[37]
Wang, Yan and Luo, Wenjie and Bai, Junjie and Cao, Yulong and Che, Tong and Chen, Ke and Chen, Yuxiao and Diamond, Jenna and Ding, Yifan and Ding, Wenhao and Feng, Liang and Heinrich, Greg and Huang, Jack and Karkus, Peter and Li, Boyi and Li, Pinyi and Lin, Tsung-Yi and Liu, ...
-
[38]
Lin, Ji and Tang, Jiaming and Tang, Haotian and Yang, Shang and Chen, Wei-Ming and Wang, Wei-Chen and Xiao, Guangxuan and Dang, Xingyu and Gan, Chuang and Han, Song , booktitle =
-
[39]
Cui, Erfei and Wang, Wenhai and Li, Zhiqi and Xie, Jiangwei and Zou, Haoming and Deng, Hanming and Luo, Gen and Lu, Lewei and Zhu, Xizhou and Dai, Jifeng , journal =
-
[40]
European Conference on Computer Vision (ECCV) , year =
Sima, Chonghao and Renz, Katrin and Chitta, Kashyap and Chen, Li and Zhang, Hanxue and Xie, Chengen and Bei. European Conference on Computer Vision (ECCV) , year =
-
[41]
Hwang, Jyh-Jing and Xu, Runsheng and Lin, Hubert and Hung, Wei-Chih and Ji, Jingwei and Choi, Kristy and Huang, Di and He, Tong and Covington, Paul and Sapp, Benjamin and Zhou, Yin and Guo, James and Anguelov, Dragomir and Tan, Mingxing , journal =
-
[42]
and Liu, Yu and Li, Hongsheng , booktitle =
Shao, Hao and Hu, Yuxuan and Wang, Letian and Waslander, Steven L. and Liu, Yu and Li, Hongsheng , booktitle =
-
[43]
and Li, Zhenguo and Zhao, Hengshuang , journal =
Xu, Zhenhua and Zhang, Yujia and Xie, Enze and Zhao, Zhen and Guo, Yong and Wong, Kwan-Yee K. and Li, Zhenguo and Zhao, Hengshuang , journal =
-
[44]
Fu, Haoyu and Zhang, Diankun and Zhao, Zongchuang and Cui, Jianfeng and Liang, Dingkang and Zhang, Chong and Zhang, Dingyuan and Xie, Hongwei and Wang, Bing and Bai, Xiang , booktitle =
-
[45]
and Zhang, Yun and Huang, Zhiyu and Zhou, Bolei and Ma, Jiaqi , booktitle =
Zhou, Zewei and Cai, Tianhui and Zhao, Seth Z. and Zhang, Yun and Huang, Zhiyu and Zhou, Bolei and Ma, Jiaqi , booktitle =
-
[46]
Tian, Xiaoyu and Gu, Junru and Li, Bailin and Liu, Yicheng and Wang, Yang and Zhao, Zhiyong and Zhan, Kun and Jia, Peng and Lang, Xianpeng and Zhao, Hang , booktitle =
-
[47]
Pan, Chenbin and Yaman, Burhaneddin and Nesti, Tommaso and Mallik, Abhirup and Allievi, Alessandro G and Velipasalar, Senem and Ren, Liu , booktitle =
-
[48]
and Lu, Yao and Liu, Frederick and Sindhwani, Vikas and Vuong, Quan Ho and Sarl
Leal, Isabel and Choromanski, Krzysztof and Jain, Deepali and Dubey, Kumar Avinava and Varley, Jake and Ryoo, Michael S. and Lu, Yao and Liu, Frederick and Sindhwani, Vikas and Vuong, Quan Ho and Sarl. IEEE International Conference on Robotics and Automation (ICRA) , year =
-
[49]
Xu, Wanshun and Zhuang, Long and Shan, Lianlei , journal =
-
[50]
Liu, Jiaming and Liu, Mengzhen and Wang, Zhenyu and An, Pengju and Li, Xiaoqi and Zhou, Kaichen and Yang, Senqiao and Zhang, Renrui and Guo, Yandong and Zhang, Shanghang , booktitle =
-
[51]
Kim, Moo Jin and Finn, Chelsea and Liang, Percy , booktitle =
-
[52]
Wen, Junjie and Zhu, Yichen and Li, Jinming and Zhu, Minjie and Wu, Kun and Xu, Zhiyuan and Liu, Ning and Cheng, Ran and Shen, Chaomin and Peng, Yaxin and Feng, Feifei and Tang, Jian , journal =
-
[53]
Song, Wenxuan and Chen, Jiayi and Ding, Pengxiang and Zhao, Han and Zhao, Wei and Zhong, Zhide and Ge, Zongyuan and Li, Zhijun and Wang, Donglin and Ma, Jun and Wang, Lujia and Li, Haoang , journal =
-
[54]
arXiv preprint arXiv:2507.14049 , year =
Budzianowski, Pawe. arXiv preprint arXiv:2507.14049 , year =
-
[55]
Song, Wenxuan and Zhao, Han and Ding, Pengxiang and Cui, Can and Lyu, Shangke and Fan, Yaning and Wang, Donglin , booktitle =
-
[56]
Zhang, Jianke and Guo, Yanjiang and Chen, Xiaoyu and Wang, Yen-Jen and Hu, Yucheng and Shi, Chengming and Chen, Jianyu , booktitle =
-
[57]
Zhang, Rongyu and Dong, Menghang and Zhang, Yuan and Heng, Liang and Chi, Xiaowei and Dai, Gaole and Du, Li and Du, Yuan and Zhang, Shanghang , booktitle =
-
[58]
Yue, Yang and Wang, Yulin and Kang, Bingyi and Han, Yizeng and Wang, Shenzhi and Song, Shiji and Feng, Jiashi and Huang, Gao , booktitle =
-
[59]
Kim, Moo Jin and Pertsch, Karl and Karamcheti, Siddharth and Xiao, Ted and Balakrishna, Ashwin and Nair, Suraj and Rafailov, Rafael and Foster, Ethan and Lam, Grace and Sanketi, Pannag and Vuong, Quan and Kollar, Thomas and Burchfiel, Benjamin and Tedrake, Russ and Sadigh, Dor...
-
[60]
Wang, Hongyu and Xiong, Chuyan and Wang, Ruiping and Chen, Xilin , journal =
-
[61]
Pertsch, Karl and Stachowicz, Kyle and Ichter, Brian and Driess, Danny and Nair, Suraj and Vuong, Quan and Mees, Oier and Finn, Chelsea and Levine, Sergey , booktitle =
-
[62]
Tan, Xudong and Yang, Yaoxin and Ye, Peng and Zheng, Jialin and Bai, Bizhe and Wang, Xinyi and Hao, Jia and Chen, Tao , journal =
-
[63]
Xu, Siyu and Wang, Yunke and Xia, Chenghao and Zhu, Dihao and Huang, Tao and Xu, Chang , booktitle =
-
[64]
Black, Kevin and Brown, Noah and Darpinian, James and Dhabalia, Karan and Driess, Danny and Esmail, Adnan and Equi, Michael and Finn, Chelsea and Fusai, Niccolo and Galliker, Manuel Y. and Ghosh, Dibya and Groom, Lachy and Hausman, Karol and Ichter, Brian and Jakubczak, Szymon...
-
[65]
Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian , booktitle =
-
[66]
arXiv preprint arXiv:2408.11743 , year=
MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models , author=. arXiv preprint arXiv:2408.11743 , year=
-
[67]
2025 , month =
NVIDIA and Yulong Cao and Riccardo de Lutio and Sanja Fidler and Guillermo Garcia Cobo and Zan Gojcic and Maximilian Igl and Boris Ivanovic and Peter Karkus and Janick Martinez Esturo and Marco Pavone and Aaron Smith and Ellie Tanimura and Michal Tyszkiewicz and Michael Watson...
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.