Pith. sign in

REVIEW 3 major objections 5 minor 65 references

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read FlashDrive reduces a 10B-parameter driving VLA's per-frame inference latency from 717 ms to 151 ms (4.7×) while keeping trajectory error nearly unchanged.

desk verdict Real measured speedups with a plausible near-lossless claim, but missing error bars and a thin empirical basis for the adaptive step policy keep it just short of fully convincing. read the letter →

arxiv 2608.12932 v1 pith:D3MLEUGM submitted 2026-08-13 cs.AI

classification cs.AI
keywords vision-language-actionmodelsautonomousdrivingefficientinferencespeculativedecodingflowmatchingKVcachereusequantizationCUDAGraphs
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the latency of vision-language-action (VLA) driving models is a cascade of four independent bottlenecks, not one bottleneck, and that each admits a separate lightweight algorithmic shortcut. It presents FlashDrive, which combines streaming KV-cache reuse across video frames, diffusion-based speculative decoding of reasoning tokens, adaptive caching of flow-matching denoising steps, and system-level CUDA Graph compilation and kernel fusion, plus W4A8 quantization. On the Alpamayo 1.5-10B model, the combined framework shortens per-frame inference from 717 ms to 151 ms (4.7×) on an RTX PRO 6000 while keeping trajectory error nearly unchanged: minADE6@6.4s shifts by about 0.08 m and minADE1 improves. The paper's central thesis is that matching each stage's specific redundancy to its own shortcut compounds gains that no single-stage optimization can achieve.

What carries the argument

The load-bearing machinery is the four-stage decomposition of VLA inference and the matching shortcut for each stage: (1) streaming KV-cache reuse, which reduces effective sequence length by 75% and requires a streaming attention mask plus pre-RoPE key storage; (2) DFlash, a diffusion-based non-autoregressive drafter trained on driving-domain chain-of-causation tokens, which produces candidate blocks in one forward pass; (3) adaptive-step flow matching, which exploits the U-shaped velocity profile to cache four of eight denoising steps; and (4) CUDA Graphs with fused QKV and gate-up projections, which removes CPU launch overhead. These shortcut identities—streaming mask, pre-RoPE caching, the diffusion drafter, and velocity caching—are what carry the 4.7× compounding gain.

What would settle it

Measure the flow-matching velocity profile on a large, diverse set of clips and compare minADE6 with the four middle denoising steps cached versus recomputed: if the U-shape does not hold on scenes with sudden braking or sharp turns, or if the minADE6 gap exceeds the reported roughly 0.04 m, the adaptive-step claim would be falsified. A second check would run a continuous rollout much longer than the training windows and test whether minADE1 stops improving as streaming KV-cache approximation accumulates.

Watch

Extended reading notes

Core claim

FlashDrive's central claim is that all four stages of a reasoning VLA for driving—visual encoding, language prefill, autoregressive decoding of chain-of-causation tokens, and flow-matching trajectory denoising—contain a distinct form of redundancy that can be removed almost without accuracy loss. Concretely, only the newest frame of a sliding window needs to be encoded; the previous frames' KV entries can be reused if keys are stored pre-RoPE and rotated on the fly at shifted positions. A two-layer diffusion drafter generates candidate reasoning blocks in parallel, and with an average accepted length of 5.6 tokens it cuts decode latency 4.7× over the unoptimized baseline. The flow-matching velocity field is U-shaped across denoising steps, sharp at the endpoints and flat in the middle, so four of eight intermediate velocity evaluations can be cached and reused. Together with CUDA Graph compilation, kernel fusion, and W4A8 quantization, these tricks reduce end-to-end latency from 716.9 ms to 151.4 ms on an RTX PRO 6000 while minADE6@6.4s degrades only from 0.767 m to 0.844 m and minADE1 improves from 1.705 m to 1.573 m.

Load-bearing premise

The near-lossless adaptive-step flow-matching result rests on the velocity profile measured from only 10 clips with 20 window inputs each; if that U-shape is not representative of deployment scenes, skipping the same four middle denoising steps for every input could produce trajectory error well beyond the reported 0.04 m minADE6 degradation.

Editorial extensions

If this is right

  • A single RTX PRO 6000 can run a 10B-parameter reasoning VLA at 6.6 Hz instead of 1.4 Hz, bringing end-to-end driving within the replanning rates used in urban driving.
  • The 4.7× speedup carries to other GPUs: 4.0× on Jetson Thor with one trajectory sample, and 9.6×–10.6× on edge through workstation GPUs with six trajectory samples.
  • W4A8 quantization is the appropriate compression regime for VLA inference because 8-bit activations accelerate the compute-bound prefill, whereas weight-only W4A16 quantization would leave prefill untouched.
  • Streaming fine-tuning of only the action expert, not the VLM, recovers the trajectory accuracy lost to KV-cache approximation and appears to act as a regularizer that improves minADE1.
  • In closed-loop AlpaSim evaluation, collision and off-road rates drop (0.19→0.15 and 0.41→0.32) while relative progress is unchanged, indicating that the latency reduction does not degrade safety.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the four speedups are roughly multiplicative, the same profile-each-stage-and-exploit-its-redundancy recipe should transfer to other cascade pipelines, such as long-horizon video-language agents, with the largest remaining gain concentrated where autoregressive decode dominates.
  • The fixed skip schedule in adaptive-step flow matching is an implementation choice rather than a learned policy; a scene-conditioned or confidence-based selection of which denoising steps to cache could preserve more accuracy in edge cases, but that variant is not tested in the paper.
  • If a future VLA uses free-form rather than templated reasoning tokens, the low entropy that makes DFlash effective would weaken, so the decode speedup would need to be re-derived for that distribution rather than assumed.
  • The streaming KV-cache approximation is only shown to be recoverable within the rollout lengths used in training; longer continuous drives may reveal where the action-expert regularizer saturates.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes FlashDrive, an algorithm-system co-design framework that accelerates vision-language-action (VLA) inference for autonomous driving by attacking four pipeline stages simultaneously. On Alpamayo 1.5-10B, the authors combine streaming KV-cache reuse (§3.1), DFlash diffusion-based speculative decoding (§3.2), adaptive-step flow matching (§3.3), CUDA Graph compilation and kernel fusion (§3.5), and W4A8 quantization (§3.4). They report a 4.7x end-to-end latency reduction from 716.9 ms to 151.4 ms on an RTX PRO 6000, consistent speedups on four other GPUs, and approximately unchanged open-loop minADE6 (0.767 to 0.844 m) with improved minADE1 (1.705 to 1.573 m). Closed-loop results in AlpaSim show improved collision and off-road rates but regressions in Dtraj and Wrong Lane. The paper also reports a 2.5x per-step rollout speedup in closed-loop simulation. The central claim is that the speedup is achieved with negligible accuracy loss, making real-time VLA driving feasible.

Significance. If the central claim holds, FlashDrive is a valuable systems contribution. The latency measurements are direct and the per-stage ablations in Table 1 add consistently to the end-to-end total, which supports the claim that the four bottlenecks compound. The code and pre-trained checkpoints are promised, and the evaluation covers multiple hardware classes including edge devices. The main caveat is that the 'near-lossless' accuracy claim currently rests on single-run point estimates without confidence intervals, and the adaptive-step flow matching component is justified by velocity profiles from only 10 clips. These issues are fixable with additional analysis and do not undermine the speedup measurements themselves.

major comments (3)
  1. [§3.3, Fig. 5, Table 1] The adaptive-step flow-matching policy skips the same four middle denoising steps for every input, but the velocity-profile evidence (Fig. 5) is from only 10 clips with 20 window inputs per clip. The paper reports no variation of the U-shaped profile across clips, and Table 1 reports only aggregate minADE6, which is a minimum over six trajectories and can therefore hide large per-sample errors. The closed-loop regressions in Table 3 (Dtraj 20.0→22.4 m, Wrong Lane 0.45→0.51) are consistent with a subset of scenes degrading more than the open-loop mean suggests. Without per-clip accuracy breakdowns or a robustness analysis of the fixed skip, the 'near-lossless' claim for this component is not established.
  2. [§4.2, Table 1] All efficiency and accuracy numbers in Table 1 are single-run point estimates. The accuracy differences central to the paper (minADE6 +0.077 m, minADE1 −0.132 m) are small enough that they could lie within run-to-run or seed-to-seed variation, yet no error bars, multiple seeds, or confidence intervals are reported. The paper should provide mean±std over at least three seeds for the final comparison (and ideally for the ablations), with bootstrapped confidence intervals over the 12k evaluation windows, before 'essentially unchanged' can be substantiated.
  3. [§4.4, Table 3] The closed-loop safety interpretation is only partially supported by the metrics. While Collision and Off Road improve, Wrong Lane increases (0.45→0.51) and Dtraj increases (20.0→22.4 m). The text argues that Wrong Lane is noisy near intersections, but no episode-level statistics or significance tests are provided. To support the conclusion that 'acceleration does not compromise driving safety,' the paper should report the distribution of event rates across episodes and, if possible, a statistical test for the differences.
minor comments (5)
  1. [Table 1] The token throughput (tokens/s) is embedded in the Decode column; consider a separate subcolumn or annotation to avoid ambiguity.
  2. [§A.1, Table A1a] The row label 'w/ f.t. AE' appears misaligned; check the formatting.
  3. [§2] The related work discusses KV caching for streaming text generation (Xiao et al., 2024), but no comparison with recent token-caching methods for VLA (e.g., VLA-Cache, Xu et al., 2025a) is given; a brief quantitative or qualitative comparison would clarify the novelty of the streaming scheme.
  4. [Table 3] The closed-loop latency in Table 3 includes simulator rendering and trajectory optimization, so it is not directly comparable to the Table 1 inference latency; please state this explicitly in the table caption.
  5. [§3.2] The draft block size of 8 yields an average accepted length of 5.6 tokens, but the acceptance rate and verification cost are not reported; reporting these would help assess the DFlash contribution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: FlashDrive's speedup and near-lossless accuracy claims are direct measurements, not derivations from fitted inputs.

full rationale

The paper's central claims are empirical measurements: end-to-end latency drops from 716.9 ms to 151.4 ms, minADE6@6.4s changes from 0.767 m to 0.844 m, and minADE1@6.4s improves from 1.705 m to 1.573 m on the externally defined Alpamayo 1.5-10B model and NVIDIA AV dataset. No claimed result is equivalent to its inputs by construction. The adaptive-step flow-matching skip pattern is chosen after inspecting the U-shaped velocity profile on 10 clips / 200 windows (Sec. 3.3, Fig. 5), but the near-lossless accuracy figure is a separate measurement on 100 clips / 12k windows (Sec. 4.1, Table 1), not an algebraic consequence of the profile; at most this is small-sample model selection, a statistical limitation rather than circularity. DFlash (Chen et al., 2026) and ParoQuant (Liang et al., 2026) are prior works with overlapping authors, but they are not cited as evidence for the speedup: the paper trains its own two-layer drafter and reports a measured accepted length of 5.6 tokens and 58.2 ms decode, and it applies ParoQuant and measures the resulting 24.6 ms reduction. These citations are therefore not load-bearing for the central result. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via citation, and no known result is renamed as a derivation. The weakest point, the representativeness of the 10-clip velocity profile, is a generalization risk, not a circular step.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

No new physical entities or latent constructs are introduced. DFlash and ParoQuant are prior methods by overlapping authors. The main free choices are hyperparameters that determine the speedup/accuracy trade-off, not fitted physical constants.

free parameters (3)
  • draft block size = 8
    Chosen by hand; ablation shows B=16 gives similar latency, so the exact value affects the decode speedup and accuracy trade-off.
  • cached flow-matching steps = 4 of 8
    Chosen by hand from the measured U-shaped velocity profile; skipping more or fewer steps changes the speedup/accuracy trade-off.
  • fused hidden-state count for drafter = 8
    The last eight target-model tokens are fused into the draft layer KV cache; a design choice to reduce verification cost.
assumptions (3)
  • domain assumption Driving-domain reasoning tokens are short, template-like, low-entropy, and intra-block correlated.
    Needed for the diffusion drafter's high acceptance rate of 5.6 tokens (§3.2). Empirically asserted from Alpamayo's chain-of-causation format.
  • domain assumption The flow-matching velocity field changes sharply at the endpoints and is nearly constant in the middle.
    Needed for adaptive step caching (§3.3). Measured on only 10 clips with 20 windows each, and assumed to hold across deployment.
  • domain assumption Streaming KV-cache approximation error mainly affects the action expert and can be compensated by fine-tuning only the action expert.
    Needed for streaming inference accuracy recovery (§3.1, Fig. 3b). Supported by ablations but only on the NVIDIA dataset.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving." pith.science (2026). https://pith.science/paper/D3MLEUGM

@misc{pith2026260812932,
  author       = {Pith},
  title        = {Pith review of: FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D3MLEUGM}},
  note         = {Machine review of arXiv:2608.12932}
}
read the original abstract

Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.

Figures

Figures reproduced from arXiv: 2608.12932 by the authors.

Figure 1
Figure 1. Reasoning VLA models for autonomous driving, such as Alpamayo 1.5, exhibit prohibitive [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Streaming inference encodes only the newest frame and reuses the KV cache from [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. (a) The streaming attention mask preserves causal attention across views while admitting [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Speculative reasoning with DFlash. A dif [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Consecutive velocities in the flow-matching process are highly redundant through the [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 54 canonical work pages

  1. [1]

    Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xi...

  2. [3]

    DFlash: Block Diffusion for Flash Speculative Decoding

    Jian Chen, Yesheng Liang, and Zhijian Liu. DFlash: Block Diffusion for Flash Speculative Decoding . In International Conference on Machine Learning (ICML), 2026

  3. [4]

    DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving

    Erfei Cui, Wenhai Wang, Zhiqi Li, Jiangwei Xie, Haoming Zou, Hanming Deng, Gen Luo, Lewei Lu, Xizhou Zhu, and Jifeng Dai. DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving . Visual Intelligence, 2025

  4. [6]

    ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

    Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Chong Zhang, Dingyuan Zhang, Hongwei Xie, Bing Wang, and Xiang Bai. ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation . In IEEE/CVF International Conference on Computer Vision (ICCV), 2025

  5. [7]

    EMMA: End-to-End Multimodal Model for Autonomous Driving

    Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, Yin Zhou, James Guo, Dragomir Anguelov, and Mingxing Tan. EMMA: End-to-End Multimodal Model for Autonomous Driving . Transactions on Machine Learning Research (TMLR), 2025

  6. [8]

    OpenVLA: An Open-Source Vision-Language-Action Model

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: An Open-Source Vision-Language-Action Model . In Conference on Robot Learning ...

  7. [9]

    Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

    Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success . In Robotics: Science and Systems (RSS), 2025

  8. [10]

    Ryoo, Yao Lu, Frederick Liu, Vikas Sindhwani, Quan Ho Vuong, Tam \'a s Sarl \'o s, Kenneth Oslund, Karol Hausman, and Kanishka Rao

    Isabel Leal, Krzysztof Choromanski, Deepali Jain, Kumar Avinava Dubey, Jake Varley, Michael S. Ryoo, Yao Lu, Frederick Liu, Vikas Sindhwani, Quan Ho Vuong, Tam \'a s Sarl \'o s, Kenneth Oslund, Karol Hausman, and Kanishka Rao. SARA-RT: Scaling Up Robotics Transformers with Self-Adaptive Robust Attention . In IEEE International Conference on Robotics and A...

Show all 65 references
  1. [11]

    ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

    Yesheng Liang, Haisheng Chen, Zihan Zhang, Song Han, and Zhijian Liu. ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference . In International Conference on Learning Representations (ICLR), 2026

  2. [12]

    AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration . In Conference on Machine Learning and Systems (MLSys), 2024

  3. [13]

    RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

    Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An, Xiaoqi Li, Kaichen Zhou, Senqiao Yang, Renrui Zhang, Yandong Guo, and Shanghang Zhang. RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation . In Advances in Neural Information Processing Sys...

  4. [14]

    PhysicalAI Autonomous Vehicles Dataset , 2025

    NVIDIA . PhysicalAI Autonomous Vehicles Dataset , 2025. URL https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles

  5. [15]

    Alpasim: A modular, lightweight, and data-driven research simulator for autonomous driving, October 2025

    NVIDIA, Yulong Cao, Riccardo de Lutio, Sanja Fidler, Guillermo Garcia Cobo, Zan Gojcic, Maximilian Igl, Boris Ivanovic, Peter Karkus, Janick Martinez Esturo, Marco Pavone, Aaron Smith, Ellie Tanimura, Michal Tyszkiewicz, Michael Watson, Qi Wu, and Le Zhang. Alpasim: A modular,...

  6. [16]

    VLP: Vision Language Planning for Autonomous Driving

    Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik, Alessandro G Allievi, Senem Velipasalar, and Liu Ren. VLP: Vision Language Planning for Autonomous Driving . In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  7. [17]

    FAST: Efficient Action Tokenization for Vision-Language-Action Models

    Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. FAST: Efficient Action Tokenization for Vision-Language-Action Models . In Robotics: Science and Systems (RSS), 2025

  8. [18]

    Waslander, Yu Liu, and Hongsheng Li

    Hao Shao, Yuxuan Hu, Letian Wang, Steven L. Waslander, Yu Liu, and Hongsheng Li. LMDrive: Closed-Loop End-to-End Driving with Large Language Models . In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  9. [19]

    DriveLM: Driving with Graph Visual Question Answering

    Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens Bei wenger, Ping Luo, Andreas Geiger, and Hongyang Li. DriveLM: Driving with Graph Visual Question Answering . In European Conference on Computer Vision (ECCV), 2024

  10. [20]

    GeRM: A Generalist Robotic Model with Mixture-of-Experts for Quadruped Robot

    Wenxuan Song, Han Zhao, Pengxiang Ding, Can Cui, Shangke Lyu, Yaning Fan, and Donglin Wang. GeRM: A Generalist Robotic Model with Mixture-of-Experts for Quadruped Robot . In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024

  11. [21]

    PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding

    Wenxuan Song, Jiayi Chen, Pengxiang Ding, Han Zhao, Wei Zhao, Zhide Zhong, Zongyuan Ge, Zhijun Li, Donglin Wang, Jun Ma, Lujia Wang, and Haoang Li. PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding . arXiv preprint arXiv:25...

  12. [22]

    Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

    Xudong Tan, Yaoxin Yang, Peng Ye, Jialin Zheng, Bizhe Bai, Xinyi Wang, Jia Hao, and Tao Chen. Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models . arXiv preprint arXiv:2505.21200, 2025

  13. [23]

    DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

    Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, Xianpeng Lang, and Hang Zhao. DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models . In Conference on Robot Learning (CoRL), 2024

  14. [24]

    BitVLA: 1-Bit Vision-Language-Action Models for Robotics Manipulation

    Hongyu Wang, Chuyan Xiong, Ruiping Wang, and Xilin Chen. BitVLA: 1-Bit Vision-Language-Action Models for Robotics Manipulation . arXiv preprint arXiv:2506.07530, 2025 a

  15. [25]

    Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail

    Yan Wang, Wenjie Luo, Junjie Bai, Yulong Cao, Tong Che, Ke Chen, Yuxiao Chen, Jenna Diamond, Yifan Ding, Wenhao Ding, Liang Feng, Greg Heinrich, Jack Huang, Peter Karkus, Boyi Li, Pinyi Li, Tsung-Yi Lin, Dongran Liu, Ming-Yu Liu, Langechuan Liu, Zhijian Liu, Jason Lu, Yunxiang...

  16. [26]

    TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

    Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, Yaxin Peng, Feifei Feng, and Jian Tang. TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation . IEEE Robotics and Automation Letters, 2025

  17. [27]

    Efficient Streaming Language Models with Attention Sinks

    Guangxuan Xiao, Yuandong Tang, Juntao Zuo, Junxian Guo, Shang Yang, Haotian Tang, Jinlong Fu, and Song Han. Efficient Streaming Language Models with Attention Sinks . In International Conference on Learning Representations (ICLR), 2024

  18. [28]

    VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching

    Siyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang, and Chang Xu. VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching . In Advances in Neural Information Processing Systems (NeurIPS), 2025 a

  19. [29]

    KV-Efficient VLA: A Method to Speed Up Vision Language Models with RNN-Gated Chunked KV Cache

    Wanshun Xu, Long Zhuang, and Lianlei Shan. KV-Efficient VLA: A Method to Speed Up Vision Language Models with RNN-Gated Chunked KV Cache . arXiv preprint arXiv:2509.21354, 2025 b

  20. [30]

    Wong, Zhenguo Li, and Hengshuang Zhao

    Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee K. Wong, Zhenguo Li, and Hengshuang Zhao. DriveGPT4: Interpretable End-to-End Autonomous Driving via Large Language Model . IEEE Robotics and Automation Letters, 2024

  21. [31]

    DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution

    Yang Yue, Yulin Wang, Bingyi Kang, Yizeng Han, Shenzhi Wang, Shiji Song, Jiashi Feng, and Gao Huang. DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution . In Advances in Neural Information Processing Systems (NeurIPS), 2024

  22. [32]

    HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

    Jianke Zhang, Yanjiang Guo, Xiaoyu Chen, Yen-Jen Wang, Yucheng Hu, Chengming Shi, and Jianyu Chen. HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers . In Conference on Robot Learning (CoRL), 2024

  23. [33]

    MoLe-VLA: Dynamic Layer-Skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

    Rongyu Zhang, Menghang Dong, Yuan Zhang, Liang Heng, Xiaowei Chi, Gaole Dai, Li Du, Yuan Du, and Shanghang Zhang. MoLe-VLA: Dynamic Layer-Skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation . In AAAI Conference on Artificial Intelligenc...

  24. [34]

    Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma

    Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma. AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning . In Advances in Neural Information Processing Systems (Neur...

  25. [35]

    Xiao, Guangxuan and Tang, Yuandong and Zuo, Juntao and Guo, Junxian and Yang, Shang and Tang, Haotian and Fu, Jinlong and Han, Song , booktitle =

  26. [36]

    Chen, Jian and Liang, Yesheng and Liu, Zhijian , booktitle =

  27. [37]

    Wang, Yan and Luo, Wenjie and Bai, Junjie and Cao, Yulong and Che, Tong and Chen, Ke and Chen, Yuxiao and Diamond, Jenna and Ding, Yifan and Ding, Wenhao and Feng, Liang and Heinrich, Greg and Huang, Jack and Karkus, Peter and Li, Boyi and Li, Pinyi and Lin, Tsung-Yi and Liu, ...

  28. [38]

    Lin, Ji and Tang, Jiaming and Tang, Haotian and Yang, Shang and Chen, Wei-Ming and Wang, Wei-Chen and Xiao, Guangxuan and Dang, Xingyu and Gan, Chuang and Han, Song , booktitle =

  29. [39]

    Cui, Erfei and Wang, Wenhai and Li, Zhiqi and Xie, Jiangwei and Zou, Haoming and Deng, Hanming and Luo, Gen and Lu, Lewei and Zhu, Xizhou and Dai, Jifeng , journal =

  30. [40]

    European Conference on Computer Vision (ECCV) , year =

    Sima, Chonghao and Renz, Katrin and Chitta, Kashyap and Chen, Li and Zhang, Hanxue and Xie, Chengen and Bei. European Conference on Computer Vision (ECCV) , year =

  31. [41]

    Hwang, Jyh-Jing and Xu, Runsheng and Lin, Hubert and Hung, Wei-Chih and Ji, Jingwei and Choi, Kristy and Huang, Di and He, Tong and Covington, Paul and Sapp, Benjamin and Zhou, Yin and Guo, James and Anguelov, Dragomir and Tan, Mingxing , journal =

  32. [42]

    and Liu, Yu and Li, Hongsheng , booktitle =

    Shao, Hao and Hu, Yuxuan and Wang, Letian and Waslander, Steven L. and Liu, Yu and Li, Hongsheng , booktitle =

  33. [43]

    and Li, Zhenguo and Zhao, Hengshuang , journal =

    Xu, Zhenhua and Zhang, Yujia and Xie, Enze and Zhao, Zhen and Guo, Yong and Wong, Kwan-Yee K. and Li, Zhenguo and Zhao, Hengshuang , journal =

  34. [44]

    Fu, Haoyu and Zhang, Diankun and Zhao, Zongchuang and Cui, Jianfeng and Liang, Dingkang and Zhang, Chong and Zhang, Dingyuan and Xie, Hongwei and Wang, Bing and Bai, Xiang , booktitle =

  35. [45]

    and Zhang, Yun and Huang, Zhiyu and Zhou, Bolei and Ma, Jiaqi , booktitle =

    Zhou, Zewei and Cai, Tianhui and Zhao, Seth Z. and Zhang, Yun and Huang, Zhiyu and Zhou, Bolei and Ma, Jiaqi , booktitle =

  36. [46]

    Tian, Xiaoyu and Gu, Junru and Li, Bailin and Liu, Yicheng and Wang, Yang and Zhao, Zhiyong and Zhan, Kun and Jia, Peng and Lang, Xianpeng and Zhao, Hang , booktitle =

  37. [47]

    Pan, Chenbin and Yaman, Burhaneddin and Nesti, Tommaso and Mallik, Abhirup and Allievi, Alessandro G and Velipasalar, Senem and Ren, Liu , booktitle =

  38. [48]

    and Lu, Yao and Liu, Frederick and Sindhwani, Vikas and Vuong, Quan Ho and Sarl

    Leal, Isabel and Choromanski, Krzysztof and Jain, Deepali and Dubey, Kumar Avinava and Varley, Jake and Ryoo, Michael S. and Lu, Yao and Liu, Frederick and Sindhwani, Vikas and Vuong, Quan Ho and Sarl. IEEE International Conference on Robotics and Automation (ICRA) , year =

  39. [49]

    Xu, Wanshun and Zhuang, Long and Shan, Lianlei , journal =

  40. [50]

    Liu, Jiaming and Liu, Mengzhen and Wang, Zhenyu and An, Pengju and Li, Xiaoqi and Zhou, Kaichen and Yang, Senqiao and Zhang, Renrui and Guo, Yandong and Zhang, Shanghang , booktitle =

  41. [51]

    Kim, Moo Jin and Finn, Chelsea and Liang, Percy , booktitle =

  42. [52]

    Wen, Junjie and Zhu, Yichen and Li, Jinming and Zhu, Minjie and Wu, Kun and Xu, Zhiyuan and Liu, Ning and Cheng, Ran and Shen, Chaomin and Peng, Yaxin and Feng, Feifei and Tang, Jian , journal =

  43. [53]

    Song, Wenxuan and Chen, Jiayi and Ding, Pengxiang and Zhao, Han and Zhao, Wei and Zhong, Zhide and Ge, Zongyuan and Li, Zhijun and Wang, Donglin and Ma, Jun and Wang, Lujia and Li, Haoang , journal =

  44. [54]

    arXiv preprint arXiv:2507.14049 , year =

    Budzianowski, Pawe. arXiv preprint arXiv:2507.14049 , year =

  45. [55]

    Song, Wenxuan and Zhao, Han and Ding, Pengxiang and Cui, Can and Lyu, Shangke and Fan, Yaning and Wang, Donglin , booktitle =

  46. [56]

    Zhang, Jianke and Guo, Yanjiang and Chen, Xiaoyu and Wang, Yen-Jen and Hu, Yucheng and Shi, Chengming and Chen, Jianyu , booktitle =

  47. [57]

    Zhang, Rongyu and Dong, Menghang and Zhang, Yuan and Heng, Liang and Chi, Xiaowei and Dai, Gaole and Du, Li and Du, Yuan and Zhang, Shanghang , booktitle =

  48. [58]

    Yue, Yang and Wang, Yulin and Kang, Bingyi and Han, Yizeng and Wang, Shenzhi and Song, Shiji and Feng, Jiashi and Huang, Gao , booktitle =

  49. [59]

    Kim, Moo Jin and Pertsch, Karl and Karamcheti, Siddharth and Xiao, Ted and Balakrishna, Ashwin and Nair, Suraj and Rafailov, Rafael and Foster, Ethan and Lam, Grace and Sanketi, Pannag and Vuong, Quan and Kollar, Thomas and Burchfiel, Benjamin and Tedrake, Russ and Sadigh, Dor...

  50. [60]

    Wang, Hongyu and Xiong, Chuyan and Wang, Ruiping and Chen, Xilin , journal =

  51. [61]

    Pertsch, Karl and Stachowicz, Kyle and Ichter, Brian and Driess, Danny and Nair, Suraj and Vuong, Quan and Mees, Oier and Finn, Chelsea and Levine, Sergey , booktitle =

  52. [62]

    Tan, Xudong and Yang, Yaoxin and Ye, Peng and Zheng, Jialin and Bai, Bizhe and Wang, Xinyi and Hao, Jia and Chen, Tao , journal =

  53. [63]

    Xu, Siyu and Wang, Yunke and Xia, Chenghao and Zhu, Dihao and Huang, Tao and Xu, Chang , booktitle =

  54. [64]

    Black, Kevin and Brown, Noah and Darpinian, James and Dhabalia, Karan and Driess, Danny and Esmail, Adnan and Equi, Michael and Finn, Chelsea and Fusai, Niccolo and Galliker, Manuel Y. and Ghosh, Dibya and Groom, Lachy and Hausman, Karol and Ichter, Brian and Jakubczak, Szymon...

  55. [65]

    Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian , booktitle =

  56. [66]

    arXiv preprint arXiv:2408.11743 , year=

    MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models , author=. arXiv preprint arXiv:2408.11743 , year=

  57. [67]

    2025 , month =

    NVIDIA and Yulong Cao and Riccardo de Lutio and Sanja Fidler and Guillermo Garcia Cobo and Zan Gojcic and Maximilian Igl and Boris Ivanovic and Peter Karkus and Janick Martinez Esturo and Marco Pavone and Aaron Smith and Ellie Tanimura and Michal Tyszkiewicz and Michael Watson...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.