A co-design of streaming KV-cache reuse, diffusion-based speculative decoding, adaptive flow-matching step caching, and CUDA Graph/kernel fusion cuts VLA autonomous-driving inference latency 4.7x with roughly unchanged open-loop trajectory error.
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
A co-design of streaming KV-cache reuse, diffusion-based speculative decoding, adaptive flow-matching step caching, and CUDA Graph/kernel fusion cuts VLA autonomous-driving inference latency 4.7x with roughly unchanged open-loop trajectory error.