A 337M action expert that reads a frozen 7B VLM's per-layer KV cache every tick, trained with randomized staleness, lifts CARLA route completion from 37% to 94% at full 20 Hz control rate.
ReasonNet: End-to-End Driving with Temporal and Global Reasoning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The large-scale deployment of autonomous vehicles is yet to come, and one of the major remaining challenges lies in urban dense traffic scenarios. In such cases, it remains challenging to predict the future evolution of the scene and future behaviors of objects, and to deal with rare adverse events such as the sudden appearance of occluded objects. In this paper, we present ReasonNet, a novel end-to-end driving framework that extensively exploits both temporal and global information of the driving scene. By reasoning on the temporal behavior of objects, our method can effectively process the interactions and relationships among features in different frames. Reasoning about the global information of the scene can also improve overall perception performance and benefit the detection of adverse events, especially the anticipation of potential danger from occluded objects. For comprehensive evaluation on occlusion events, we also release publicly a driving simulation benchmark DriveOcclusionSim consisting of diverse occlusion events. We conduct extensive experiments on multiple CARLA benchmarks, where our model outperforms all prior methods, ranking first on the sensor track of the public CARLA Leaderboard.
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving
A 337M action expert that reads a frozen 7B VLM's per-layer KV cache every tick, trained with randomized staleness, lifts CARLA route completion from 37% to 94% at full 20 Hz control rate.