Pith. sign in

hub Canonical reference

Vision-language-action models for autonomous driving: Past, present, and future

Canonical reference. 100% of citing Pith papers cite this work as background.

18 Pith papers citing it
Background 100% of classified citations

hub tools

citation-role summary

background 7

citation-polarity summary

years

2026 18

roles

background 6

polarities

background 6

representative citing papers

Grounding Driving VLA via Inverse Kinematics

cs.CV · 2026-05-20 · conditional · novelty 7.0

By adding future visual state prediction and a dedicated inverse kinematics diffusion network that uses only visual boundary conditions, a 0.5B driving VLA recovers visual grounding and matches 7-8B models on NAVSIM-v2 and nuScenes.

EventDrive: Event Cameras for Vision-Language Driving Intelligence

cs.CV · 2026-06-16 · unverdicted · novelty 6.0

EventDrive supplies a multi-task benchmark and EventDrive-VLM architecture that fuses event data, RGB, and language supervision, reporting gains in temporal precision and motion awareness for driving intelligence.

Post-Training in End-to-End Autonomous Driving

cs.CV · 2026-07-09 · unverdicted · novelty 4.0

Post-training for end-to-end autonomous driving is surveyed and grouped into four supervision-based families to address limits of open-loop imitation.

citing papers explorer

Showing 18 of 18 citing papers.