A hybrid planner that adds a vision-language model reading raw multi-view images to a real-time trajectory planner achieves modest but consistent gains on nuPlan, with an adaptive gate that cuts VLM inference frequency.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VLMPlanner: Integrating Visual Language Models with Motion Planning
A hybrid planner that adds a vision-language model reading raw multi-view images to a real-time trajectory planner achieves modest but consistent gains on nuPlan, with an adaptive gate that cuts VLM inference frequency.