A Survey on Multimodal Large Language Models for Autonomous Driving

Ao Liu; Can Cui; Chao Zheng; Erlong Li; Jianguo Cao; Jintai Chen; Juanwu Lu; Kaizhao Liang; Kuei-Da Liao; Kun Tang

arxiv: 2311.12320 · v1 · pith:RQW6UXLInew · submitted 2023-11-21 · 💻 cs.AI

A Survey on Multimodal Large Language Models for Autonomous Driving

Can Cui , Yunsheng Ma , Xu Cao , Wenqian Ye , Yang Zhou , Kaizhao Liang , Jintai Chen , Juanwu Lu

show 13 more authors

Zichong Yang Kuei-Da Liao Tianren Gao Erlong Li Kun Tang Zhipeng Cao Tong Zhou Ao Liu Xinrui Yan Shuqi Mei Jianguo Cao Ziran Wang Chao Zheng

This is my paper

classification 💻 cs.AI

keywords drivingmodelsautonomouslargesystemslanguagellmsmultimodal

0 comments

read the original abstract

With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and control tools as humans. In recent months, LLMs have shown widespread attention in autonomous driving and map systems. Despite its immense potential, there is still a lack of a comprehensive understanding of key challenges, opportunities, and future endeavors to apply in LLM driving systems. In this paper, we present a systematic investigation in this field. We first introduce the background of Multimodal Large Language Models (MLLMs), the multimodal models development using LLMs, and the history of autonomous driving. Then, we overview existing MLLM tools for driving, transportation, and map systems together with existing datasets and benchmarks. Moreover, we summarized the works in The 1st WACV Workshop on Large Language and Vision Models for Autonomous Driving (LLVM-AD), which is the first workshop of its kind regarding LLMs in autonomous driving. To further promote the development of this field, we also discuss several important problems regarding using MLLMs in autonomous driving systems that need to be solved by both academia and industry.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
cs.CV 2026-05 unverdicted novelty 7.0

DriveSpatial benchmark shows the best of 15 VLMs trails humans by 28.4 points on spatiotemporal driving tasks, with cognitive scene construction as the main failure mode.
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
cs.CV 2026-05 unverdicted novelty 7.0

DriveSpatial benchmark shows the strongest of 15 VLMs trails humans by 28.4 points on spatiotemporal tasks, with cognitive scene construction as the primary weakness.
LiloDriver: A Lifelong Learning Framework for Closed-loop Motion Planning in Long-tail Autonomous Driving Scenarios
cs.RO 2025-05 unverdicted novelty 6.0

LiloDriver uses LLMs and memory-augmented planning in a four-stage pipeline to outperform rule-based and learning-based methods on both common and rare scenarios in the nuPlan benchmark.
Large Language Models in Transportation Systems Management and Operations: From Text Reasoning to Multi-modal Decision Support
cs.AI 2026-05 unverdicted novelty 2.0

A survey synthesizing LLM and MM-LLM uses in transportation operations, mobility services, and decision support while noting challenges like data heterogeneity and real-time needs.