REVIEW 3 major objections 4 minor 3 cited by
STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Driving vision-language models fail at spatio-temporal scene understanding, a new 971-question benchmark argues.
desk verdict A well-built benchmark that fills a real gap, but the headline claim that driving VLMs lack spatio-temporal reasoning outruns the confounded evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the STSBench scenario catalog: 43 textual scenario definitions (e.g., lane change, overtake, wait for pedestrian to cross), each paired with negative scenarios that do not occur, so that every question has distractors requiring discrimination of close alternatives. Mining heuristics use ground-truth 3D bounding boxes, tracks, class labels, ego-motion, and HD maps to detect these patterns automatically; a lightweight human verification interface lets drivers confirm positives and reject false negatives; verified samples are then turned into multiple-choice questions of the form 'which of the following best describes ...' with one correct answer among at least four choices. The evaluation protocol compares three model families under adapted prompts: LLMs receive ground-truth trajectories, off-the-shelf VLMs receive single-view image sequences with camera metadata, and driving-expert VLMs receive full multi-view video, making the comparison hinge on how each model integrates spatial and temporal evidence.
What would settle it
Give the best driving-expert VLM perfect perception (for example, replace its visual input with ground-truth object tracks and rendered bounding boxes while keeping the language model frozen) and re-run STSnu; if its agent-to-agent accuracy jumps to the level of the trajectory-fed LLM, the paper's conclusion that the model lacks spatio-temporal reasoning would collapse, because the failure would trace to the vision-to-language interface rather than to reasoning.
Extended reading notes
Core claim
The paper's central claim is that spatio-temporal reasoning is the capability currently missing in driving vision-language models. Using STSnu, the paper shows that when an LLM is given perfect trajectories it can identify ego maneuvers and interactions at 57.08% average accuracy (GPT-4o), while the best driving-expert VLM (DriveMM) reaches only 39.51%; agent-to-agent interactions, where neither participant is the ego vehicle, are hardest for all models. The paper interprets this gap as evidence that visual models have not learned to jointly reason over spatially distributed multi-view inputs and temporally extended dynamics. It therefore argues that end-to-end driving models need architectural mechanisms that explicitly model spatio-temporal relationships, not just better perception heads or larger training sets.
Load-bearing premise
The evaluation assumes that the accuracy gap between LLMs fed ground-truth trajectories and VLMs fed raw images is caused by differences in spatio-temporal reasoning ability rather than by perception quality, prompt formatting, or model scale.
Editorial extensions
If this is right
- Any claim that an end-to-end driving VLM understands a scene should be backed by interaction-level questions like those in STSnu; waypoint or ego-action accuracy alone is insufficient.
- Training data and objectives for driving VLMs should explicitly include third-party interactions, since agent-to-agent scenarios are where all evaluated models drop hardest.
- Injecting perception-derived trajectories or 3D object states into the language model may be a more direct route to spatio-temporal reasoning than asking the vision encoder to infer them from raw pixels.
- STSBench can be re-instantiated on other datasets with ground-truth annotations, producing comparable interaction benchmarks across different sensor setups without per-dataset manual annotation.
- Benchmark distractors should remain semantically close (e.g., overtake versus pass) because the results show models frequently confuse such distinctions.
Reading between the lines
- A plausible extension is to use STSnu as a training signal: fine-tune a driving VLM with a trajectory-infusion module and re-measure; the paper's data predict a large gain if the gap is mainly perceptual rather than reasoning-based.
- The multiple-choice format may reward elimination strategies rather than true identification; a follow-up variant that asks the model to justify its choice or to detect two simultaneous scenarios in one scene would test whether the apparent reasoning is robust.
- Because NuScenes is recorded in Boston and Singapore and contains mostly lawful behavior, applying STSBench to more diverse or adversarial recordings could reveal even larger failures than the 57% ceiling suggests.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces STSBench, a framework for automatically mining defined traffic scenarios (ego, agent, ego-to-agent, and agent-to-agent) from datasets with rich ground-truth annotations, together with a lightweight human-verification interface and automatic multiple-choice question generation. Applied to the NuScenes validation split, the framework yields STSnu, comprising 43 scenario types and 971 human-verified multiple-choice questions. The authors evaluate three LLMs, three off-the-shelf VLMs, and three driving-expert VLMs on STSnu, and report that LLMs given ground-truth trajectories substantially outperform the visual models. They interpret this as evidence that current driving VLMs lack spatio-temporal understanding of dynamic traffic scenes, particularly for third-party agent interactions.
Significance. The framework addresses a genuine evaluation gap: most existing driving-VLM benchmarks focus on single-image or monocular ego-centric tasks, whereas STSnu targets multi-view video, third-party interactions, and holistic spatio-temporal reasoning. The release of code and data, the reporting of inter-reviewer agreement (85.6% positive agreement and 20.8% disagreement on negatives), the analysis of multiple-choice letter distribution, and the ablations of query frame and chain-of-thought for OmniDrive are concrete strengths that make the benchmark resource potentially useful to the community. If the benchmark is adopted, the per-scenario results in Tables 8-11 will be a valuable reference. However, the paper's headline conclusion goes beyond what the current experimental design can establish, because the central comparison between LLMs and VLMs is confounded by differences in the information given to each model class, in prompt formatting, and in model scale.
major comments (3)
- [Section 4, Table 2, and Appendix D, Table 13] The claim that VLMs 'do not have a spatio-temporal understanding of dynamic traffic scenes' is not established by the reported comparison, because the LLM baseline receives ground-truth trajectories (GPS positions, LiDAR coordinates, velocities) while the VLMs receive raw images with per-model adapted prompts; the accuracy gap can therefore be explained by perception quality, prompt compatibility, or model scale rather than by reasoning ability alone. The authors' own control experiment in Table 13, which feeds ground-truth information to three expert VLMs, yields mixed results (DriveMM improves from 39.5 to 48.5, OmniDrive stays flat at 28.4, Senna collapses to 3.2), and it does not cover the off-the-shelf VLMs or the exact LLM prompt format, so it does not resolve the confound. I recommend adding an oracle-perception or text-input condition for the same VLMs (e.g., providing the same trajectory and GPS text used for the LLMs) and/or softening the conclusion to the supported claim that current driving VLMs perform poorly on these questions under their native input formats.
- [Section 4 and Tables 8-11] No confidence intervals or significance tests are reported, despite small per-scenario sample sizes (e.g., Table 8 shows scenario groups with 10-37 questions) and differences between model accuracies that are often small relative to the implied sampling noise (e.g., Table 2: InternVL 2.5 8B at 46.07% vs. DriveMM at 39.51% vs. OmniDrive at 29.33%). The phrase 'outperforms ... by a significant margin' in Section 4 requires a formal test, such as McNemar's test or bootstrap confidence intervals; without this, the model ranking and the 'critical shortcomings' narrative are not quantitatively supported.
- [Section 4 and Appendix A.4/E] The evaluation protocol adapts prompts and input formats per model 'in order to get better performance' (Appendix A.4), which makes cross-model accuracy differences at least partly attributable to how well each prompt matches the model's training distribution rather than to spatio-temporal reasoning. I recommend reporting results with a shared, minimally-adapted prompt as the primary protocol, with the per-model optimized prompts as a secondary analysis, and documenting the variance induced by prompt changes, for example by running each model with two or three variants on a subset of the benchmark.
minor comments (4)
- [Table 9 caption] The caption of Table 9 says 'for ego-to-agent scenarios', but the table lists single-agent scenarios such as jaywalking, walking, standing, crossing, left turn, and overtaking ego; the caption should read 'agent scenarios' to match the main text and the scenario categories.
- [Figures and Appendix E] There are several typos in the figures and appendix prompts, including 'U-Tuen' in Figure 2, 'caputred' in Figures 27-34, 'assistent' and 'specilized' in Figures 35-38, and 'whic is a pedestrian' in Figure 42; these should be corrected for a polished final version.
- [References] Reference [15] (HiLM-D) uses the same arXiv identifier 2308.12966 as reference [4] (Qwen-VL), which appears to be an incorrect duplicate; please verify and replace it with the correct identifier for the HiLM-D technical report.
- [Section 3.3] The verification protocol accepts positive samples by majority voting but keeps only negatives with full agreement across all reviewers; this asymmetry in quality thresholds is not analyzed, and a brief discussion of its potential effect on benchmark difficulty and on the false-negative rate would improve transparency.
Circularity Check
No circularity: STSnu is built from NuScenes ground-truth annotations and human verification, and the model evaluations are independent of the benchmark construction.
full rationale
The paper's derivation chain is: define a scenario catalog, mine scenarios from NuScenes 3D tracks, ego-motion, and HD maps, verify them with human reviewers, generate fixed multiple-choice questions, and then evaluate models on those questions. The central findings (e.g., Table 2 showing that driving expert VLMs score poorly) are measurements on this externally constructed benchmark, not quantities fitted to any evaluated model. The LLM baselines receive ground-truth trajectories while VLMs receive images, so the accuracy gap may reflect perception quality, prompt adaptation, or model scale differences; however, that is an experimental confound relevant to the strength of the conclusion, not circularity, because no model output is used to define the questions or answers. The Appendix D experiment that provides ground-truth information to expert VLMs (Table 13) is an additional probe and shows mixed results, which further indicates the benchmark is not engineered to force a particular outcome. No self-citation is load-bearing, no uniqueness theorem is imported from prior work by the authors, and no ansatz is smuggled in via citation. The optional sub-sampling of scenarios based on occlusion and distance shapes benchmark difficulty, but that is a design choice rather than a circular step under the stated criteria.
Assumptions & free parameters
free parameters (2)
- Scenario mining heuristic thresholds (e.g., lane-crossing detection, speed change thresholds)
- Sub-sampling thresholds for occlusion rate, distance to ego, and spatial distribution
assumptions (3)
- domain assumption NuScenes ground-truth annotations (3D boxes, tracks, ego motion, HD maps) are accurate and complete enough for scenario mining.
- domain assumption Human reviewers can reliably distinguish true scenarios from false positives and negatives after brief training.
- domain assumption Multiple-choice questions with one correct answer from a fixed scenario catalog are a valid test of spatio-temporal reasoning.
Cite this review
Pith. "Pith review of STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving." pith.science (2026). https://pith.science/paper/RQAUSEHM
@misc{pith2026250606218,
author = {Pith},
title = {Pith review of: STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving},
year = {2026},
howpublished = {\url{https://pith.science/paper/RQAUSEHM}},
note = {Machine review of arXiv:2506.06218}
}
read the original abstract
We introduce STSBench, a scenario-based framework to benchmark the holistic understanding of vision-language models (VLMs) for autonomous driving. The framework automatically mines pre-defined traffic scenarios from any dataset using ground-truth annotations, provides an intuitive user interface for efficient human verification, and generates multiple-choice questions for model evaluation. Applied to the NuScenes dataset, we present STSnu, the first benchmark that evaluates the spatio-temporal reasoning capabilities of VLMs based on comprehensive 3D perception. Existing benchmarks typically target off-the-shelf or fine-tuned VLMs for images or videos from a single viewpoint and focus on semantic tasks such as object recognition, dense captioning, risk assessment, or scene understanding. In contrast, STSnu evaluates driving expert VLMs for end-to-end driving, operating on videos from multi-view cameras or LiDAR. It specifically assesses their ability to reason about both ego-vehicle actions and complex interactions among traffic participants, a crucial capability for autonomous vehicles. The benchmark features 43 diverse scenarios spanning multiple views and frames, resulting in 971 human-verified multiple-choice questions. A thorough evaluation uncovers critical shortcomings in existing models' ability to reason about fundamental traffic dynamics in complex environments. These findings highlight the urgent need for architectural advances that explicitly model spatio-temporal reasoning. By addressing a core gap in spatio-temporal evaluation, STSBench enables the development of more robust and explainable VLMs for autonomous driving.
Figures
Figures from the paper (43 more)
Forward citations
Cited by 3 Pith papers
-
STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision
Using automotive radar Doppler as supervision, STAR-VLM enables a vision-language model to estimate metric radial velocity and motion state of objects from video, outperforming zero-shot task-specific baselines on a n...
-
ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
Real adverse-weather camera–LiDAR–radar MCQs expose VLM failures from observability estimation through spatial grounding to trajectory safety, partially mitigated by SFT+RL.
-
CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis
CARLA-GS is a modular pipeline that uses an LLM for semantic trajectory planning, CARLA for physics execution, and 3D Gaussian Splatting for photorealistic rendering to synthesize autonomous driving corner cases.
Reference graph
Works this paper leans on
-
[1]
LLaMA 3.2: Open Foundation and Instruction Models, 2024
Meta AI. LLaMA 3.2: Open Foundation and Instruction Models, 2024. 9
work page 2024
-
[2]
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, R...
work page 2022
-
[3]
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
Hidehisa Arai, Keita Miwa, Kento Sasaki, Yu Yamaguchi, Kohei Watanabe, Shunsuke Aoki, and Issei Yamamoto. CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving. In WACV, 2024
work page 2024
-
[5]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-VL Technical Report. a...
arXiv 2025
-
[6]
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuScenes: A multimodal dataset for autonomous driving. In CVPR, 2020
work page 2020
-
[7]
MAPLM: A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding
Xu Cao, Tong Zhou, Yunsheng Ma, Wenqian Ye, Can Cui, Kun Tang, Zhipeng Cao, Kaizhao Liang, Ziran Wang, James Rehg, and Chao Zheng. MAPLM: A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding. In CVPR, 2024
work page 2024
-
[8]
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Science China Information Sciences , 67(12):220101, 2024
2024
-
[9]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In CVPR, 2024
2024
Show all 76 references
-
[10]
The Cityscapes Dataset for Semantic Urban Scene Understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding. In CVPR, 2016
2016
-
[11]
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. In NeurIPS, 2023
2023
-
[12]
Deepseek-v3 technical report
DeepSeek-AI. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[13]
Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch,...
2024
-
[14]
Talk2Car: Taking control of your self-driving car
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic, Luc Van Gool, and Marie-Francine Moens. Talk2Car: Taking control of your self-driving car. In EMNLP-IJCNLP, 2019. 10
2019
-
[15]
HiLM-D: Towards High- Resolution Understanding in Multimodal Large Language Models for Autonomous Driving
Xinpeng Ding, Jianhua Han, Hang Xu, Wei Zhang, and Xiaomeng Li. HiLM-D: Towards High- Resolution Understanding in Multimodal Large Language Models for Autonomous Driving. arXiv preprint arXiv:2308.12966, 2023
2023 arXiv
-
[16]
Holistic Autonomous Driving Understanding by Bird’s-Eye-View Injected Multi-Modal Large Models
Xinpeng Ding, Jinahua Han, Hang Xu, Xiaodan Laing, Xu Hang, Wei Zhang, and Xiaomeng Li. Holistic Autonomous Driving Understanding by Bird’s-Eye-View Injected Multi-Modal Large Models. In CVPR, 2024
2024
-
[17]
CARLA: An Open Urban Driving Simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. CARLA: An Open Urban Driving Simulator. In CoRL, 2017
2017
-
[18]
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Chong Zhang, Dingyuan Zhang, Hongwei Xie, Bing Wang, and Xiang Bai. ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation. arXiv preprint arXiv:2503.19755, 2025
2025 arXiv
-
[19]
Are we ready for autonomous driving? The KITTI vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? The KITTI vision benchmark suite. In CVPR, 2012
2012
-
[20]
DriveMLLM: A Benchmark for Spatial Understanding with Multimodal Large Language Models in Autonomous Driving
Xianda Guo, Zhang Ruijun, Duan Yiqun, He Yuhang, Chenming Zhang, and Long Chen. DriveMLLM: A Benchmark for Spatial Understanding with Multimodal Large Language Models in Autonomous Driving. arXiv preprint arXiv:2411.13112, 2024
2024 arXiv
-
[21]
Planning-oriented Autonomous Driving
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, Lewei Lu, Xiaosong Jia, Qiang Liu, Jifeng Dai, Yu Qiao, and Hongyang Li. Planning-oriented Autonomous Driving. In CVPR, 2023
2023
-
[22]
Drivemm: All-in-one large multimodal model for autonomous driving
Zhijian Huang, Chengjian Fen, Feng Yan, Baihui Xiao, Zequn Jie, Yujie Zhong, Xiaodan Liang, and Lin Ma. Drivemm: All-in-one large multimodal model for autonomous driving. arXiv preprint arXiv:2412.07689, 2024
2024 arXiv
-
[23]
Making Large Language Models Better Planners with Reasoning-Decision Alignment
Zhijian Huang, Tao Tang, Shaoxiang Chen, Sihao Lin, Zequn Jie, Lin Ma, Guangrun Wang, and Xiaodan Liang. Making Large Language Models Better Planners with Reasoning-Decision Alignment. In ECCV, 2024
2024
-
[24]
Emma: End-to-end multimodal model for autonomous driving
Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, Yin Zhou, James Guo, Dragomir Anguelov, and Mingxing Tan. Emma: End-to-end multimodal model for autonomous driving. arXiv preprint arXiv:2410.23262, 2024
-
[25]
NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets Using Markup Annotations
Yuichi Inoue, Yuki Yada, Kotaro Tanahashi, and Yu Yamaguchi. NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets Using Markup Annotations. In WACVW, 2024
2024
-
[26]
Drivelmm-o1: A step-by-step reasoning dataset and large multimodal model for driving scenario understanding
Ayesha Ishaq, Jean Lahoud, Ketan More, Omkar Thawakar, Ritesh Thawkar, Dinura Dis- sanayake, Noor Ahsan, Yuhao Li, Fahad Shahbaz Khan, Hisham Cholakkal, Ivan Laptev, Rao Muhammad Anwer, and Salman Khan. Drivelmm-o1: A step-by-step reasoning dataset and large multimodal model f...
2025 arXiv
-
[27]
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving
Xiaosong Jia, Zhenjie Yang, Qifeng Li, Zhiyuan Zhang, and Junchi Yan. Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving. In NeurIPS, 2024
2024
-
[28]
V AD: Vectorized Scene Representation for Efficient Autonomous Driving
Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang. V AD: Vectorized Scene Representation for Efficient Autonomous Driving. In ICCV, 2023
2023
-
[29]
Senna: Bridging Large Vision-Language Models and End-to- End Autonomous Driving
Bo Jiang, Shaoyu Chen, Bencheng Liao, Xingyu Zhang, Wei Yin, Qian Zhang, Chang Huang, Wenyu Liu, and Xinggang Wang. Senna: Bridging Large Vision-Language Models and End-to- End Autonomous Driving. arXiv preprint arXiv:2410.22313, 2024. 11
-
[30]
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Bu Jin, Yupeng Zheng, Pengfei Li, Weize Li, Yuhang Zheng, Sujie Hu, Xinyu Liu, Jinwei Zhu, Zhijie Yan, Haiyang Sun, Kun Zhan, Peng Jia, Xiaoxiao Long, Yilun Chen, and Hao Zhao. TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes. In ECCV, 2024
2024
-
[31]
Textual Explana- tions for Self-Driving Vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata. Textual Explana- tions for Self-Driving Vehicles. In ECCV, 2018
2018
-
[32]
Junnan Li, Dongxu Li, Caiming Xiong, and S. Hoi. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. In ICML, 2022
2022
-
[33]
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023
2023
-
[34]
Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving
Tengpeng Li, Hanli Wang, Xianfei Li, Wenlong Liao, Tao He, and Pai Peng. Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving. arXiv preprint arXiv:2501.08861, 2025
2025 arXiv
-
[35]
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
Yanze Li, Wenhua Zhang, Kai Chen, Yanxin Liu, Pengxiang Li, Ruiyuan Gao, Lanqing Hong, Meng Tian, Xinhai Zhao, Zhenguo Li, et al. Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases. In WACV, 2025
2025
-
[36]
BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Yu Qiao, and Jifeng Dai. BEVFormer: Learning Bird’s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers. In ECCV, 2022
2022
-
[37]
Zhiqi Li, Zhiding Yu, Shiyi Lan, Jiahan Li, Jan Kautz, Tong Lu, and Jose M. Alvarez. Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? . In CVPR, 2024
2024
-
[38]
Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. In NeurIPS, 2023
2023
-
[39]
Improved Baselines with Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved Baselines with Visual Instruction Tuning. In CVPR, 2024
2024
-
[40]
Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, 2024
2024
-
[41]
Dolphins: Multimodal Language Model for Driving
Yingzi Ma, Yulong Cao, Jiachen Sun, Marco Pavone, and Chaowei Xiao. Dolphins: Multimodal Language Model for Driving. In ECCV, 2024
2024
-
[42]
DRAMA: Joint Risk Localization and Captioning in Driving
Srikanth Malla, Chiho Choi, Isht Dwivedi, Joon Hee Choi, and Jiachen Li. DRAMA: Joint Risk Localization and Captioning in Driving. In WACV, 2023
2023
-
[43]
One Million Scenes for Autonomous Driving: ONCE Dataset
Jiageng Mao, Minzhe Niu, Chenhan Jiang, Xiaodan Liang, Yamin Li, Chaoqiang Ye, Wei Zhang, Zhenguo Li, Jie Yu, and Chunjing Xu. One Million Scenes for Autonomous Driving: ONCE Dataset. In NeurIPS, 2021
2021
-
[44]
LingoQA: Video Question Answering for Autonomous Driving
Ana-Maria Marcu, Long Chen, Jan Hünermann, Alice Karnsund, Benoit Hanotte, Prajwal Chidananda, Saurabh Nair, Vijay Badrinarayanan, Alex Kendall, Jamie Shotton, and Oleg Sinavski. LingoQA: Video Question Answering for Autonomous Driving. In ECCV, 2024
2024
-
[45]
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
Ming Nie, Renyuan Peng, Chunwei Wang, Xinyue Cai, Jianhua Han, Hang Xu, and Li Zhang. Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving. In ECCV, 2024
2024
-
[46]
GPT-4 Technical Report
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and et al. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774, 2024
2024 arXiv
-
[47]
VLP: Vision Language Planning for Autonomous Driving
Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik, Alessandro G Allievi, Senem Velipasalar, and Liu Ren. VLP: Vision Language Planning for Autonomous Driving. In CVPR, 2024. 12
2024
-
[48]
NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario
Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao, and Yu-Gang Jiang. NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving Scenario. In AAAI, 2024
2024
-
[49]
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. In ICML, 2021
2021
-
[50]
Rerun: A Visualization SDK for Multimodal Data, 2024
Rerun Development Team. Rerun: A Visualization SDK for Multimodal Data, 2024. URL https://www.rerun.io. Available from https://www.rerun.io/ and https://github.com/rerun- io/rerun
2024
-
[51]
Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning
Enna Sachdeva, Nakul Agarwal, Suhas Chundi, Sean Roelofs, Jiachen Li, Mykel Kochender- fer, Chiho Choi, and Behzad Dariush. Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning. In WACV, 2024
2024
-
[52]
Waslander, Yu Liu, and Hong- sheng Li
Hao Shao, Yuxuan Hu, Letian Wang, Guanglu Song, Steven L. Waslander, Yu Liu, and Hong- sheng Li. LMDrive: Closed-Loop End-to-End Driving with Large Language Models. In CVPR, 2024
2024
-
[53]
DriveLM: Driving with Graph Visual Question Answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Ping Luo, Andreas Geiger, and Hongyang Li. DriveLM: Driving with Graph Visual Question Answering. In ECCV, 2024
2024
-
[54]
InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
Ruiqi Song, Xianda Guo, Hangbin Wu, Qinggong Wei, and Long Chen. InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving. arXiv preprint arXiv:2503.13047, 2025
2025
-
[55]
Scalability in Perception for Autonomous Driving: Waymo Open Dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in Perception for Autonomous Driving: Waymo Open Dataset. In CVPR, 2020
2020
-
[56]
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
Kexin Tian, Jingrui Mao, Yunlong Zhang, Jiwan Jiang, Yang Zhou, and Zhengzhong Tu. NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving. arXiv preprint arXiv:2504.03164, 2025
2025 arXiv
-
[57]
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Zhiyong Zhao, Yang Wang, Kun Zhan, Peng Jia, Xianpeng Lang, and Hang Zhao. DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models. In CoRL, 2024
2024
-
[58]
Object Referring in Videos With Language and Human Gaze
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool. Object Referring in Videos With Language and Human Gaze. In CVPR, 2018
2018
-
[59]
Exploring Object- Centric Temporal Modeling for Efficient Multi-View 3D Object Detection
Shihao Wang, Yingfei Liu, Tiancai Wang, Ying Li, and Xiangyu Zhang. Exploring Object- Centric Temporal Modeling for Efficient Multi-View 3D Object Detection. In ICCV, 2023
2023
-
[60]
Shihao Wang, Zhiding Yu, Xiaohui Jiang, Shiyi Lan, Min Shi, Nadine Chang, Jan Kautz, Ying Li, and Jose M. Alvarez. OmniDrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning. arXiv preprint arXiv:2405.01533, 2024
2024 arXiv
-
[61]
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. NeurIPS, 2022
2022
-
[62]
Argoverse 2: Next Generation Datasets for Self-driving Perception and Forecasting
Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, Deva Ramanan, Peter Carr, and James Hays. Argoverse 2: Next Generation Datasets for Self-driving Perception and Fo...
2021
-
[63]
Katharina Winter, Mark Azer, and Fabian B. Flohr. BEVDriver: Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving. arXiv preprint arXiv:2503.03074, 2025. 13
2025 arXiv
-
[64]
Language Prompt for Autonomous Driving
Dongming Wu, Wencheng Han, Tiancai Wang, Yingfei Liu, Xiangyu Zhang, and Jianbing Shen. Language Prompt for Autonomous Driving. In AAAI, 2025
2025
-
[65]
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Wenwei Zhang, Qi Alfred Chen, Ziwei Liu, and Liang Pan. Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives. arXiv preprint arXiv:2501.04003, 2025
2025 arXiv
-
[66]
Explainable Object-Induced Action Decision for Autonomous Vehicles
Yiran Xu, Xiaoyin Yang, Lihang Gong, Hsuan-Chu Lin, Tz-Ying Wu, Yunsheng Li, and Nuno Vasconcelos. Explainable Object-Induced Action Decision for Autonomous Vehicles . In CVPR, 2020
2020
-
[67]
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee K Wong, Zhenguo Li, and Hengshuang Zhao. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters , 2024
2024
-
[68]
BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning. In CVPR, June 2020
2020
-
[69]
Are Vision LLMs Road- Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
Tong Zeng, Longfeng Wu, Liang Shi, Dawei Zhou, and Feng Guo. Are Vision LLMs Road- Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding. arXiv preprint arXiv:2504.14526, 2025
2025 arXiv
-
[70]
Sigmoid Loss for Language Image Pre-Training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid Loss for Language Image Pre-Training. In ICCV, October 2023
2023
-
[71]
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning, 2025
Rui Zhao, Qirui Yuan, Jinyu Li, Haofeng Hu, Yun Li, Chengyuan Zheng, and Fei Gao. Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning, 2025
2025
-
[72]
Large Language Models Are Not Robust Multiple Choice Selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large Language Models Are Not Robust Multiple Choice Selectors. In ICLR, 2024
2024
-
[73]
Doe-1: Closed-Loop Autonomous Driving with Large World Model
Wenzhao Zheng, Zetian Xia, Yuanhui Huang, Sicheng Zuo, Jie Zhou, and Jiwen Lu. Doe-1: Closed-Loop Autonomous Driving with Large World Model. arXiv preprint arXiv: 2412.09627 , 2024
2024 arXiv
-
[74]
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
Xin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen, Yikang Ding, Dingyuan Zhang, Feiyang Tan, Hengshuang Zhao, and Xiang Bai. HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation. arXiv preprint arXiv:2501.14729 , 2025
2025 arXiv
-
[75]
Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma, and Alois C. Knoll. OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model. arXiv preprint arXiv:2503.23463, 2025
2025
-
[76]
Embodied Understanding of Driving Scenarios
Yunsong Zhou, Linyan Huang, Qingwen Bu, Jia Zeng, Tianyu Li, Hang Qiu, Hongzi Zhu, Minyi Guo, Yu Qiao, and Hongyang Li. Embodied Understanding of Driving Scenarios. In ECCV, 2024. 14 Appendix Table of Contents A Benchmark details 16 A.1 STSnu statistics . . . . . . . . . . . ....
2024
-
[77]
Ego turning left,
occlusion, 2) distance to the ego-vehicle, and 3) spatial distribution. Based on these criteria, we retain scenarios that are highly visible, occur in the near surrounding of the ego-vehicle, and are spatially well-distributed around it. The first criterion is straightforward ...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.