REVIEW 10 cited by
Talk2Car: Taking Control of Your Self-Driving Car
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A long-term goal of artificial intelligence is to have an agent execute commands communicated through natural language. In many cases the commands are grounded in a visual environment shared by the human who gives the command and the agent. Execution of the command then requires mapping the command into the physical visual space, after which the appropriate action can be taken. In this paper we consider the former. Or more specifically, we consider the problem in an autonomous driving setting, where a passenger requests an action that can be associated with an object found in a street scene. Our work presents the Talk2Car dataset, which is the first object referral dataset that contains commands written in natural language for self-driving cars. We provide a detailed comparison with related datasets such as ReferIt, RefCOCO, RefCOCO+, RefCOCOg, Cityscape-Ref and CLEVR-Ref. Additionally, we include a performance analysis using strong state-of-the-art models. The results show that the proposed object referral task is a challenging one for which the models show promising results but still require additional research in natural language processing, computer vision and the intersection of these fields. The dataset can be found on our website: http://macchina-ai.eu/
Forward citations
Cited by 10 Pith papers
-
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
An emotion-aware vision-language-action driving model estimates VAD emotion from commands and uses it to improve grounding and waypoint planning.
-
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
Box-QAymo introduces a box-referring VQA benchmark for autonomous driving, with hierarchical binary, attribute, and motion reasoning questions built from Waymo data and crowd-sourced labels.
-
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
A new 80K-clip dataset of unstructured driving scenarios with Q&A annotations improves VLA performance on NeuroNCAP and nuScenes benchmarks.
-
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
DVBench introduces 10,000 expert-annotated questions on crash and near-crash driving videos and reports that no tested vision LLM exceeds 40 percent accuracy under its strict GroupEval scoring.
-
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
Joint training on driving-knowledge QA (LingoQA, DRAMA) and CARLA trajectory data yields a VLM (WiseAD) that improves closed-loop driving score by 11.9% over trajectory-only training.
-
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
A hierarchical Mamba adapter (C-Mamba and Q-Mamba) improves multimodal LLM video understanding in autonomous driving, achieving SOTA 66.9% mIoU on DRAMA risk localization.
-
Autoware.Flex: Human-Instructed Dynamically Reconfigurable Autonomous Driving Systems
A human-instructed driving layer that uses an LLM with a retrieval-augmented Autoware knowledge base to translate natural language into validated AutoIR configuration commands.
-
Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.
-
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
TellTrack improves referring multi-object tracking by adding collaborative query matching, direct query-level language infusion, and a reordered cross-modal encoder, achieving SOTA HOTA on Refer-KITTI and Refer-KITTI-V2.
-
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.
Discussion (0). Continue with ORCID to comment.