Pith. sign in

REVIEW 10 cited by

Talk2Car: Taking Control of Your Self-Driving Car

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.10838 v2 pith:FW25LWG5 submitted 2019-09-24 cs.AI cs.CLcs.RO

classification cs.AIcs.CLcs.RO
keywords commandcommandsdatasetlanguagenaturalobjectactionagent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A long-term goal of artificial intelligence is to have an agent execute commands communicated through natural language. In many cases the commands are grounded in a visual environment shared by the human who gives the command and the agent. Execution of the command then requires mapping the command into the physical visual space, after which the appropriate action can be taken. In this paper we consider the former. Or more specifically, we consider the problem in an autonomous driving setting, where a passenger requests an action that can be associated with an object found in a street scene. Our work presents the Talk2Car dataset, which is the first object referral dataset that contains commands written in natural language for self-driving cars. We provide a detailed comparison with related datasets such as ReferIt, RefCOCO, RefCOCO+, RefCOCOg, Cityscape-Ref and CLEVR-Ref. Additionally, we include a performance analysis using strong state-of-the-art models. The results show that the proposed object referral task is a challenging one for which the models show promising results but still require additional research in natural language processing, computer vision and the intersection of these fields. The dataset can be found on our website: http://macchina-ai.eu/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving

    cs.CV 2025-12 conditional novelty 6.0 of 10

    An emotion-aware vision-language-action driving model estimates VAD emotion from commands and uses it to improve grounding and waypoint planning.

  2. Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Box-QAymo introduces a box-referring VQA benchmark for autonomous driving, with hierarchical binary, attribute, and motion reasoning questions built from Waymo data and crowd-sourced labels.

  3. Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 80K-clip dataset of unstructured driving scenarios with Q&A annotations improves VLA performance on NeuroNCAP and nuScenes benchmarks.

  4. Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding

    cs.CV 2025-04 conditional novelty 6.0 of 10

    DVBench introduces 10,000 expert-annotated questions on crash and near-crash driving videos and reports that no tested vision LLM exceeds 40 percent accuracy under its strict GroupEval scoring.

  5. WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Joint training on driving-knowledge QA (LingoQA, DRAMA) and CARLA trajectory data yields a VLM (WiseAD) that improves closed-loop driving score by 11.9% over trajectory-only training.

  6. H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A hierarchical Mamba adapter (C-Mamba and Q-Mamba) improves multimodal LLM video understanding in autonomous driving, achieving SOTA 66.9% mIoU on DRAMA risk localization.

  7. Autoware.Flex: Human-Instructed Dynamically Reconfigurable Autonomous Driving Systems

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A human-instructed driving layer that uses an LLM with a retrieval-augmented Autoware knowledge base to translate natural language into validated AutoIR configuration commands.

  8. Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.

  9. Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking

    cs.CV 2024-12 conditional novelty 4.0 of 10

    TellTrack improves referring multi-object tracking by adding collaborative query matching, direct query-level language infusion, and a reordered cross-modal encoder, achieving SOTA HOTA on Refer-KITTI and Refer-KITTI-V2.

  10. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.

Pith tools