Pith. sign in

REVIEW 4 cited by

On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05332 v2 pith:SFHPEHCT submitted 2023-11-09 cs.CV cs.AIcs.CLcs.RO

classification cs.CVcs.AIcs.CLcs.RO
keywords drivingautonomousreasoninggpt-4vscenebeencausalcomplex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The pursuit of autonomous driving technology hinges on the sophisticated integration of perception, decision-making, and control systems. Traditional approaches, both data-driven and rule-based, have been hindered by their inability to grasp the nuance of complex driving environments and the intentions of other road users. This has been a significant bottleneck, particularly in the development of common sense reasoning and nuanced scene understanding necessary for safe and reliable autonomous driving. The advent of Visual Language Models (VLM) represents a novel frontier in realizing fully autonomous vehicle driving. This report provides an exhaustive evaluation of the latest state-of-the-art VLM, GPT-4V(ision), and its application in autonomous driving scenarios. We explore the model's abilities to understand and reason about driving scenes, make decisions, and ultimately act in the capacity of a driver. Our comprehensive tests span from basic scene recognition to complex causal reasoning and real-time decision-making under varying conditions. Our findings reveal that GPT-4V demonstrates superior performance in scene understanding and causal reasoning compared to existing autonomous systems. It showcases the potential to handle out-of-distribution scenarios, recognize intentions, and make informed decisions in real driving contexts. However, challenges remain, particularly in direction discernment, traffic light recognition, vision grounding, and spatial reasoning tasks. These limitations underscore the need for further research and development. Project is now available on GitHub for interested parties to access and utilize: \url{https://github.com/PJLab-ADG/GPT4V-AD-Exploration}

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Study of Implementation Bugs in Multi-modal Agents

    cs.SE 2026-07 accept novelty 7.0 of 10

    First systematic taxonomy of 158 multi-modal agent bugs plus a runtime analyzer that recovers most open issues and surfaces 31 new ones.

  2. MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A skill-by-skill mixture-of-experts router lets a sub-3B vision-language model beat much larger models on autonomous-driving and robot-reasoning benchmarks.

  3. SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation

    cs.AI 2025-07 conditional novelty 5.0 of 10

    SafeDrive228K is a 228K-example multimodal QA benchmark for traffic safety, and a graph-based RAG method improves VLM accuracy on it by 4.7 to 14.6 points across five models.

  4. ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving

    cs.CV 2025-07 unverdicted novelty 1.0 of 10

    A workshop report documenting the ECCV 2024 W-CODA event, its accepted papers, speakers, and the dual-track corner case understanding and generation challenge.

Pith tools