Pith. sign in

REVIEW 2 cited by

Probing Multimodal LLMs as World Models for Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05956 v2 pith:RF2MKADZ submitted 2024-05-09 cs.RO cs.CV

classification cs.ROcs.CV
keywords modelsdrivingdynamicenvironmentsmllmsmultimodalworldability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We provide a sober look at the application of Multimodal Large Language Models (MLLMs) in autonomous driving, challenging common assumptions about their ability to interpret dynamic driving scenarios. Despite advances in models like GPT-4o, their performance in complex driving environments remains largely unexplored. Our experimental study assesses various MLLMs as world models using in-car camera perspectives and reveals that while these models excel at interpreting individual images, they struggle to synthesize coherent narratives across frames, leading to considerable inaccuracies in understanding (i) ego vehicle dynamics, (ii) interactions with other road actors, (iii) trajectory planning, and (iv) open-set scene reasoning. We introduce the Eval-LLM-Drive dataset and DriveSim simulator to enhance our evaluation, highlighting gaps in current MLLM capabilities and the need for improved models in dynamic real-world environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An asynchronous dual-system architecture forecasts large-model features into the current frame and adds a small-model update to drive in near real time, scoring 69.53 on Bench2Drive at 50 ms.

  2. GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoDrive conditions a frozen video diffusion model on a 3D-rendered version of the requested ego trajectory, cutting trajectory-following error by 42% versus Vista while using 99.7% less training data.

Pith tools