Pith. sign in

REVIEW 4 cited by

Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.19838 v2 pith:GGSMQB6X submitted 2024-03-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords autonomousdrivingem-vlm4admodelslanguagemodelsystemsanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-Language Models (VLMs) and Multi-Modal Language models (MMLMs) have become prominent in autonomous driving research, as these models can provide interpretable textual reasoning and responses for end-to-end autonomous driving safety tasks using traffic scene images and other data modalities. However, current approaches to these systems use expensive large language model (LLM) backbones and image encoders, making such systems unsuitable for real-time autonomous driving systems where tight memory constraints exist and fast inference time is necessary. To address these previous issues, we develop EM-VLM4AD, an efficient, lightweight, multi-frame vision language model which performs Visual Question Answering for autonomous driving. In comparison to previous approaches, EM-VLM4AD requires at least 10 times less memory and floating point operations, while also achieving higher CIDEr and ROUGE-L scores than the existing baseline on the DriveLM dataset. EM-VLM4AD also exhibits the ability to extract relevant information from traffic views related to prompts and can answer questions for various autonomous driving subtasks. We release our code to train and evaluate our model at https://github.com/akshaygopalkr/EM-VLM4AD.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Starting from a benign driving clip, SafeGen injects a realistic pedestrian and diffuses the video toward a collision, raising the failure score of three driving VLMs by 24.25% on average and improving later fine-tuni...

  2. SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation

    cs.AI 2025-07 conditional novelty 5.0 of 10

    SafeDrive228K is a 228K-example multimodal QA benchmark for traffic safety, and a graph-based RAG method improves VLM accuracy on it by 4.7 to 14.6 points across five models.

  3. DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A synthetic dataset and visual prompting framework improve VLM-based driving risk prediction, with the main evidence coming from an unreleased private test set.

  4. Automated Vehicles Should be Connected with Natural Language

    cs.MA 2025-06 conditional novelty 3.0 of 10

    A vision paper recommending natural language as the universal communication medium for connected and automated vehicles.

Pith tools