Pith. sign in

REVIEW 9 cited by

VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04821 v2 pith:2TZRMTGU submitted 2024-08-09 cs.RO

classification cs.RO
keywords vlm-mpccontroldrivingautonomouscontrollerenvironmentmodelcomponents
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Motivated by the emergent reasoning capabilities of Vision Language Models (VLMs) and their potential to improve the comprehensibility of autonomous driving systems, this paper introduces a closed-loop autonomous driving controller called VLM-MPC, which combines the Model Predictive Controller (MPC) with VLM to evaluate how model-based control could enhance VLM decision-making. The proposed VLM-MPC is structured into two asynchronous components: The upper layer VLM generates driving parameters (e.g., desired speed, desired headway) for lower-level control based on front camera images, ego vehicle state, traffic environment conditions, and reference memory; The lower-level MPC controls the vehicle in real-time using these parameters, considering engine lag and providing state feedback to the entire system. Experiments based on the nuScenes dataset validated the effectiveness of the proposed VLM-MPC across various environments (e.g., night, rain, and intersections). The results demonstrate that the VLM-MPC consistently maintains Post Encroachment Time (PET) above safe thresholds, in contrast to some scenarios where the VLM-based control posed collision risks. Additionally, the VLM-MPC enhances smoothness compared to the real-world trajectories and VLM-based control. By comparing behaviors under different environmental settings, we highlight the VLM-MPC's capability to understand the environment and make reasoned inferences. Moreover, we validate the contributions of two key components, the reference memory and the environment encoder, to the stability of responses through ablation tests.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching

    cs.CV 2025-05 conditional novelty 6.0 of 10

    GeoVLM reranks top-10 candidates from a pretrained cross-view encoder by fusing image and text embeddings, improving top-1 retrieval on VIGOR, CVUK, and University-1652.

  2. DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Fine-tuning multimodal language models on a new SOTIF-focused driving dataset improves question answering and captioning, but the open-ended gains are measured by an LLM judge with no independent human scoring.

  3. LVLM-MPC Collaboration for Autonomous Driving: A Safety-Aware and Task-Scalable Control Architecture

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A system that lets a large vision-language model propose driving maneuvers while a model predictive controller verifies and safely executes them, rejecting or assisting unsafe lane changes.

  4. On-Board Vision-Language Models for Personalized Autonomous Vehicle Motion Control: System Design and Real-World Validation

    cs.AI 2024-11 conditional novelty 5.0 of 10

    An on-board fine-tuned vision-language model with a retrieval-augmented memory converts natural-language driving commands and camera views into MPC and PID controller parameters, reducing takeover rates by up to 76.9 ...

  5. VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving

    cs.RO 2025-07 conditional novelty 4.0 of 10

    VLM-UDMC uses a vision-language model to switch safety cost functions in a model predictive controller and a multi-kernel LSTM to predict traffic trajectories, reporting improved urban driving metrics in CARLA and cam...

  6. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

  7. Simulating the Unseen: Crash Prediction Must Learn from What Did Not Happen

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Crash prediction should learn from near-miss events and synthetic counterfactual scenarios, not just recorded crashes.

  8. A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

    cs.RO 2025-12 conditional novelty 3.0 of 10

    A position/review paper argues data-driven model predictive control is the best route to safe, adaptive, human-like autonomous-driving motion planning, but provides no new derivation or experiment.

  9. Research on Driving Scenario Technology Based on Multimodal Large Lauguage Model Optimization

    cs.CV 2025-05 conditional novelty 3.0 of 10

    On a private driving-scenario test set, a pipeline combining dynamic prompts, synthetic data, distillation with LoRA, and AWQ quantization raises average accuracy of a 7B vision-language model from 0.542 to 0.894.

Pith tools