REVIEW 4 cited by
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on on-board multi-view images and scene reasoning text, but this approach often lacks the holistic and nuanced scene recognition and powerful spatial awareness required for autonomous driving, especially in complex situations. To address this gap, we propose a novel vision-language framework tailored for autonomous driving, called LMAD. Our framework emulates modern end-to-end driving paradigms by incorporating comprehensive scene understanding and a task-specialized structure with VLMs. In particular, we introduce preliminary scene interaction and specialized expert adapters within the same driving task structure, which better align VLMs with autonomous driving scenarios. Furthermore, our approach is designed to be fully compatible with existing VLMs while seamlessly integrating with planning-oriented driving systems. Extensive experiments on the DriveLM and nuScenes-QA datasets demonstrate that LMAD significantly boosts the performance of existing VLMs on driving reasoning tasks,setting a new standard in explainable autonomous driving.
Forward citations
Cited by 4 Pith papers
-
OpenLongTail: Generative Scaling of Long-Tail Driving Data
Pose-informed diffusion with Plücker rays, depth warps, and cross-view memory converts monocular long-tail videos into multi-view assets that improve closed-loop driving robustness nearly to ground-truth multi-view levels.
-
Spatial-aware Vision Language Model for Autonomous Driving
LVLDrive fuses LiDAR point clouds into a pre-trained vision-language driving model through a zero-initialized gated Q-Former and a spatial-aware QA dataset, improving open-loop planning and 3D grounding over image-onl...
-
Can VLMs Unlock Semantic Anomaly Detection? A Framework for Structured Reasoning
SAVANT reformulates semantic anomaly detection as layered consistency verification, raising VLM recall by 18.5% on real driving images and enabling a fine-tuned 7B open model to reach 90.8% recall and 93.8% accuracy.
-
Can VLMs Unlock Semantic Anomaly Detection? A Framework for Structured Reasoning
SAVANT boosts VLM recall for semantic anomaly detection in driving images by 18.5% via structured reasoning and enables fine-tuning a 7B open model to 90.8% recall and 93.8% accuracy.
Discussion (0). Sign in to comment.