Pith. sign in

REVIEW 4 cited by

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.08725 v1 pith:RMPZYYQE submitted 2025-05-13 cs.CV

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

classification cs.CV
keywords lvlmsdrivemonkeyunderstandingautonomousdatasetdiversedrivingexisting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The Large Visual-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on front-view perspectives and partial objects within scenes, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D object localization and instruction understanding. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to improve 3D perception. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a 9.86% notable improvement on the 3D visual grounding task. The dataset and code will be released at https://github.com/zc-zhao/DriveMonkey.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    HERMES++ unifies 3D scene understanding and future geometry prediction in driving scenes via BEV representations, LLM-enhanced queries, a temporal link, and joint geometric optimization.

  2. GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

    cs.CV 2025-12 unverdicted novelty 6.0

    GaussianDWM uses 3D Gaussians with embedded linguistic features, language-guided sampling, and dual-condition generation for unified scene understanding and multi-modal output in driving world models.

  3. An interactive enhanced driving dataset for autonomous driving

    cs.CV 2026-02 conditional novelty 5.0

    A fused, interaction-labeled dataset of 7.31M driving segments with synthetic BEV videos and VQA pairs for training/evaluating driving VLMs.

  4. MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning

    cs.CV 2025-12 conditional novelty 5.0

    An autonomous-driving vision-language model that uses online RL over discrete language actions, with a separate action expert mapping decisions to trajectories, reports DS 78.04 and SR 55.09% on Bench2Drive.