Pith. sign in

REVIEW 2 cited by

Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00988 v1 pith:655PGJ7A submitted 2024-01-02 cs.CV

classification cs.CV
keywords nuinstructdrivingmllmsmulti-viewtasksbev-inmllmexistinglarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rise of multimodal large language models (MLLMs) has spurred interest in language-based driving tasks. However, existing research typically focuses on limited tasks and often omits key multi-view and temporal information which is crucial for robust autonomous driving. To bridge these gaps, we introduce NuInstruct, a novel dataset with 91K multi-view video-QA pairs across 17 subtasks, where each task demands holistic information (e.g., temporal, multi-view, and spatial), significantly elevating the challenge level. To obtain NuInstruct, we propose a novel SQL-based method to generate instruction-response pairs automatically, which is inspired by the driving logical progression of humans. We further present BEV-InMLLM, an end-to-end method for efficiently deriving instruction-aware Bird's-Eye-View (BEV) features, language-aligned for large language models. BEV-InMLLM integrates multi-view, spatial awareness, and temporal semantics to enhance MLLMs' capabilities on NuInstruct tasks. Moreover, our proposed BEV injection module is a plug-and-play method for existing MLLMs. Our experiments on NuInstruct demonstrate that BEV-InMLLM significantly outperforms existing MLLMs, e.g. around 9% improvement on various tasks. We plan to release our NuInstruct for future research development.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A protocol model writes editable JSON layer specs, then a diffusion model fills in a matching background, producing multi-layer posters that beat several commercial tools in a small benchmark.

  2. 2nd Place Solution for CVPR2024 E2E Challenge: End-to-End Autonomous Driving Using Vision Language Model

    cs.CV 2025-09 conditional novelty 3.0 of 10

    A single-camera vision-language-model system scored 0.8747 on the CVPR 2024 E2E driving benchmark, the best camera-only result.

Pith tools