Pith. sign in

Paper Citation Record · LEDGER

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

As of 21 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 6 inbound Pith citation observations for arXiv:2505.08725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08725 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:53:10.078556Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:58:26.525113Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T17:10:25.159616Z

Reference resolution

100 of 103 outbound references displayed

  • verified exact2
  • verified fuzzy46
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af0e2db5-97df-48e6-ba34-eef2b8fb4744 · outbound

This paper cites Visual instruction tuning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Visual instruction tuning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.606344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.606344Z digest=sha256:5525765caea1555d632d8442e4e87c38380ced0c9a061fafa4cdf3fa4d00fb7b

Observation 02b1dff0-eea1-4800-95eb-7dd44a64bac4 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.610547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.610547Z digest=sha256:32e3ad73b8066f4f41d68209420332fba8864c4dfc1411b7fb37f9efaaa06f8c

Observation a87bd67b-a921-4263-8acb-e453cf170b50 · outbound

This paper cites GPT-4 Technical Report.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.614593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.614593Z digest=sha256:60d94f0c9851908c9c01a581745aaf2f219fc9ca8a998c28ce015070246a0ec2

Observation 15eb341e-5305-44cb-8388-04eba5b584f1 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.618465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.618465Z digest=sha256:8e6b48ce860a117a9826334ee4412ab5341a8e9a44a6b0dea419ebd161faa7a6

Observation 5ef1f138-a931-4f50-93bd-25247ad7119c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.622381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.622381Z digest=sha256:5101526296a1d0b59ba3568cc50b069fa72bc35139ad51c1df8480aecf86254d

Observation b3f29d88-3a74-43d1-bdc2-cf4ea0137144 · outbound

This paper cites Qwen2.5-VL Technical Report.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.626041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.626041Z digest=sha256:4a5bb3951d95b958491cd103983cf69af220d0d7660e6ff4b09b9da6d4853436

Observation a200733a-a8c2-4235-be19-f139feb09e1d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Gemini: A Family of Highly Capable Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.630196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.630196Z digest=sha256:b0c26707f18249a8511ffe9f60ee6e457751221695ba280fd04dd16037d8484e

Observation 0ffcde34-759e-40dc-b65a-8b9352e44612 · outbound

This paper cites Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Navcot: Boosting llm-based vision-and-language navigation via learning disentangled reasoning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.633985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.633985Z digest=sha256:cc8b90304e701003f22e3a40902d4570b0e730c7bd59b694d20e89c4b5fb36e8

Observation 241fb771-5cb1-4263-a7f6-dce7ab92e8d5 · outbound

This paper cites InternLM2 Technical Report.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving InternLM2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.637512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.637512Z digest=sha256:fab418eb215840a3f60631c18c1fc145892f6093441c6649903a105e665ab6ff

Observation fc2a629a-9792-4350-b954-26c79c47f028 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LLaMA: Open and Efficient Foundation Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.741034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.741034Z digest=sha256:9a61ef92955e1772f366774a8a7324f0c60dee8a9639ac14fd250ac49dbcfbc2

Observation b4472e26-9493-49ce-8ba0-a4559cdb6f64 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.744895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.744895Z digest=sha256:ed996d67e8d44f77dd4b4847e4cc03012c8e7b815f4ff60fb898e93cd1a1770e

Observation 150e5978-133c-46b3-901e-938a75c08f7c · outbound

This paper cites Holistic autonomous driving understanding by bird’s-eye-view injected multi- modal large models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Holistic autonomous driving understanding by bird’s-eye-view injected multi- modal large models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.748705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.748705Z digest=sha256:62412396c664f29cda39cfc86962fc82a7205a6167cf8c760a9b0660cf9f0789

Observation da94fea8-9b4a-4abd-a4ed-6ac2c33d19fd · outbound

This paper cites Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.752104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.752104Z digest=sha256:695cd9e0eb6a7d715c5c668efefb9a0575f0a076bd33a31bfe9b6f0f9ac31fb8

Observation 13a885fc-63c3-4161-8387-c7ea46159ffd · outbound

This paper cites Embodied understanding of driving scenarios,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Embodied understanding of driving scenarios,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.755779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.755779Z digest=sha256:ddb4108f10cf78659d1d481f9b83919b7b15ae1e9de046a967b5f7343e23406a

Observation 72c513d5-8c01-4cc3-b493-8013c377c597 · outbound

This paper cites Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Omnidrive: A holistic llm-agent framework for autonomous driving with 3d perception, reasoning and planning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.759471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.759471Z digest=sha256:261616f37bd802fcbc34e49644128f2237b98050f0dbca0af9321be194a55dc5

Observation 0c2f36b6-edc5-45b5-8277-056f7ea8f29c · outbound

This paper cites Drivevlm: The convergence of autonomous driving and large vision-language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivevlm: The convergence of autonomous driving and large vision-language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.763107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.763107Z digest=sha256:eab4934d7090e81056b3e9479a2c5526ed0389edfb1e42d85dcb7397ef27ee32

Observation e0c4f03a-4367-47b4-bd00-9f2f6084ae81 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.766656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.766656Z digest=sha256:128ccdef89249ec3517b1b1e1f89bd40870b81d85593f01af1d8a05c7e334136

Observation b1266064-9caa-47b3-af39-d24e24de19d3 · outbound

This paper cites WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.771166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.771166Z digest=sha256:2d57926ba3402bc1e0dad20fd353df87bf756dc06f01cdf59871e914d0f3a914

Observation 80804ced-713e-4aeb-bfe1-5f89af80d114 · outbound

This paper cites RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.775204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.775204Z digest=sha256:951c73e17d97fae1f1a408128b90c8222840c11cf1de6da1148eecd34b2347d8

Observation 057b03da-0b8b-4083-97b3-9f1a6283553d · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lingoqa: Visual question answering for autonomous driving,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.779230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.779230Z digest=sha256:726d2a87cc5affcc7d510f79b93251faa5e3b4debdd0b0e16c16606f269a6772

Observation c0a35f8a-2bbb-4e2e-acbd-042353accd74 · outbound

This paper cites Automated evaluation of large vision-language models on self-driving corner cases,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Automated evaluation of large vision-language models on self-driving corner cases,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.782949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.782949Z digest=sha256:d0a8d30af1a5cdd378f450dae83842e1e75286533ee868004e8c7d2f525a9ac0

Observation 1a66bee5-b1e7-4988-905e-38537054d902 · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.786674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.786674Z digest=sha256:94b0c4ed16fa05c72f8f36844e39f24e62a9ee570489d0561e0344261e8cc826

Observation 7e553ca0-123d-4a98-a893-c797c35e4a75 · outbound

This paper cites Drivelm: Driving with graph visual question answering,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivelm: Driving with graph visual question answering,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.790238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.790238Z digest=sha256:e8863adb75a7fe9f137b1b213905f52f2c63288ee76580749bd1790d3111e748

Observation edc6e766-13cd-4d68-81b5-1b6ea6c962a3 · outbound

This paper cites Language prompt for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Language prompt for autonomous driving,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.793816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.793816Z digest=sha256:3f961f9ed4ef630188be59357d7cd4556f470f02d152f6c758735fea672d4f52

Observation 35ee3faa-f307-44e3-9b6c-d28be498676b · outbound

This paper cites Language-Image Models with 3D Understanding.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Language-Image Models with 3D Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.797343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.797343Z digest=sha256:81f2427b184a62bb5b72dbd0ce3301b329f4392becf153f71439ec8719dd47f9

Observation 9d731add-91c8-40da-af82-0540c830deeb · outbound

This paper cites EMMA: End-to-End Multimodal Model for Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving EMMA: End-to-End Multimodal Model for Autonomous Driving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.801297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.801297Z digest=sha256:2bb5e71607cbd2f8414394d799273373ae3196a55e0fa0b6f9770764ad1f2e31

Observation c816d584-4c43-467a-938d-9a22b791c108 · outbound

This paper cites Talk2car: Taking control of your self-driving car,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Talk2car: Taking control of your self-driving car,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.804694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.804694Z digest=sha256:cb268f3de522c8e1b9c4992fac91fa789f519929dc1b5553733e7b91f77a142e

Observation cc639249-fb3a-406e-bdc0-a2413c90e37a · outbound

This paper cites Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.808274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.808274Z digest=sha256:baa6c964e7f552eb7275a4e271761433137267dea5ad50d0bb235ec8c97f9c8f

Observation 7ab54154-d3f1-4af2-8559-f7a4be381719 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Improved baselines with visual instruction tuning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.811720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.811720Z digest=sha256:1724f1d559c530c3a5da00f11730c2018df9662a6714a7b68b8bff759a502a16

Observation a236f099-dca0-40ab-ade2-e099f95af1b5 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.815225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.815225Z digest=sha256:2167752be204a76508e9980778649ca38e395b066ac0ea69e449fb275f30514b

Observation b77ab0f0-04c0-49da-9d91-68a022e90d8e · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.818802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.818802Z digest=sha256:4da2a6c2735b264db7c7f7ace275b7c5ae214a86b25768d2793e34507f177754

Observation 80c44b41-cc05-47fd-8529-3f803263c0e4 · outbound

This paper cites Pix2seq: A language modeling framework for object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Pix2seq: A language modeling framework for object detection,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.822281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.822281Z digest=sha256:3357e028b5c6c5e2693c9cd513ede4c5419e5531ed12bfbd5a72240fc4d40b51

Observation 509210da-73bd-4da1-97e3-0c23bb99be10 · outbound

This paper cites Petr: Position embedding transformation for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Petr: Position embedding transformation for multi-view 3d object detection,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.825816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.825816Z digest=sha256:d8fd9cdb8d97cecffb6e93f217170cc6cd6637acda8bc11fe1ce57e50009118e

Observation 05ce9192-cab8-4dea-84d2-aec9baa3765b · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.829420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.829420Z digest=sha256:79b88d56a7dafe6ac04657d24dc1dcd450bc2340e786448d6590d7dc7726f2a6

Observation 69ec23d1-b54e-4a2f-a360-7286e37d526f · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.833021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.833021Z digest=sha256:3ff8fc6b43a1870a021541deef938c38f5a38a46642aaa511f545bc3c0c0b226

Observation f18875db-46df-4a97-8b72-06eda01de60c · outbound

This paper cites Grounding human-to-vehicle advice for self-driving vehicles,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grounding human-to-vehicle advice for self-driving vehicles,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.836649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.836649Z digest=sha256:399437480b5a13e1ef5c995e36da99cdb4f6b579bba18e2de34c36547fa4e1b7

Observation 9b1f4411-3818-4d74-b3eb-0d99dd8081e6 · outbound

This paper cites HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.840222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.840222Z digest=sha256:0ac48fcff8422517850f9c80420042b9c7638f33bbe49fe778eebfa92ef87175

Observation 81c503aa-8639-4c2a-a326-8e878e824e3e · outbound

This paper cites Drama: Joint risk localization and captioning in driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drama: Joint risk localization and captioning in driving,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.844453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.844453Z digest=sha256:d2112c039b9de4959d5fb3ebd16303877f4b6d191a2adae1acf0989c44df368a

Observation fa4099c1-778f-4a97-bf85-ad9358af885d · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.848095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.848095Z digest=sha256:b984b41d9f7bcd1a19701aae9406dd1c5e69875c40fbe53a7c8fcc19fbbabd9d

Observation 076f3704-9b96-4ee8-ad67-a00a728329a6 · outbound

This paper cites Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.083936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.851562Z digest=sha256:c291d669a929e07b6ed5680cb313b77170f3c9cf139d9e67194ee2c0f6788c84

Observation 8c525c83-1dde-4239-b171-898e93d02cdc · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.855200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.855200Z digest=sha256:57bdfd0861de1062af44e2185f1f72e6004758a6823032e203ed1d78aa929d19

Observation b0351992-9a9b-40fc-a141-ae809611e947 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.063220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.858638Z digest=sha256:cb60c76d93fccaf7ce1ad877be642b6234568dd89dc577d34f131f89b4cf97bd

Observation 72584531-c02e-4537-af12-46682891f22e · outbound

This paper cites Llava- next: Improved reasoning, ocr, and world knowledge,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Llava- next: Improved reasoning, ocr, and world knowledge,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.049928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.862308Z digest=sha256:8a5fdc461ba39f2a4bb49b316c34fb10e81d396b5468d80e0072c6554d9b4dcc

Observation 15b41a1e-7211-4ac7-856c-dfebccaf96f4 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Learning transferable visual models from natural language supervision,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.037147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.865801Z digest=sha256:58b80e43c32af935b9f16d8cae20be0521792e72266dfd3e3901ed9ee11f3ee1

Observation 62406d58-a7b0-40aa-b1d7-0805bdf5e0b1 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Sigmoid loss for language image pre-training,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.025373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.869255Z digest=sha256:a9ea93af9589612b8b354fd707b06d92e3a902d3b5a4ecf9921f1eaae070e3b5

Observation 569009a9-a579-4d5d-99f5-95ba72d36488 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Dinov2: Learning robust visual features without supervision,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:11.013827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.872768Z digest=sha256:e6a9aa053d9de1c5ef0b31436300dd14a3013b9205b7c92f377038b3be880be4

Observation 2e91f92a-2861-4361-a337-d7a9656543d7 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Vision-language models for vision tasks: A survey,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.876303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.876303Z digest=sha256:cdc110ba59ae268817ff5382c7dae47ae9e1fbc9a2888d4f9c557b4408dbbae0

Observation ecec2b5d-db96-4866-afa2-18be6461c90a · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Monkey: Image resolution and text label are important things for large multi-modal models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.995209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.880018Z digest=sha256:bb46d2355334e298939f9b5b84e07ce36a5eb2cabdc402234020d4b6de3356f4

Observation 9e14063f-7254-4a27-a907-eb4952f5b386 · outbound

This paper cites Uni-moe: Scaling unified multimodal llms with mixture of experts,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Uni-moe: Scaling unified multimodal llms with mixture of experts,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.983833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.884186Z digest=sha256:c165c1c1baf8833a4d1adf0b93d61d910da8e0bd6eb1dde794e7f61de2478346

Observation d7a730da-cb98-432f-a98d-07c5897e15e7 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.972612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.888040Z digest=sha256:aed6df09c852895e88e59f87ccc66e52f8f0a6cbba98636252faa9937f88b9f5

Observation 9f16af07-e6cc-4f25-b1e9-18f1ae6e4f7d · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Ferret: Refer and ground anything anywhere at any granularity,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.961906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.891792Z digest=sha256:f0b718a191cfac37755c7b6edf0cc1546aa2b9712fc8fe0e03af3ef92c471515

Observation 96470717-e76d-424b-bbc2-d0e1043c5001 · outbound

This paper cites Ferret-v2: An improved baseline for referring and grounding with large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Ferret-v2: An improved baseline for referring and grounding with large language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.950815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.895836Z digest=sha256:dff1fb5986b9e246c54886b9174e6c664c955c2a45917782791e4a1315e3130f

Observation daad5a10-c313-49e9-a894-0087c9786e26 · outbound

This paper cites Relationlmm: Large multimodal model as open and versatile visual relationship generalist,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Relationlmm: Large multimodal model as open and versatile visual relationship generalist,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.940250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.899785Z digest=sha256:d8cb606d99b547625a72e4cdb4738ce91a80e8a37c23f093ad2b64f91fe7c818

Observation 0cf9eba9-e938-443c-9dbd-99c564403b34 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Groma: Localized visual tokenization for grounding multimodal large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.929883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.903100Z digest=sha256:451226243ff68fac274cc734b97ea0bb29adf25e87425df34a2792962d21f6a0

Observation 40a5d163-4b92-494d-9f53-009a63553b0c · outbound

This paper cites Segment anything,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Segment anything,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.919161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.906547Z digest=sha256:f5466bc4cc611e44a92d60ef88414b4590fe68c77cdbe6207c8fd72c900b0c4d

Observation 513d95c0-5368-4b54-845a-297bb16acb97 · outbound

This paper cites Masked-attention mask transformer for universal image segmenta- tion,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Masked-attention mask transformer for universal image segmenta- tion,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.909006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.910064Z digest=sha256:555665bbef89b608d86ae064dce7d99a30fbb8c188b803278f94e0320214beb6

Observation 8890a925-f7ab-44b7-8f6b-fe1c213b118e · outbound

This paper cites Lisa: Reasoning segmentation via large language model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lisa: Reasoning segmentation via large language model,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.898061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.913619Z digest=sha256:cf80020d14b8884c47a600d1bc3dce70ac59fa28f11389af8b4853f04753602f

Observation 7293f22d-464a-45d6-8aa3-e02dc3fedd46 · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Psalm: Pixelwise segmentation with large multi-modal model,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.886398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.917950Z digest=sha256:f52c4000842fe4aa19b24a76fc86ebc36472dadce518accf01c6e01c6b2a14dd

Observation 95e58821-a84a-4ab3-a11c-13d0928aae29 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.874930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.921672Z digest=sha256:f68527918109eaa37a6bd22a41bb36490bec80380ab52e086320394c0f51c3f5

Observation 355d12dc-ed99-4b68-89bb-ebbfffa8cd80 · outbound

This paper cites Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Jm3d & jm3d- llm: Elevating 3d representation with joint multi-modal cues,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.863558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.925108Z digest=sha256:04647689197cbcbfba298cf0947d0b8e063db7803514c92e451cecf2a91f46ad

Observation 08b527f2-7726-4b34-bb2f-561516633621 · outbound

This paper cites Tod3cap: Towards 3d dense captioning in outdoor scenes,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Tod3cap: Towards 3d dense captioning in outdoor scenes,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.852934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.928355Z digest=sha256:b54b0977d8f0e8f807a6f1d725c3a0fe4a6dd62c6fb86b12e1886a63707e8fc4

Observation 56bc0e69-faaa-47f5-92dc-7ef0993400a4 · outbound

This paper cites Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.931808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.931808Z digest=sha256:06d58712b2a4919df35476a102125e5aeaef742b8e9ec957644c450796601093

Observation eeeb4bff-6e2e-4bb1-a37f-74f3fd7af46d · outbound

This paper cites Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.935135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.935135Z digest=sha256:755d17d035929c6d13fbd169a22ecf290b5805889370c7ef8ebcab342873faf4

Observation a968a4f6-892f-47ca-8685-b032c9d935de · outbound

This paper cites Gpt-driver: Learning to drive with gpt,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Gpt-driver: Learning to drive with gpt,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.841943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.938984Z digest=sha256:aa1e8854f7a1cb30f209c13b651e98c8c814a2fe37422fdedb3fc441bac9b331

Observation badf7126-f77f-42d7-9abf-e50ca4687240 · outbound

This paper cites Making large language models better planners with reasoning-decision alignment,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Making large language models better planners with reasoning-decision alignment,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.830737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.942801Z digest=sha256:57bcede85d91a42d7a5d434400b18f506155e3599820cb6030eaecd3f6db3165

Observation 120ab713-3f78-47f2-ab94-13d325e1a21d · outbound

This paper cites Distilling multi- modal large language models for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Distilling multi- modal large language models for autonomous driving,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.819866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.946579Z digest=sha256:18f48369a805155542691973eb0994807cad7ef7844bb07d0e8c7b131b3d356a

Observation c75dc518-574d-49aa-b789-f44dc378b1e9 · outbound

This paper cites Generative plan- ning with 3d-vision language pre-training for end-to-end autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Generative plan- ning with 3d-vision language pre-training for end-to-end autonomous driving,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.808574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.950415Z digest=sha256:2a35b3a3a8d3663e90eeea7c10338909be760a392d92fc84faa7db8809be9fc5

Observation 78706104-26be-41cf-a195-1bdadec92af9 · outbound

This paper cites World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:53:10.287648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.955286Z digest=sha256:c267a19ff8f61c88d57be399df9e0248aa4c9545c7e3117ac9969c2ba9a4a55a

Observation 70d8969c-ee11-425b-a5c5-65bb897228a4 · outbound

This paper cites LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:53:10.269257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.960199Z digest=sha256:35d233545bf4f96b7a8fddbd79d051eacadcc4c123ee0cb6423c59d88abd2849

Observation 451ac5ed-dfca-4ea8-828a-f01a0dca9c0d · outbound

This paper cites Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.797193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.963841Z digest=sha256:56a92b27dbedf31252ad2c6ebffb292c1c067de4cced8bfafdde553dd579a656

Observation 34407945-3071-440e-b026-cf31d84494f1 · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.967292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.967292Z digest=sha256:31f3091ada585ddae607509194c6c809d12f6f220a3aafa740755e073323ae78

Observation b970df43-224a-4f28-a776-41ddd56c168c · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lmdrive: Closed-loop end-to-end driving with large language models,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.785869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.970639Z digest=sha256:3dde45ffa3cf681fc8d038293ff73a36ad767104fa17626bcfd9e4616ff54dcb

Observation 337b915b-67f3-424f-92a9-99247d4290c6 · outbound

This paper cites ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.974080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.974080Z digest=sha256:58b593bf11f528812d6b3a5b2072df92f64bc99578dcfac85d200bb3fe82b26c

Observation 4663e554-d863-45b0-9644-94e37527db6d · outbound

This paper cites Graph-detr4d: Spatio-temporal graph modeling for multi- view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Graph-detr4d: Spatio-temporal graph modeling for multi- view 3d object detection,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:09.978067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:09.978067Z digest=sha256:8d8277d5f00a5fed83f3f0965ac89652a73b2aa6db5a736f7c1f5904b8cc3210

Observation d988ba15-8d25-4f99-ac2c-019b819fbe97 · outbound

This paper cites Physically realizable adversarial creating attack against vision-based bev space 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Physically realizable adversarial creating attack against vision-based bev space 3d object detection,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.768257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.981826Z digest=sha256:c6804f4cef767618a21c558b36c334765f655cee48ab592e4a01ce8506de7dbc

Observation af0a7518-2426-4606-a6e0-fcd1212d1553 · outbound

This paper cites Notice of violation of ieee publication principles: Recent advances in 3d object detection in the era of deep neural networks: A survey,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Notice of violation of ieee publication principles: Recent advances in 3d object detection in the era of deep neural networks: A survey,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.756585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.985380Z digest=sha256:df029d6ae7a716272824da6e1daff599e202afa624cbd9e8c18ce8653530cf69

Observation 36daa477-5ef5-4f81-9a69-6658308eb72a · outbound

This paper cites Obmo: One bounding box multiple objects for monocular 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Obmo: One bounding box multiple objects for monocular 3d object detection,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.744729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.988903Z digest=sha256:564a87e9119b91e873972090bab3992d6731e4707ad62a4abde9cb39278cf6aa

Observation a91342ba-febe-4693-80fa-49e53d5a7757 · outbound

This paper cites Stereoscopic vision recalling memory for monocular 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Stereoscopic vision recalling memory for monocular 3d object detection,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.733383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.992940Z digest=sha256:306edb8af1b0622cc60678a1e0a684e04d2d7438e70c1953d357b8fa1673ff74

Observation 643cd81b-df84-4ca9-aa99-79e395cff86f · outbound

This paper cites X-view: Non-egocentric multi-view 3d object detector,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving X-view: Non-egocentric multi-view 3d object detector,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.722634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:09.996622Z digest=sha256:c7e460b60bbddc71d1943973883da4549e77afcdf95b869b015f015efc0d3d3c

Observation fcfde415-e9c7-4027-abf6-73bcda875309 · outbound

This paper cites Do-sa&r: Distant object augmented set abstraction and regression for point-based 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Do-sa&r: Distant object augmented set abstraction and regression for point-based 3d object detection,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.712088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.000788Z digest=sha256:0288ad3e4a8951a00a5537e5bf558dcaeae7dcd1be7a4ace86d3f90c0b2808b0

Observation 2ed55f55-0093-499b-b0f8-c1f22cb69acb · outbound

This paper cites 3d cascade rcnn: High quality object detection in point clouds,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving 3d cascade rcnn: High quality object detection in point clouds,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.701126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.004869Z digest=sha256:83432bbee5aee2136dd7e76b669ee18a96faf23efa7bd4d11b8f768792f5c8a4

Observation 14f4dbbe-42fb-4eba-906d-4882482e3aab · outbound

This paper cites Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.690077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.009010Z digest=sha256:12596f57631467f2f009d5d2ac485c3fc8494d7833170743985aee4493c525da

Observation 90da7222-d46c-4883-bc41-1523eb343adc · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.013003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.013003Z digest=sha256:3ea3137200bae852cca9e02485fff14594d12c0cb36d8db500ae0aa589bac019

Observation c9199523-fb97-4b64-b608-9838b3aad4fa · outbound

This paper cites Petrv2: A unified framework for 3d perception from multi-camera images,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Petrv2: A unified framework for 3d perception from multi-camera images,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.679295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.017715Z digest=sha256:97897c3efc9675f1b55333e0dfc9ae8eaeef26d591939751e08fc957e894b4d2

Observation fb7f6392-75cd-4790-b24d-2cc0897914e9 · outbound

This paper cites Cape: Camera view position embedding for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Cape: Camera view position embedding for multi-view 3d object detection,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.667963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.021479Z digest=sha256:50657e4a821b7cb246157b10b147a663c3a64141fa4854a02967f0fec5b48d02

Observation 503998d9-5ce9-4875-a161-f574d5ed3826 · outbound

This paper cites Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.656426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.025262Z digest=sha256:5e0c94974e96e2ed6be56126b0d9381cc7f799252ed1e86ab4be5167453f3f79

Observation 9f196b30-cefa-46da-8530-3ee3c469f82b · outbound

This paper cites Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.645704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.028772Z digest=sha256:a6dbd35d9d5a1a7f0865aaf87419d91695a63722b5a68b416457fccb647f4397

Observation a44a9a7c-1fa1-4556-948d-3fd3fe9ba1a9 · outbound

This paper cites Open: Object-wise position embedding for multi-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Open: Object-wise position embedding for multi-view 3d object detection,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.633351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.032996Z digest=sha256:fda51a5c663cd61bd9fea1981718c8fbd0fa8a8f6653d2f8702b131500c7f638

Observation 586804eb-24b7-4830-8eea-790c4d954b48 · outbound

This paper cites Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.621918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.037169Z digest=sha256:5f9cccaa8b9ab1c354869022cdcc85a4092c34fdd90d029da99b3a49986f132c

Observation 4d4d5945-1cba-42e6-a12a-42b464cfeb3b · outbound

This paper cites Far3d: Expanding the horizon for surround-view 3d object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Far3d: Expanding the horizon for surround-view 3d object detection,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.610574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.041250Z digest=sha256:3f2df5efefddd9c041838d8fd99556010c1dc23395098717fde18ab169d83690

Observation 31c094b9-5d34-429f-8a0c-d46f01ac88d0 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.044871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.044871Z digest=sha256:184b090086ebbba7bf8f906c8dd02937b5366be37467fa7e20952a9a259deaa8

Observation eaa0a421-08be-4a19-b4e9-0b8ef187121e · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving nuscenes: A multimodal dataset for autonomous driving,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.048887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.048887Z digest=sha256:e7aa9ef6b37c8c6a85965ce33602a0c66c02c35e2fbd5c3044e55ff51f2c6ad1

Observation 6e332cea-28fc-470f-9fc2-fc05ce598759 · outbound

This paper cites Grit: Faster and better image captioning transformer using dual visual features,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grit: Faster and better image captioning transformer using dual visual features,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.591964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.052761Z digest=sha256:57c62f8970f04abfc04f88794c748adf5a94f17690173f230097df857f96c5b4

Observation 7038bcaa-0acd-43ba-8388-8ab3c01b7303 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.056914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.056914Z digest=sha256:58763512e0639396ab230b6a3a3a033d37f8dd0689b261969ca44606eddff865

Observation a5ce548a-ed9b-412f-9b67-fa8a857000bb · outbound

This paper cites Glamm: Pixel grounding large multimodal model,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Glamm: Pixel grounding large multimodal model,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.578397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.060557Z digest=sha256:e1edc88ac44ce3c2425b7b982215b86084cee52f320cf9f8a361b8ee70169646

Observation 6112512b-2920-4f8f-85a2-90a350a19ca2 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Lora: Low-rank adaptation of large language models,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.566036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.064117Z digest=sha256:2e676e4e8c91540d727288e13b4e0e35a457e1564538010d1851f8ada643b559

Observation 6fd371b2-07d9-4065-acd9-4430f1899597 · outbound

This paper cites Focal loss for dense object detection,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Focal loss for dense object detection,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.553864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.067413Z digest=sha256:1ca0935a47576b813bd6439111ff3362a07c8a3cbe61f5adfca7acc6025c8c8f

Observation f57b63ec-c6e1-4abf-b89d-cc156f3b585d · outbound

This paper cites The hungarian method for the assignment problem,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving The hungarian method for the assignment problem,

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.071387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.071387Z digest=sha256:7f5b35d594741508e92f953fed73015d518b14a12bf274d9a0b20bd1db518a33

Observation a83705ec-c02e-4de1-ac68-1ebdfa929b7b · outbound

This paper cites Decoupled Weight Decay Regularization.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Decoupled Weight Decay Regularization

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T21:53:10.074870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:53:10.074870Z digest=sha256:c7e00b93b20da55d43330acdd66f2b2edf7e765064022a7e342eaef27e3006dd

Observation 3ef449d7-7ebf-4564-a457-296bffa66e82 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving Cider: Consensus- based image description evaluation,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:53:10.534683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:53:10.078556Z digest=sha256:5cd37fa4535b23bb5b173c8bc7a9812e7d396ab78b22645f38264eb6af57492e

Pith citing papers

Observation 155de858-f381-424b-992f-64d33929750e · inbound

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation cites this paper.

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:26.525113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:26.525113Z digest=sha256:a8f41622a8fbaa230770318e149f9afd0f0d5edcc8723975479e8236fe8a641d

Observation e782e27f-1a09-42ab-bbfc-eb2796f514da · inbound

A Survey on Vision-Language-Action Models for Autonomous Driving cites this paper.

A Survey on Vision-Language-Action Models for Autonomous Driving Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:04.895175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:04.895175Z digest=sha256:8c6dbedfaa2b4af9fdf5934973741b02b24dd11bc1fd8b0f15700b66175c0696

Observation 1e6634ed-9689-4ee4-83a8-c1f5137017f3 · inbound

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning cites this paper.

MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:29.207219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:29.207219Z digest=sha256:0a51073f40d4c3705c1af728e55e20fa27cf2321590b91689f86c2eea220429d

Observation a561ba29-651a-49e2-bdc8-6f643e12bc5e · inbound

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation cites this paper.

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:10:25.162426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T17:06:34.398973Z digest=sha256:d72cb67bfb5e44c49cc705394418fb4764dcd129947aad9ddcaaf1789ad23e9c

Observation 42e96adf-e3de-4b3a-932b-4faf347be58e · inbound

An interactive enhanced driving dataset for autonomous driving cites this paper.

An interactive enhanced driving dataset for autonomous driving Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T21:19:52.534231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:19:52.534231Z digest=sha256:31364dfc3a3aede9bdb11e17ca97c7e1b0e4ab1bb9cbcd8c01af74e9a8427ce0

Observation 57171dc6-a398-4e4b-ab3f-c89d872d27c7 · inbound

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation cites this paper.

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.079222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T05:31:59.676725Z digest=sha256:4e6adebba536b55c2523c06b2fd54520c74d2c6a5f230896954de17e3f2884c3