Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.02615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02615 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:23:14.558335Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0cbb2dd-c1e8-4c7c-8218-59c0e5f81b4c · outbound

This paper cites Spherical transformer for lidar-based 3d recognition,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Spherical transformer for lidar-based 3d recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:18.242083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:12.510042Z digest=sha256:6c576371589cf1e0b83989339c2afb1bb55eaa75e6c29ad642080adf97b29f74

Observation db2325a5-b133-499e-b29b-043b829e37ba · outbound

This paper cites Rea- son2drive: Towards interpretable and chain-based reasoning for au- tonomous driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Rea- son2drive: Towards interpretable and chain-based reasoning for au- tonomous driving,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.994358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:12.552518Z digest=sha256:0c6ddc78a546cf97f96d89ee04f9e6494491db15844fdb560b93033e306049ab

Observation ebd7634f-7846-439d-a3ec-c488704bf2aa · outbound

This paper cites GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:12.643347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:12.643347Z digest=sha256:07a28df3aa1b5d9b2ba28d70f84cbc78518807e62a4a61e1e2614d5ecba79f50

Observation b21e0480-4d1f-4f23-b446-a5dcb535547a · outbound

This paper cites DriveVLM: The convergence of autonomous driving and large vision-language models,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models DriveVLM: The convergence of autonomous driving and large vision-language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.756695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:12.700488Z digest=sha256:da04e9c064b7ab9564003aaee85d701cec297974cebde4a14f6f1c50482edc4c

Observation a2ca9fb8-e7f4-4487-acec-05afb4933ac4 · outbound

This paper cites BLIP: Bootstrapping language- image pre-training for unified vision-language understanding and generation,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models BLIP: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.438610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:12.731307Z digest=sha256:d0198de431d1c4214794b6337d6c23212825bf3f874a170a821e6317f9991362

Observation 6963edf6-5390-4462-8e3d-3c80c5aa53aa · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:17.195480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:12.811578Z digest=sha256:9ad8aeb8a0909f24d01348a61f1933de77841b8a5c12e05e76849e1c18c46ca6

Observation 597f786b-fe0c-447d-9e0e-11ca408ee3c9 · outbound

This paper cites Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser, “Openscene: 3d scene understanding with open vocab- ularies,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser, “Openscene: 3d scene understanding with open vocab- ularies,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.935102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:12.896521Z digest=sha256:4dac1109b987eb2ad429fdc7bab581d2b7794a0bda72a73e4e65756febaaeaf5

Observation e316f44c-cf65-42b6-8a72-d5771e19f43d · outbound

This paper cites Pla: Language-driven open-vocabulary 3d scene understanding,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Pla: Language-driven open-vocabulary 3d scene understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.597432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:13.020546Z digest=sha256:5d6aeb13cff0afd1a36c4dc5ca44b602a43c8e2ac87a301e0187f1527f6bfc0a

Observation 8e082adb-b2e0-4caf-8991-74438cb51821 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:13.119412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:13.119412Z digest=sha256:9c690b415310d946383f12002e73cd4504696688e50a9918d34342d03922c5b4

Observation 0c7d06ac-d4f1-4a68-9d5a-24fd23cf89eb · outbound

This paper cites The traffic scene un- derstanding and prediction based on image captioning,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models The traffic scene un- derstanding and prediction based on image captioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.384009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:13.240502Z digest=sha256:9611ae69b82ec70be1a1653a9d7934580eee2d3f5f3c6b9a16a0fba990b181e7

Observation 101ebe9c-9183-4f49-82aa-f422d7fb56a6 · outbound

This paper cites Delving into clip latent space for video anomaly recognition,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Delving into clip latent space for video anomaly recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.108961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:13.339597Z digest=sha256:0b7139ac369742f0d0468fd168390355b1def6e0f96525a8b37f67a9ff600f53

Observation 212ca42e-a757-42f2-ac76-e26492b69d59 · outbound

This paper cites Learning to prompt for vision-language models,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Learning to prompt for vision-language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:13.482123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:13.482123Z digest=sha256:982e3ebab6992af518acfb547ed7910f52041002fa061a223a13fc1f86c4a31a

Observation 08761c9a-b025-4ba5-8f76-416a3fc6726a · outbound

This paper cites LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:13.595461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:13.595461Z digest=sha256:4c6fd177ebeb435e36cf925bb9b6eaa1a553170df3aab669790f73502ae86c0d

Observation 43ab1c19-b7c8-48aa-aa60-4a48018b4346 · outbound

This paper cites Vectornet: Encoding hd maps and agent dynamics from vectorized representation,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Vectornet: Encoding hd maps and agent dynamics from vectorized representation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:16.011931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:13.727279Z digest=sha256:9a8ad1a2dc440e2905d09878b81171b3d1040dae9439ee37dbfbb777ee286ef7

Observation bfe27062-e1d8-437a-a555-578abb92c5fb · outbound

This paper cites Reimagining an autonomous vehicle.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Reimagining an autonomous vehicle

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:23:14.984276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:13.817011Z digest=sha256:bd7e2545243a415934b9ca1a15417233d19801ca02d9181862b97b2b2fabcdaa

Observation a0a3c807-ec43-42a4-8d44-2a30e46a5e91 · outbound

This paper cites Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.867593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:13.891608Z digest=sha256:6eaeb882bce6f0f4c1a82fb2336620ab5fe30af2e75970912cca11e5fb529042

Observation 26ff000f-8917-4c67-bca6-31f7d4ad6930 · outbound

This paper cites From automation to autonomy and autonomous vehicles: Challenges and opportunities for human-computer interaction,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models From automation to autonomy and autonomous vehicles: Challenges and opportunities for human-computer interaction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.758020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:13.955573Z digest=sha256:ca5623ef957cca6dda54bbe972dd51e2976666414c3f6e6ef979345d1897a1db

Observation fd326759-a542-4496-b394-370e4317e8fb · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.026613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.026613Z digest=sha256:38aa4c0b69c5355368ffdad6a42f79c379a8213327d6922bcf449c46fd903be6

Observation 009525c1-7108-4b0f-841b-2898fddf9f13 · outbound

This paper cites Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.644833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:14.068576Z digest=sha256:b9944c9e942d96596450206290ea8651462867f12693fd4b9e9b813a65071b28

Observation 439f4f6a-e764-44cf-bd1c-228f1e3b279a · outbound

This paper cites DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.113366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.113366Z digest=sha256:b1678e8b7f828c5cb85252ce6d22c0933c1711baebacea0abd5c22e4a75d97ea

Observation c9dcfaf7-e65c-424a-92cf-24e977e57cb6 · outbound

This paper cites Driving with llms: Fusing object- level vector modality for explainable autonomous driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Driving with llms: Fusing object- level vector modality for explainable autonomous driving,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.445095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:14.182507Z digest=sha256:14f96817479332f946a71157510745efad223c4c2f46ddf164a2f5c599fef296

Observation a06e06f1-6544-4233-b29f-5e7294f6bfdf · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models GPT-Driver: Learning to Drive with GPT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.218939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.218939Z digest=sha256:acd3216d39468b4d5b70469f380448fcaeea000c9b9c73b980f68aa1ee945995

Observation 9bc840b0-96eb-4faf-bba9-7f50d9c1cd97 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.317834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.317834Z digest=sha256:c8541eb2f64bf500d6b71595df8fd47f693a04beccc6e172714e1daf6654d74e

Observation f106eb8c-399d-4736-be2e-3d6bc7a9dd8d · outbound

This paper cites RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Lan- guage Model,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Lan- guage Model,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.375874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.375874Z digest=sha256:921d95432d7d407347b41eecaffb88689ff070492de6162d7917d3f822fbd992

Observation 9962cfb5-f5b4-4242-9d78-35efb8d775fc · outbound

This paper cites DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.454649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.454649Z digest=sha256:a7b5d4636a43c08b160c8c422fe51cd652b35e8d17c516e9fe98d8fc7f8e82ad

Observation ecc36b9e-663c-4f8c-9e43-53724ef1cebb · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Lmdrive: Closed-loop end-to-end driving with large language models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:23:14.503544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:23:14.503544Z digest=sha256:93516a2a61d8ff7accee2367a1f136aec4f86df049787d45868c73480c4e7a3f

Observation 0de58213-75e5-4885-9290-f651cc754242 · outbound

This paper cites Lingoqa: Visual question answering for au- tonomous driving,.

Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models Lingoqa: Visual question answering for au- tonomous driving,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:23:15.227697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:23:14.558335Z digest=sha256:0b530339693d2e84fba17942e943a84ee9098f4064a33f9887ce335c88f3f72f

Pith citing papers

No inbound Pith citation observations are available.