Pith. sign in

Paper Citation Record · LEDGER

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

As of 14 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2607.08745.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08745 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T02:01:58.009647Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfc42bf9-3ec6-45a5-889c-71e1d2271daa · outbound

This paper cites Lang, et al.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Lang, et al

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.250635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:103756c9e3533be8bddee88fad960282dcc2429f370c1c649adb3fbfa38dac49

Observation 84dceef0-7407-4cb7-ab56-0ba4ba2d5ef8 · outbound

This paper cites Argoverse: 3d tracking and forecasting with rich maps.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Argoverse: 3d tracking and forecasting with rich maps

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.232453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:bea388f4c22dee7aa62e4549500fd811920df93e010f806e1f766c8ea346301f

Observation 72eab2a4-57cf-49eb-84f9-1550dc09715b · outbound

This paper cites Drivingvqa: A dataset for interleaved visual chain-of-thought in real-world driving scenarios.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Drivingvqa: A dataset for interleaved visual chain-of-thought in real-world driving scenarios

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.236499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:93aec053bed89426f36c7454d5f0c571c28e97a036b62e92bb28ba2a87d6bc77

Observation 2313295b-6e32-4099-a1e9-35ef28916285 · outbound

This paper cites The cityscapes dataset for semantic urban scene understand- ing.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding The cityscapes dataset for semantic urban scene understand- ing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.234732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:9fb8afdae61ce9ead7e631136ac8e89334b0ed1dc257fa89432ab612c84442c5

Observation 8ef455ac-2c4a-430e-b50e-76d1b6c44a56 · outbound

This paper cites Vision-based traffic accident detection and anticipation: A survey.IEEE Transactions on Circuits and Systems for Video Technology, 34(4):1983–1999, 2023.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Vision-based traffic accident detection and anticipation: A survey.IEEE Transactions on Circuits and Systems for Video Technology, 34(4):1983–1999, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.238761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:8c5971facf4a952a7a1057f05ce1dce8ddd0a2c219ef4e0f0387eecffcd5189c

Observation 62fc1f11-523c-44c2-941d-3f156605b111 · outbound

This paper cites Are we ready for autonomous driving? the kitti vision benchmark suite.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Are we ready for autonomous driving? the kitti vision benchmark suite

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.244268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:2fb781a573f7c4ddca0508968345bd73ea0797f29cc8870d55be40c95b2a6277

Observation 690730ff-9ec8-4863-96eb-ff3c45598f35 · outbound

This paper cites Computer vision for autonomous vehicles: Prob- lems, datasets and state of the art.Foundations and Trends in Computer Graphics and Vision, 12(1–3):1–308, 2017.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Computer vision for autonomous vehicles: Prob- lems, datasets and state of the art.Foundations and Trends in Computer Graphics and Vision, 12(1–3):1–308, 2017

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.246423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:389ae1ff96ba4d1efaa181f2940e6c714978c64df7294bf8f25675eca341b2c7

Observation d6c2bce7-542b-489d-9e11-bd69cd7452c4 · outbound

This paper cites Nuplanqa: A large-scale dataset and benchmark for multi- view driving scene understanding in multi-modal large lan- guage models.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Nuplanqa: A large-scale dataset and benchmark for multi- view driving scene understanding in multi-modal large lan- guage models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.248714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:166f95d705945737b0cd4086cf296d4810ea4064aa1ac21dd568452ed0f7133c

Observation 52d3d124-b2db-4145-8703-6638226949f2 · outbound

This paper cites Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.252557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:d2d0761a5505e57341bd9e5aba3b3ce13bcbcb6a3a0317f91072c5e676a28cbc

Observation 28a46e76-4422-4f62-ab7e-dc603ce7575a · outbound

This paper cites Scala- bility in perception for autonomous driving: Waymo open dataset.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Scala- bility in perception for autonomous driving: Waymo open dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.240411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:c49b42af75de960f4f5980ddb9cd9f84bd93c070e359cbeb1e2cc36f394c7fd8

Observation 00067a6a-50f3-49fc-835b-5ee56923127b · outbound

This paper cites Embodied scene understanding for vi- sion language models via metavqa.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Embodied scene understanding for vi- sion language models via metavqa

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.242260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:8e0a122d916d0d835b7ee19881d3f958952ebe34509a43a084639ad54346ae22

Observation 7d2d75f2-f885-4192-b229-100c998caab2 · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T02:06:42.238542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T02:01:58.009647Z digest=sha256:da423639d188af8512781156be11100cbb4ef0f41fef31a0903bdfe95f9f5228

Pith citing papers

No inbound Pith citation observations are available.