Pith. sign in

Paper Citation Record · LEDGER

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding

As of 20 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.09815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09815 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:52:02.847623Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T06:14:11.109840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:14:18.588892Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved13
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38fd8538-24f6-4d2c-aae3-bd2b897c11cf · outbound

This paper cites Vru-cipi: Crossing intention prediction at intersections for improving vulnerable road users safety.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Vru-cipi: Crossing intention prediction at intersections for improving vulnerable road users safety

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.168594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.531647Z digest=sha256:3a1fc90035e1bf6d99a1e8fd76aa11c8fe0dcc8e2ac530e214eb3c5f7ca7819f

Observation a286f007-514b-417a-ba85-c1b0ad9a845a · outbound

This paper cites Video-to-text pedestrian monitoring (vtpm): Leveraging large language models for privacy-preserve pedestrian activity monitoring at intersections.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-to-text pedestrian monitoring (vtpm): Leveraging large language models for privacy-preserve pedestrian activity monitoring at intersections

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.141310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.537855Z digest=sha256:4dc65ebe9086973cfeb2a6d9ac32cd7b3f0bd24b0f3c14078cde960ff1d6aabf

Observation c9eec18b-ec51-4a31-9852-9f02a56849cb · outbound

This paper cites Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.545707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.545707Z digest=sha256:8b74d778550fa376632e197b1d07a165d780a88df578ac9d6c6fb78b98525a89

Observation 4c7d569a-bd47-4b21-9ea6-6e2ada4e1c66 · outbound

This paper cites Vrucrosssafe for crossing intention prediction of vulnerable road users for improving safe crossing at in- tersections.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Vrucrosssafe for crossing intention prediction of vulnerable road users for improving safe crossing at in- tersections

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.119283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.552861Z digest=sha256:4e98d18a5c783de7f8112609c2cbdb491290bcc0833bfa6a3b03359665196c34

Observation 900be371-71f9-4061-af6e-6b844eb9e059 · outbound

This paper cites Evaluating the safety impact of mid-block pedes- trian signals (mps).

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Evaluating the safety impact of mid-block pedes- trian signals (mps)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.096318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.558371Z digest=sha256:6c41a7b00c94af219dca3c4a3f99f947c95fec0a9cd9c3796780d6793a2f9e04

Observation d48006a3-202e-439e-8f47-478ccb091f61 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Spice: Semantic propositional image cap- tion evaluation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.073991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.565180Z digest=sha256:15c69662923c4d869fe1487358ee5e6719c15c27d135c8db54d92b1daf4e53e3

Observation 3416a708-a128-4df1-b6b7-69178b9aaa4e · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.572056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.572056Z digest=sha256:1bcb862d31390471eb60cea2e6b62f85b6b0132c455735d0aac2a3466a29ce9d

Observation d0d84b7b-0a81-4c40-8868-ba56219f95f6 · outbound

This paper cites Ex- panding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Ex- panding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.056688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.581554Z digest=sha256:3a4ce782e0af41b878aa344ace2d3591d200e3a8ec881184d71d4221f05f9033

Observation 1222fda1-ac94-414e-ba6d-3b64c5e3a3c1 · outbound

This paper cites Analysis of automatic evaluation metric on low-resourced language: Bertscore vs bleu score.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Analysis of automatic evaluation metric on low-resourced language: Bertscore vs bleu score

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.030930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.587138Z digest=sha256:4d163b3d803ff0afde3772f9ae64fd09e4033e193554b2172d2f4ca1eba78bd8

Observation 96043b8d-26ac-469c-bf94-cff9df10cbbe · outbound

This paper cites Dada-2000: Can driving accident be pre- dicted by driver attention? analyzed by a benchmark, 2019.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Dada-2000: Can driving accident be pre- dicted by driver attention? analyzed by a benchmark, 2019

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:04.003953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.592310Z digest=sha256:f76826179f0a101abc6e405ebb6a9f57a93f19ba511ee7e0d2d78bd74433e307

Observation a6e8ee17-cba3-4894-8f62-5c425b77185e · outbound

This paper cites Cognitive accident prediction in driving scenes: A multimodality benchmark, 2023.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Cognitive accident prediction in driving scenes: A multimodality benchmark, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.981806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.601084Z digest=sha256:1fdd5a1243fcf2fbc9d4355a0e7ffd5380c457b7275fea213b997bfa8195fdcf

Observation b4f48d70-7253-4fe7-bebe-0ff6d0042f61 · outbound

This paper cites Abductive ego-view accident video understanding for safe driving per- ception, 2024.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Abductive ego-view accident video understanding for safe driving per- ception, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.955901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.606670Z digest=sha256:789f26db39b32692243b3ac71610ce3f7b97936719ada93e770a13f6877459fe

Observation b45b3eae-6f09-4bed-a23c-72ea6f3cc09f · outbound

This paper cites Pedestrian traffic fatalities by state: 2019 preliminary data, 2020.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Pedestrian traffic fatalities by state: 2019 preliminary data, 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.931712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.614469Z digest=sha256:2d21fad42c182c28a7d563773c516b0704337044fcd195ac1ff4576b240f3d38

Observation bcbfc582-c3a6-4e7a-a638-26e887c54f5f · outbound

This paper cites Multi-frame, lightweight & efficient vision-language models for question answering in autonomous driving, 2024.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Multi-frame, lightweight & efficient vision-language models for question answering in autonomous driving, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.907323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.619633Z digest=sha256:fb3cc077bf83966f86638702c146c2e91a68b5f86dc4a93d3814df3b9a6f6f6e

Observation c6c2d9de-2839-414b-b9cf-b4ca06f49391 · outbound

This paper cites Cipf: Crossing intention prediction network based on feature fusion modules for improving pedestrian safety.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Cipf: Crossing intention prediction network based on feature fusion modules for improving pedestrian safety

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.876626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.626061Z digest=sha256:9d72adf8d604940a3822f24c9e795c018ecd791a64609e7923a77c0225fedc72

Observation 437af149-efc2-48e1-91b8-ea87e6358e35 · outbound

This paper cites An attention-guided multistream feature fusion network for early localization of risky traffic agents in driving videos.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding An attention-guided multistream feature fusion network for early localization of risky traffic agents in driving videos

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.846650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.630519Z digest=sha256:621bcb9c7c665a77598c7298969d3a24763b6563fa70740aa3453114cceb937f

Observation c69289b2-c008-457b-9f99-2ad02ae87c93 · outbound

This paper cites Textual explanations for self-driving ve- hicles.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Textual explanations for self-driving ve- hicles

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.819327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.635560Z digest=sha256:d2dd6e88272265ee337b9592939dd890074b25e66e27f324d5259d895f6830cf

Observation 8f6a7a2c-9d2d-4365-af24-9d981b308623 · outbound

This paper cites Pedes- trian crossing direction prediction at intersections for pedes- trian safety.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Pedes- trian crossing direction prediction at intersections for pedes- trian safety

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.797623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.641083Z digest=sha256:d158a79a7c4a53c4b4f0fcc44ccf2bdd718d316bc25323cacbec746c84ff8b5d

Observation 66c24616-51c5-4880-bfb7-f9bc7114862a · outbound

This paper cites Meteor: an automatic met- ric for mt evaluation with high levels of correlation with hu- man judgments.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Meteor: an automatic met- ric for mt evaluation with high levels of correlation with hu- man judgments

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.779038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.645935Z digest=sha256:f99a9a85062a81155f1c6122913d24b089b541d8ba951796036787f70b160bc8

Observation e6d39b54-35fc-42bc-884e-64849a02c3e4 · outbound

This paper cites Llava-onevision: Easy visual task transfer,.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Llava-onevision: Easy visual task transfer,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.650906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.650906Z digest=sha256:07e3a30897869abe79b88c2a06cb9595447b4eecfbb2b32037cc0bee3fd64010

Observation b4889e41-dde1-40b2-84fe-6b05226fc6c1 · outbound

This paper cites Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.744812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.655775Z digest=sha256:256fb22cbdd7f4724b6bc8b80ee02eeb8b613532f25dfd50a52cdfc2faf489e2

Observation a1e39bb1-e5e9-4e0d-b595-34030f2e210b · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding ROUGE: A package for automatic evaluation of summaries

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.725604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.660510Z digest=sha256:ab23eb37f54d6189239c78ead8e0e1d5979405fabd814cda4b1e23711bd4c187

Observation 8df008b4-e0f6-46d6-b7fd-79f86aee2def · outbound

This paper cites Aligning llm with human travel choices: a persona-based embedding learning approach, 2025.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Aligning llm with human travel choices: a persona-based embedding learning approach, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.706129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.664906Z digest=sha256:d59ce96e4cf419e263a55bd62b064e64c463962bced6db4df734a3907dd9b2fb

Observation 4f4e8c86-8a73-4805-9d35-81fd31d78cfe · outbound

This paper cites Toward llm- agent-based modeling of transportation systems: A concep- tual framework.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Toward llm- agent-based modeling of transportation systems: A concep- tual framework

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.681997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.669408Z digest=sha256:8ddbdb5f9e35fea5b1432574d7773c90bab8dd86c161c03a3d106dda5f376f61

Observation cb349dfc-58fd-4f56-abd9-84981299b6a4 · outbound

This paper cites Video-xl-pro: Reconstructive token compression for extremely long video understanding, 2025.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-xl-pro: Reconstructive token compression for extremely long video understanding, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.657476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.673938Z digest=sha256:d5c79e0905aff44ffc852f722b9a6c9073f855f9dd68db4e235e617731861d60

Observation 9af2e01a-8dec-450a-9ea8-adc7cca15778 · outbound

This paper cites A simulation-based frame- work for urban traffic accident detection.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding A simulation-based frame- work for urban traffic accident detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.637337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.678862Z digest=sha256:b7984cc5b8dc75a09932fdb5eb8c3c1c54013bae2fb512fb0c307daac2c6fb9a

Observation 75f5519a-8631-45c4-bc38-d279deceaf54 · outbound

This paper cites Dolphins: Multimodal language model for driving.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Dolphins: Multimodal language model for driving

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.619441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.684273Z digest=sha256:5c6c0fde25c9071c75b6f251d66c77e8c7c12fffafdf1663f535531b089f084d

Observation f356e168-ccf3-40f2-b327-152b3748334d · outbound

This paper cites Drama: Joint risk localization and captioning in driving.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Drama: Joint risk localization and captioning in driving

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.690495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.690495Z digest=sha256:2919f0ead4382068359941d0d0febb882a8c852d7db7f9b60440adb41b86ded9

Observation 3a135f0c-687e-4e62-9b96-62710e03f9e5 · outbound

This paper cites LingoQA: Visual Question Answering for Autonomous Driving.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding LingoQA: Visual Question Answering for Autonomous Driving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.695281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.695281Z digest=sha256:7ede7cc25448ea238b698e1a211a34d6f4c1f94012ad19a2df888aaac920471c

Observation 16337822-e145-47c9-aed5-43ffa7e03484 · outbound

This paper cites Gpt-4 technical report, 2024.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Gpt-4 technical report, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.588197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.700318Z digest=sha256:90b674b2c9c175a76bc27e21799cec6af0fdcbd3d7248ba37f87b03e6aed09a1

Observation 11413f66-0aef-43f1-b08d-9d71f1e4d99f · outbound

This paper cites Idd-x: A multi-view dataset for ego-relative important object localization and explanation in dense and unstructured traffic.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Idd-x: A multi-view dataset for ego-relative important object localization and explanation in dense and unstructured traffic

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.572315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.706429Z digest=sha256:4286436f41a982638dec13170eb774f89e3c51b32e055fc20f917b3e0865472d

Observation 3b07107c-ab52-4cce-8d92-bd624950b985 · outbound

This paper cites Traffic-Domain Video Question Answering with Automatic Captioning.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Traffic-Domain Video Question Answering with Automatic Captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.711142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.711142Z digest=sha256:e0a504b9c418b597e022129e54af494adcfb39f6353dd35ac991802b5730aaa1

Observation 32d4c2d8-6703-42a8-9cd3-9e62123a26fc · outbound

This paper cites NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.716601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.716601Z digest=sha256:22b7dd7bd00b98eaa9c8455c91d430b237878fceb3542b40c03f7f59cc7136ea

Observation 7daffbfd-aaa1-4595-a611-50f5b666af8f · outbound

This paper cites Comet: A neural framework for mt evaluation, 2020.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Comet: A neural framework for mt evaluation, 2020

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.554026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.721525Z digest=sha256:efceea20cd2f09bebe4d782d272a8c79e360b91a4cb076d38dce1e7eaf83ee38

Observation 80e23840-b2fd-4f2a-91aa-eb47000d2447 · outbound

This paper cites Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.538447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.728469Z digest=sha256:579d85970656dcf8eace5383ff5d3c56002e91de6d1c89aeffa1fbf453bcf092

Observation 336768a4-ab1f-4dac-be11-3dc0bd9dbc10 · outbound

This paper cites Mobile-videogpt: Fast and accurate video understanding language model.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Mobile-videogpt: Fast and accurate video understanding language model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.522934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.733454Z digest=sha256:9e0d8f66445dfcb17986d9cd2d2e674877d47973990498388f05d91fd1b335a8

Observation dab9472d-2449-4130-acf1-d9b6d776d216 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.738153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.738153Z digest=sha256:6afe78df0b177697d31b1c43e96bb680d6abf8263262d4323da1cd68847dcd4b

Observation ae7b551c-13a4-44c9-ac95-77ba4b1a8538 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding DriveLM: Driving with Graph Visual Question Answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.743420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.743420Z digest=sha256:38b5a52728efa5ce2be33853cbd40172ee6ec678a66eddde71086295d9f27195

Observation 7c694c99-57e3-41b4-a3eb-6f139c2149ad · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2025.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Gemini: A family of highly capable multi- modal models, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.508237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.749174Z digest=sha256:14dab257a5d6400eb366f388b9a0f3a0342d089707ff86b3aba5ada5cbc7417f

Observation 685bdf84-233c-4eba-b433-f5fe17aa4fdf · outbound

This paper cites Qwen2.5-vl, 2025.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Qwen2.5-vl, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.492546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.754171Z digest=sha256:4551192c854b970ea46c3474cc6d977b32aca41d774b0ea72c9341a353a6376a

Observation 82fbeec4-4dcb-4c56-86ba-962cabc0eb05 · outbound

This paper cites Temporal stability of factors af- fecting injury severity in rear-end and non-rear-end crashes: A random parameter approach with heterogeneity in means and variances.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Temporal stability of factors af- fecting injury severity in rear-end and non-rear-end crashes: A random parameter approach with heterogeneity in means and variances

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.339225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.758892Z digest=sha256:dacde4d66ca7dc9fe99edcad9d3321ad1c721e4aafcc657b143417629665fd78

Observation c24b6dd6-2bae-4b08-8e43-92737626ad03 · outbound

This paper cites Effects of speed difference on injury severity of freeway rear-end crashes: Insights from correlated joint random parameters bivariate probit models and temporal instability.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Effects of speed difference on injury severity of freeway rear-end crashes: Insights from correlated joint random parameters bivariate probit models and temporal instability

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.323571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.763278Z digest=sha256:8a1ae46a2a2e0c3e94dde756042f24054378336c61a4def965fb548684f5981c

Observation 41cfab0a-69d1-4015-bbcb-7c15b58792bb · outbound

This paper cites an unresolved cited work.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.306708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.767659Z digest=sha256:b7ed5cb8ccd9d341e4d05e08dc95e4543d6a174f19aef7ec058821fa35b96b8b

Observation ece9ad69-d10b-45f1-805d-21a54e3633bf · outbound

This paper cites Tunnel crash severity and congestion duration joint evaluation based on cross-stitch networks.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Tunnel crash severity and congestion duration joint evaluation based on cross-stitch networks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.290523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.772368Z digest=sha256:71554744e6c664f43fffebfc1aae4d9197eaebbbec8110e5edf08f7a9fd668b4

Observation 92ee0e55-b875-4625-925e-3d6e5ea59736 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.777498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.777498Z digest=sha256:37201af6d09fbd7dd6ba6bf1b498a5e638450946e519b2bdb98f822ac51edc12

Observation a5f0adfb-f05a-4258-b9b6-f92db2fe9a35 · outbound

This paper cites Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving, 2023.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.274491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.785623Z digest=sha256:95e2aa6993cb2ec8e346ea49ce62882537b5024ce28570ef0717cbdfc7196889

Observation 9ca26b4d-d584-4b4f-82d4-e1a36294fa85 · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events, 2021.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.255075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.790612Z digest=sha256:8b8e29326a08ca01f62f4e1c1afb1d193b664594d2834abc8dd2e833b064e9a1

Observation 0c910f5c-6e54-498e-99b8-df0fbcb643b8 · outbound

This paper cites Explainable object-induced action decision for autonomous vehicles.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Explainable object-induced action decision for autonomous vehicles

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.235014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.796026Z digest=sha256:46b2a1ddd3bdfd8726866e8b4d071a4eac701d700d318d37e2e776ccb8bfdb6c

Observation d6cbc2ca-0415-4034-9640-498b2112f7e1 · outbound

This paper cites Crandall, and Ella M.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Crandall, and Ella M

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.219329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.801047Z digest=sha256:de9f576b82e6f94589539c378dc3f598a0a478a813be4560ec33951e45745d65

Observation d68548f2-2f28-4ccd-9ae7-f12fafcc7a7d · outbound

This paper cites Crandall.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Crandall

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.202439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.806325Z digest=sha256:f64b15a20e8715dd7d78cc9b5c78bcad82cbca9004350c38c0f8be63337fd730

Observation 6e8c4e6c-c1c2-416a-a55b-57ecf3b82c3a · outbound

This paper cites same seman- tics, different structure.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding same seman- tics, different structure

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.187076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.811518Z digest=sha256:46aeed6c4ee0dded9039eadfc66a09cfc32a955e4a6aea41f3f998abcaaca714

Observation abc1d762-fc31-46b7-8a7e-647025476b46 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:02.823886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:02.823886Z digest=sha256:a209b51251de4a7e7c0db3556cff647209068b4be22298ee4d04cdd73e69cb88

Observation 6342a1e9-37af-4a02-ac42-f6e3fba4d624 · outbound

This paper cites Traffic Accident Bench- mark for Causality Recognition.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Traffic Accident Bench- mark for Causality Recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.147044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.828834Z digest=sha256:33f18ec4a05abffd411963d4627e22e6025834ca016ab865c08f1eb7aef969f2

Observation 1b469906-e568-4ae3-baf5-eeb8da683ab3 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video instruction tuning with synthetic data, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.126898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.833604Z digest=sha256:7e1f1c457ed840e73dbfecb42fba62db21f043aaf5f24b8aaeca71a73c68c354

Observation a29d88a2-c1c7-4c82-9408-08847403189d · outbound

This paper cites an unresolved cited work.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.106328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.838046Z digest=sha256:5de5fac400c7d8bef2a92fe4224b6f8aedf84d0b26af4a6337bd0cfd7a055d17

Observation 7d3019b9-0d0d-45dd-92a4-62edd4999733 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.086374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.842449Z digest=sha256:3f698c7b4c260a719ae7ba50e5aa88d547df804271a3c01b8efc4e48afb13b1e

Observation 44828347-caa7-490c-a762-eae894dfb419 · outbound

This paper cites sunny day.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding sunny day

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.063471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.847623Z digest=sha256:8b42113473b9e22125680d07c34309c76161b2435876d04dfbb98add2040dafc

Observation 39424a1c-65d8-4bea-98d6-86b4f1e89ebf · outbound

This paper cites an unresolved cited work.

VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Unresolved cited work

Reference 684

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T17:52:03.166377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:52:02.818786Z digest=sha256:848aef6f7eaab243fb402b8caf7d07c857d575674d596ea37ffcb6fc86b1250e

Pith citing papers

Observation a0715a46-5816-47e9-9d8a-debf9fbb15c5 · inbound

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA cites this paper.

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:18.592347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:14:11.109840Z digest=sha256:a1656cfd76a613a574a389d13770cbe763a5fe5ea7c14daadbcb091727bfedee