Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:52:02.847623Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.09815.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:52:02.847623Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T06:14:11.109840Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T06:14:18.588892Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 38fd8538-24f6-4d2c-aae3-bd2b897c11cf · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Vru-cipi: Crossing intention prediction at intersections for improving vulnerable road users safety
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a286f007-514b-417a-ba85-c1b0ad9a845a · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-to-text pedestrian monitoring (vtpm): Leveraging large language models for privacy-preserve pedestrian activity monitoring at intersections
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9eec18b-ec51-4a31-9852-9f02a56849cb · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c7d569a-bd47-4b21-9ea6-6e2ada4e1c66 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Vrucrosssafe for crossing intention prediction of vulnerable road users for improving safe crossing at in- tersections
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 900be371-71f9-4061-af6e-6b844eb9e059 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Evaluating the safety impact of mid-block pedes- trian signals (mps)
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d48006a3-202e-439e-8f47-478ccb091f61 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Spice: Semantic propositional image cap- tion evaluation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3416a708-a128-4df1-b6b7-69178b9aaa4e · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d84b7b-0a81-4c40-8868-ba56219f95f6 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Ex- panding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1222fda1-ac94-414e-ba6d-3b64c5e3a3c1 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Analysis of automatic evaluation metric on low-resourced language: Bertscore vs bleu score
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96043b8d-26ac-469c-bf94-cff9df10cbbe · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Dada-2000: Can driving accident be pre- dicted by driver attention? analyzed by a benchmark, 2019
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6e8ee17-cba3-4894-8f62-5c425b77185e · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Cognitive accident prediction in driving scenes: A multimodality benchmark, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4f48d70-7253-4fe7-bebe-0ff6d0042f61 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Abductive ego-view accident video understanding for safe driving per- ception, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b45b3eae-6f09-4bed-a23c-72ea6f3cc09f · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Pedestrian traffic fatalities by state: 2019 preliminary data, 2020
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcbfc582-c3a6-4e7a-a638-26e887c54f5f · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Multi-frame, lightweight & efficient vision-language models for question answering in autonomous driving, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6c2d9de-2839-414b-b9cf-b4ca06f49391 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Cipf: Crossing intention prediction network based on feature fusion modules for improving pedestrian safety
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 437af149-efc2-48e1-91b8-ea87e6358e35 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding An attention-guided multistream feature fusion network for early localization of risky traffic agents in driving videos
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c69289b2-c008-457b-9f99-2ad02ae87c93 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Textual explanations for self-driving ve- hicles
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f6a7a2c-9d2d-4365-af24-9d981b308623 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Pedes- trian crossing direction prediction at intersections for pedes- trian safety
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66c24616-51c5-4880-bfb7-f9bc7114862a · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Meteor: an automatic met- ric for mt evaluation with high levels of correlation with hu- man judgments
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6d39b54-35fc-42bc-884e-64849a02c3e4 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Llava-onevision: Easy visual task transfer,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4889e41-dde1-40b2-84fe-6b05226fc6c1 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1e39bb1-e5e9-4e0d-b595-34030f2e210b · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding ROUGE: A package for automatic evaluation of summaries
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8df008b4-e0f6-46d6-b7fd-79f86aee2def · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Aligning llm with human travel choices: a persona-based embedding learning approach, 2025
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f4e8c86-8a73-4805-9d35-81fd31d78cfe · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Toward llm- agent-based modeling of transportation systems: A concep- tual framework
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb349dfc-58fd-4f56-abd9-84981299b6a4 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-xl-pro: Reconstructive token compression for extremely long video understanding, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9af2e01a-8dec-450a-9ea8-adc7cca15778 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding A simulation-based frame- work for urban traffic accident detection
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75f5519a-8631-45c4-bc38-d279deceaf54 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Dolphins: Multimodal language model for driving
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f356e168-ccf3-40f2-b327-152b3748334d · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Drama: Joint risk localization and captioning in driving
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a135f0c-687e-4e62-9b96-62710e03f9e5 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding LingoQA: Visual Question Answering for Autonomous Driving
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16337822-e145-47c9-aed5-43ffa7e03484 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Gpt-4 technical report, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11413f66-0aef-43f1-b08d-9d71f1e4d99f · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Idd-x: A multi-view dataset for ego-relative important object localization and explanation in dense and unstructured traffic
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b07107c-ab52-4cce-8d92-bd624950b985 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Traffic-Domain Video Question Answering with Automatic Captioning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32d4c2d8-6703-42a8-9cd3-9e62123a26fc · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7daffbfd-aaa1-4595-a611-50f5b666af8f · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Comet: A neural framework for mt evaluation, 2020
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80e23840-b2fd-4f2a-91aa-eb47000d2447 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 336768a4-ab1f-4dac-be11-3dc0bd9dbc10 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Mobile-videogpt: Fast and accurate video understanding language model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dab9472d-2449-4130-acf1-d9b6d776d216 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae7b551c-13a4-44c9-ac95-77ba4b1a8538 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding DriveLM: Driving with Graph Visual Question Answering
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c694c99-57e3-41b4-a3eb-6f139c2149ad · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Gemini: A family of highly capable multi- modal models, 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 685bdf84-233c-4eba-b433-f5fe17aa4fdf · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Qwen2.5-vl, 2025
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82fbeec4-4dcb-4c56-86ba-962cabc0eb05 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Temporal stability of factors af- fecting injury severity in rear-end and non-rear-end crashes: A random parameter approach with heterogeneity in means and variances
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c24b6dd6-2bae-4b08-8e43-92737626ad03 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Effects of speed difference on injury severity of freeway rear-end crashes: Insights from correlated joint random parameters bivariate probit models and temporal instability
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41cfab0a-69d1-4015-bbcb-7c15b58792bb · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ece9ad69-d10b-45f1-805d-21a54e3633bf · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Tunnel crash severity and congestion duration joint evaluation based on cross-stitch networks
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92ee0e55-b875-4625-925e-3d6e5ea59736 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f0adfb-f05a-4258-b9b6-f92db2fe9a35 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving, 2023
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ca26b4d-d584-4b4f-82d4-e1a36294fa85 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events, 2021
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c910f5c-6e54-498e-99b8-df0fbcb643b8 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Explainable object-induced action decision for autonomous vehicles
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6cbc2ca-0415-4034-9640-498b2112f7e1 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Crandall, and Ella M
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d68548f2-2f28-4ccd-9ae7-f12fafcc7a7d · outbound
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e8c4e6c-c1c2-416a-a55b-57ecf3b82c3a · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding same seman- tics, different structure
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abc1d762-fc31-46b7-8a7e-647025476b46 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6342a1e9-37af-4a02-ac42-f6e3fba4d624 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Traffic Accident Bench- mark for Causality Recognition
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b469906-e568-4ae3-baf5-eeb8da683ab3 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Video instruction tuning with synthetic data, 2024
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a29d88a2-c1c7-4c82-9408-08847403189d · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d3019b9-0d0d-45dd-92a4-62edd4999733 · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44828347-caa7-490c-a762-eae894dfb419 · outbound
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39424a1c-65d8-4bea-98d6-86b4f1e89ebf · outbound
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding Unresolved cited work
Reference 684
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0715a46-5816-47e9-9d8a-debf9fbb15c5 · inbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.