Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T06:14:11.109840Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2606.30220.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T06:14:11.109840Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T10:15:13.957224Z
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1e364f0f-4414-484a-8593-95681bd6caf2 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation daa1d942-08ea-4bb0-810c-d6dd88364ec3 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Aofei Chang, Le Huang, Parminder Bhatia, Taha Kass-Hout, Fenglong Ma, and Cao Xiao
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5fc94880-bb2f-4090-b09a-692695eff023 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1d7aac15-7aa4-4732-8fc6-f2d2ddde606e · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA train on the test set
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83c8b099-376e-4940-80a7-e7f2feba3c14 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c3ab84cc-4b58-4d19-a6bd-b6b31f6be870 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Words or vision: Do vision-language models have blind faith in text? In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e2e139-f95f-4e12-9a41-46bbe689c212 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Freeman, Frédo Durand, Eli Shechtman, and Xun Huang
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8eb521de-6f96-46a2-b234-cb41c556b7b4 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed81f1fa-7da7-435f-9647-6c9504e00fd6 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Do Vision Language Models Need to Process Image Tokens?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 68879b39-a139-4fad-a962-e2e03b9870cd · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Accidentbench: Benchmarking multimodal understanding and reasoning in vehicle accidents and beyond
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a0715a46-5816-47e9-9d8a-debf9fbb15c5 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 468a8dbc-ae18-4519-a684-ffc81b210f9d · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 81997d68-dc64-4062-a352-15a48ce81e85 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA LLaVA-OneVision: Easy Visual Task Transfer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bae40ca8-26ad-4ca5-92d8-138f00f4da29 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Understanding language prior of lvlms by contrasting chain-of-embedding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 365b3d41-34c0-4c51-bf09-e04b99c0b6c3 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d595c44-01a6-4ad5-bd4e-a7b700ece113 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e61ca1cd-6bcb-4d80-80f7-7efbe79f15d5 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28ac769d-8bc6-4dff-b4c4-b4d879a81dd3 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef53ed70-ff35-49b9-ac16-710152423cf0 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Beyond accuracy: Evaluating visual grounding in multimodal medical reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bf63b4d9-f67e-4861-b864-62478fbce937 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Watch Before You Answer: Learning from Visually Grounded Post-Training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b1a90ab7-aeb1-4544-b5e2-f6ecaa9d7cee · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA MLVU: Benchmarking Multi-task Long Video Understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 242a7a44-f182-4792-8574-ec9c43acc111 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Dataset Videos Avg
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b2e5574-9611-4483-aefc-3b6e98ab426b · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2800445-960d-429d-99a4-d755904df3dc · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d99861-7fcb-4e90-a1e9-4d1d78276ab7 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9db1599-f84a-4884-942b-65b412814a24 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA In multimodal systems this manifests asmodality collapse, where language dominates prediction regardless of visual input (Sim et al., 2025; Deng et al., 2025)
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f20afc-500f-476b-88b3-3e6da75e40cc · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Existing auditing approaches either apply binary filtering by removing questions a blind model answers correctly (Asadi et al., 2026; Zhang et al.,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94cb3ce-3bcb-4be5-a8b0-c016635fa455 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Binary filtering discards visually ambiguous but grounded questions; dataset-level metrics cannot identify which individual questions are shortcut-prone
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc854ff6-84d6-4973-b38b-1027dca88007 · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Videos can reach approximately 16 minutes, posing a practical challenge for models under a fixed frame budget
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07379e89-a66b-4d84-8454-3d7042e13d9b · outbound
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA A substantial share of what these benchmarks measure is answerable without any video
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9a72ae-3a9b-467b-b250-5d8640e3197a · inbound
Visual Credit Audit for Multimodal Spatial Reasoning From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08613a70-6534-43ac-805b-5d14fc85d25f · inbound
Visual Credit Audit for Multimodal Spatial Reasoning From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.