Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2406.14515.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.815801Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T23:12:46.636950Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 38f029b6-d9a7-4cfe-a8a0-5b6ca001afd8 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b8e1403e-d56f-470d-9a35-365fd639fabf · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d7487486-8b20-42e7-96d7-c58ac5d1533d · inbound
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06089e05-5ea8-49b4-8d4e-daea623a9032 · inbound
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4429b2ad-bd1d-413d-a6a4-054da3e47464 · inbound
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4550844a-d677-4d83-8c5f-d04588f1f544 · inbound
LinVT: Empower Your Image-level Large Language Model to Understand Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bbc87ab-4cc4-4e82-851a-e02b4159ce21 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 95f27273-2fd1-4cab-9b94-5590a755724b · inbound
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0475e5cf-6afb-480f-b64b-86e9bec5031a · inbound
Neptune: The Long Orbit to Benchmarking Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c530d3e2-5ced-4dde-9fcb-81f40a943c60 · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 216eef99-e433-4e15-ba2d-d43f10802772 · inbound
VCA: Video Curious Agent for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4427d20-f72d-465b-945a-e268ff6e010c · inbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f54430-2512-4e8a-a767-9125548171aa · inbound
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation be82f068-0bf9-414b-8437-76fa6d796041 · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 59ef57de-f721-4e37-b3cd-80215076d60d · inbound
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcd3f53a-9953-4561-8386-c85bc699baec · inbound
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ec1e6ef3-fe23-4b98-b54c-6247ef07148a · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fd4d1967-f873-4294-9409-954ecfc5c1bf · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ccd52aed-5402-4615-9018-e7e59f80c64b · inbound
A Benchmark for Crime Surveillance Video Analysis with Large Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3129f62a-b5cf-49d4-8aaa-2cf1e10f6c44 · inbound
Qwen2.5-VL Technical Report MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a6b6f22f-c849-40be-aee7-4c23dfdc372a · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bab3da26-6402-45d4-aecf-bf17ed42d31b · inbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade2840b-5cb6-4e3e-9f89-dc08e6d45abd · inbound
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a386e1c8-8df6-49f4-a812-084e235ae6a5 · inbound
VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69cae93-61b1-4404-b6c9-e5551be0bc90 · inbound
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5bb7880-9522-4bdf-bebb-2110cf0c0af6 · inbound
R^3-VQA: "Read the Room" by Video Social Reasoning MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b05d6f78-5e5f-4c83-9fe6-b3e44bc82297 · inbound
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b773435-f45d-44d3-a2a5-28ef34363b23 · inbound
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a93478-79eb-43dc-9b5d-4b9eed8e11e0 · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a3c23f-0848-4e87-9a93-13c1267286ba · inbound
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b04f4a6-1cdd-49c6-a73d-34ebafa03d10 · inbound
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419a10bf-2fc6-4e6e-a497-af875e6cbe22 · inbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ab95b4-f970-4028-bb65-c12a90eb0e61 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7943e501-1c3d-4caa-adb1-bb1002222058 · inbound
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59fb7299-bb2e-4356-abf1-6aca0a378afb · inbound
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 08fc3bc3-eab0-45af-a13f-60e86cd4ff4e · inbound
Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.