Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:09.637593Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 8 inbound Pith citation observations for arXiv:2505.14321.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:09.637593Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:55.001360Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:07:28.700018Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4c33aa65-0e2e-42db-9161-a9511e63d799 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50720574-c593-4f33-843e-abb42ec207a6 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Egoschema: A diagnostic benchmark for very long-form video language understanding.NeurIPS, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5c37511-b97e-44fe-8da5-08490f695976 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? NExT-QA: Next phase of question-answering to explaining temporal actions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d333cd39-fecf-4ff0-bd47-6f21b014e69d · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c1e67a-f6d3-49aa-bdf7-7ade136da3ed · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? MLVU: Benchmarking Multi-task Long Video Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1747ad97-fddd-4dfb-b21d-bbf26f516b93 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LVBench: An Extreme Long Video Understanding Benchmark
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3f18659-5bfb-4b0e-a23b-f1fb93941091 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748– 42761, 2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff29bc6e-46c5-44b8-b2d6-8873122790c4 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Palm: Scaling language modeling with pathways.JMLR, 2023
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eedf7de9-9ee6-4b19-a0b1-7a1e74ef8153 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9860e96a-faf4-4a6f-b7e1-5c126b77a85f · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? GPT-4 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05209c6a-ebe1-463a-a8dd-5a27d1eb520d · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Improved Baselines with Visual Instruction Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f978819-19ad-46ba-927d-11dabe00f256 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaV A- NeXT: Improved reasoning, ocr, and world knowledge, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1db472a9-81e2-4a50-8841-9cefcd3006ee · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3ce082d-2d69-45ce-a096-ac5d33a01150 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-ChatGPT: Towards detailed video understanding via large vision and language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d045ce46-38f7-49f5-a98b-3709d6d129a6 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Learning transferable visual models from natural language supervision
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a759a2-488b-4a4f-812a-0eb18daddc67 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaV A-NeXT: A strong zero-shot video understanding model, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae24c0ee-71d4-42d7-b8f4-2ced1bbe7b05 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c74869-a7ea-414b-91da-4bde70336a44 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73db5c87-2b89-4b6e-abd9-9da8866b837f · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15403d8c-9c3f-4dc4-87cf-6aab6ed4073b · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ce74f51-9a08-4df3-8d8f-bffae76d590a · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18674d71-7b64-49eb-a2e6-63ad3c5caffa · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video question answering via gradually refined attention over appearance and motion
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b72c423-7680-4d6e-a3c0-1b7b844ae2b1 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? ActivityNet-QA: A dataset for understanding complex web videos via question answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f4aaf61-ea0e-4a96-bc14-66fe79aab981 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13727959-433c-4784-a7c7-20bf3e898152 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LIME: Less Is More for MLLM Evaluation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a46c902-f8b5-4033-88b7-c49360ce7c9a · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d3ba9d7-f312-453b-8ac2-1aca202aa967 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1591ffc-cd7f-4534-bac7-83080ef6bccc · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d88534-9569-4b16-a312-cc1126394219 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd5c3eea-f837-4aaa-9500-c77784b901eb · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9223dad6-8fbc-4ace-ba0a-386a2e5366f4 · outbound
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18307915-d59b-4e62-944f-aa7c483f1624 · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afbff76-9efc-4df2-9eb1-4cf3663958cc · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af70d98-2788-4b20-85b2-abe90fc7600a · inbound
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b721bc2-2993-4738-92b6-4fd426d49e1c · inbound
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 537c814d-9e7f-4b8f-819c-53b57e0df066 · inbound
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efdfaa89-8f69-4236-a9c6-9184a07e4464 · inbound
An Attribute-Based Measure of Video Complexity Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52eb70fb-bb63-4e0c-b641-84a3ea0666ea · inbound
Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd04825c-b05d-4ba5-9fad-add7deb39880 · inbound
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.