Pith. sign in

Paper Citation Record · LEDGER

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 8 inbound Pith citation observations for arXiv:2505.14321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14321 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:09.637593Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:55.001360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:28.700018Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c33aa65-0e2e-42db-9161-a9511e63d799 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:06.956393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:06.956393Z digest=sha256:746b204ba84e0fbaf6f32cdb89e4bfb13024b85a95e22f475371b934e107d32f

Observation 50720574-c593-4f33-843e-abb42ec207a6 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.NeurIPS, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Egoschema: A diagnostic benchmark for very long-form video language understanding.NeurIPS, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:12.356042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:07.010477Z digest=sha256:d02cf42f132013d4368b6e1b72d8880d9389ccf7f2410f21322776a555a32873

Observation e5c37511-b97e-44fe-8da5-08490f695976 · outbound

This paper cites NExT-QA: Next phase of question-answering to explaining temporal actions.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? NExT-QA: Next phase of question-answering to explaining temporal actions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:12.088606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:07.083589Z digest=sha256:f0e5e41570fc0f5535e9d5048096f3bba8863bb9d755b2a1075972d33e317344

Observation d333cd39-fecf-4ff0-bd47-6f21b014e69d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.210366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.210366Z digest=sha256:d552c179d95bd55650f71bd8628ebea0f0cbc1e69ab073df34b89d4974ef98bd

Observation 37c1e67a-f6d3-49aa-bdf7-7ade136da3ed · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? MLVU: Benchmarking Multi-task Long Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.306705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.306705Z digest=sha256:e69bafcb6c7ff6ec4a04c0bdbf19568fe395e73897a3b18a9ac709e83e5f8403

Observation 1747ad97-fddd-4dfb-b21d-bbf26f516b93 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LVBench: An Extreme Long Video Understanding Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.407582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.407582Z digest=sha256:1f918c8920992549ed424999e919c3574f93fb2b2b864d89c3b29920b57b7d21

Observation f3f18659-5bfb-4b0e-a23b-f1fb93941091 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748– 42761, 2023.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Perception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748– 42761, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.817503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:07.469778Z digest=sha256:30cfe520f862e58c15489025bb108551ead4ef09489d628bf3fb07742a496dae

Observation ff29bc6e-46c5-44b8-b2d6-8873122790c4 · outbound

This paper cites Palm: Scaling language modeling with pathways.JMLR, 2023.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Palm: Scaling language modeling with pathways.JMLR, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.568082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:07.584674Z digest=sha256:da300ede2a4a52bbfc2bdc84ae1a436822318cd88bf3c3b6646de2ab00fbe3c1

Observation eedf7de9-9ee6-4b19-a0b1-7a1e74ef8153 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.658408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.658408Z digest=sha256:8932ae733cc387b6a3a364f56444c2de5c793831adbc68894933cd4d6f43a5e0

Observation 9860e96a-faf4-4a6f-b7e1-5c126b77a85f · outbound

This paper cites GPT-4 Technical Report.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.716508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.716508Z digest=sha256:72397a94bfb92954b2f6e4a9435487a58e2fa11abd1671df5270f37da75e4c09

Observation 05209c6a-ebe1-463a-a8dd-5a27d1eb520d · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Improved Baselines with Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:07.780818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:07.780818Z digest=sha256:5d7ab2cc1b93aacfcb4f18ecd1a0deffc5c175d449bf6f251cc98d85c824e211

Observation 1f978819-19ad-46ba-927d-11dabe00f256 · outbound

This paper cites LLaV A- NeXT: Improved reasoning, ocr, and world knowledge, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaV A- NeXT: Improved reasoning, ocr, and world knowledge, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.296636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:07.876391Z digest=sha256:f4d66749f280b2ffe004b58f2cfba5e548b5f6da56e2e8cf39b397f02aa26c08

Observation 1db472a9-81e2-4a50-8841-9cefcd3006ee · outbound

This paper cites Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:41:09.987240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:07.984130Z digest=sha256:796b288097c7a333500e53b2d6b90da9b7e7ec8a7ab96d49ef992b6a067e59e0

Observation d3ce082d-2d69-45ce-a096-ac5d33a01150 · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-ChatGPT: Towards detailed video understanding via large vision and language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:11.091970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:08.081997Z digest=sha256:617706d0c71a52254d3f58fd9f3fed5a9e45aae3773a53f11809395c219cfee5

Observation d045ce46-38f7-49f5-a98b-3709d6d129a6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Learning transferable visual models from natural language supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.147677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.147677Z digest=sha256:d78e7ca4cb14157121effd31ba9f2e585048b3e29d73a2bb772694b286462eaf

Observation 77a759a2-488b-4a4f-812a-0eb18daddc67 · outbound

This paper cites LLaV A-NeXT: A strong zero-shot video understanding model, 2024.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaV A-NeXT: A strong zero-shot video understanding model, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.817989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:08.205977Z digest=sha256:cf045fe25684259205111a647554d2e2b70b7d55c039d2feff29b02f1ad6aa83

Observation ae24c0ee-71d4-42d7-b8f4-2ced1bbe7b05 · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.331967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.331967Z digest=sha256:55aa5cb238372632091dcdd21a7b358dc3487970d6b22dd40d416d358c09234b

Observation 33c74869-a7ea-414b-91da-4bde70336a44 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.446932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.446932Z digest=sha256:bbfe85b877bb47646a707db3e1f66a455dff6607eb3328eff4f17c8477585a13

Observation 73db5c87-2b89-4b6e-abd9-9da8866b837f · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.568366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.568366Z digest=sha256:368fd70f7c62b377a8a21ec1c8f0d7eba4e4e5511d31b61a29fd60389c0aebf4

Observation 15403d8c-9c3f-4dc4-87cf-6aab6ed4073b · outbound

This paper cites Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.631715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:08.690765Z digest=sha256:8e4a6fc2287c79dcae7ba04780dce97c99bf9cd9cb835f93dd17df71ffc567e2

Observation 1ce74f51-9a08-4df3-8d8f-bffae76d590a · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:08.771392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:08.771392Z digest=sha256:6f4de5a95a4e7803feabe5acf2afce2160485fa8744e2046efaa178b4a6d0df8

Observation 18674d71-7b64-49eb-a2e6-63ad3c5caffa · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video question answering via gradually refined attention over appearance and motion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.460870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:08.863605Z digest=sha256:ce893a6381b9504d7adebd797e01f49362dcf724b9e1d7c69cb4a71791a81748

Observation 7b72c423-7680-4d6e-a3c0-1b7b844ae2b1 · outbound

This paper cites ActivityNet-QA: A dataset for understanding complex web videos via question answering.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? ActivityNet-QA: A dataset for understanding complex web videos via question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:41:10.295260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:41:08.944154Z digest=sha256:a6630805e08641c2d63881ac16860012df48d22f55cecb8ea2e772590d34f09f

Observation 8f4aaf61-ea0e-4a96-bc14-66fe79aab981 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.035873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.035873Z digest=sha256:183ece530189fe4e47e2d0778428520cc594d1a50aebacd602168c6bf184d80b

Observation 13727959-433c-4784-a7c7-20bf3e898152 · outbound

This paper cites LIME: Less Is More for MLLM Evaluation.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LIME: Less Is More for MLLM Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.118503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.118503Z digest=sha256:126a6cdc334ce7bd6421f499ea53831cfbe50051cd369184e9214ae78b864844

Observation 7a46c902-f8b5-4033-88b7-c49360ce7c9a · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.214825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.214825Z digest=sha256:1430ec99035b010524ed9df02e1578fbf872e5f6ce12e5a827931c02fad33f94

Observation 8d3ba9d7-f312-453b-8ac2-1aca202aa967 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.310216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.310216Z digest=sha256:5ac3ad9fe1ae46caae8d8f00ad00b3f5c1878fe928a538e62c772adebfb688ff

Observation f1591ffc-cd7f-4534-bac7-83080ef6bccc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.367899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.367899Z digest=sha256:d8b0adb194974fb98f00d7523fa0b576d5b4b6d1fdf539d76e5fd624e835d636

Observation d0d88534-9569-4b16-a312-cc1126394219 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.462250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.462250Z digest=sha256:7d3154b2878d82134fba46bce128d1d65303936dc03a7dd6d30a196f555b9fe6

Observation cd5c3eea-f837-4aaa-9500-c77784b901eb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.539358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.539358Z digest=sha256:ea973392f65bed6d5fceda4e61bb87ebf436e4a807604fb0f37e07bc9256ba16

Observation 9223dad6-8fbc-4ace-ba0a-386a2e5366f4 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.637593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.637593Z digest=sha256:83e11fc6c13a70b73b07bb1603f6d30c16f4acf7437e072acd666911fe734c19

Pith citing papers

Observation 18307915-d59b-4e62-944f-aa7c483f1624 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T13:01:57.956043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:01:57.956043Z digest=sha256:1cc63a7f206a3b6af247aaf1a430624a6525ea69538f33c95e155e3767eda938

Observation 9afbff76-9efc-4df2-9eb1-4cf3663958cc · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.806787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.806787Z digest=sha256:d5de637e48663ab3b078a877d5a7a7afe827e1463c9ad74b7ddc0f22270f7a53

Observation 4af70d98-2788-4b20-85b2-abe90fc7600a · inbound

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees cites this paper.

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:02.252494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:38:27.081358Z digest=sha256:618d5eb8369ac68846dbdf08b6b934cb7de954323f71128c1f0b638fca7e7442

Observation 8b721bc2-2993-4738-92b6-4fd426d49e1c · inbound

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark cites this paper.

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:22.517885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:18:00.436343Z digest=sha256:7ef92c74a99e1c9903e9a383ab4073c9dc25b1baf8fd701df80a3478a9baf7ae

Observation 537c814d-9e7f-4b8f-819c-53b57e0df066 · inbound

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark cites this paper.

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.422359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T00:24:25.203292Z digest=sha256:c9deddbc52f0dbe8a07142d5788021a416d171f2d2bd83871dcd59beef214156

Observation efdfaa89-8f69-4236-a9c6-9184a07e4464 · inbound

An Attribute-Based Measure of Video Complexity cites this paper.

An Attribute-Based Measure of Video Complexity Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.110328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T19:00:54.718177Z digest=sha256:e6502b13f5d51d354998c8841442dbab3989e63fd072ee7ff1d96c335ff53720

Observation 52eb70fb-bb63-4e0c-b641-84a3ea0666ea · inbound

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis cites this paper.

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:07:28.701935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:25:15.493925Z digest=sha256:fcd1fc6c9f1ca596f144af45d6687498d2a7d3dd38881db72e6b3fdb3389094f

Observation cd04825c-b05d-4ba5-9fad-add7deb39880 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:55.001360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:55.001360Z digest=sha256:8311674cd42a0d82dc974a2c890f54a5ba29bf05edad1de3be2987eed7d71a1b