Pith. sign in

Paper Citation Record · LEDGER

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.06361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06361 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:55.008240Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7c84f27-de22-4830-958c-c31c17eba640 · outbound

This paper cites L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:24:56.000894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.144137Z digest=sha256:c9eb227e22b9d23b2252ee0b817d6cff1674e308a52483de7796f99f4617bd8f

Observation 07944c51-e15a-4a6c-a568-eb97ff54b32d · outbound

This paper cites Qwen3-VL Technical Report.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.231517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.231517Z digest=sha256:a0a7aceacf5ea6ca3a58c3be102985667a01ed2e52db0029f66142d591da08e8

Observation 8c0f946c-1274-42ff-8e08-c1fa16448b78 · outbound

This paper cites MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.386597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.386597Z digest=sha256:3aa38a0e25e614f6b5d6ca177c73b5c3b21e03aedd1047ceb8001ba84071afb2

Observation 599d86f9-2b43-4dd5-94d9-a12d01fc3ef8 · outbound

This paper cites M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.691178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.488003Z digest=sha256:adc24ac098b24a6660c5623272866d1ceb86425bf5df69dff35641be1da693fa

Observation 671a9d24-71dd-4040-a285-22661e74e1c7 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.643870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.643870Z digest=sha256:c239aee96a7a9f2caa5d820ba80530c26a25487a729953326d32d7fa7f978c7c

Observation 0805b9db-5acd-4cf9-8d63-735f9a1322c1 · outbound

This paper cites P.; Li, W.; and Gong, S.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping P.; Li, W.; and Gong, S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.672317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.734869Z digest=sha256:f0a654b9e592d13a6d16dbb1985709b3e9d3f07b547b9b5563715acd57e883c9

Observation 07139ee8-aa73-416a-984a-5340b765adc4 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.655945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.830178Z digest=sha256:b68c07680bb90a2f340cb3d3daf4494818ce4af1347d6c49091b813a33183711

Observation cfcddce4-ae72-4b94-b290-f290aef5a431 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.914333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.914333Z digest=sha256:87c9ca0ff6f37c39555baba4c6a51f5a1702f21aa7252b385e71e86ef9aa23bf

Observation 823347c9-3cd6-4b20-86b5-a49d03ba5008 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.639322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.012924Z digest=sha256:cf2eafb0df4c87c49fe7904488fb79956a84fbc5bcfcf05fcb2b07b9c6e94d3a

Observation 92ed64fa-c936-432c-a326-d779fce80713 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.620228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.214323Z digest=sha256:bb5419fbba42aa1e8d396902aa27598a3e9395b053ec11e5403c0a974aecd6c0

Observation 2f4fecf4-dedb-4797-aab3-cfc5ecd2ec38 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.296769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.296769Z digest=sha256:fc15ac9ed7af494658ddb01d598cf27a75581084ac61cf7123b3d3b027b8dce4

Observation 30ee07bf-454e-4f99-823a-6742c436921c · outbound

This paper cites W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.602422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.418701Z digest=sha256:1bd9e2b7e4d1cb70583fe70bc748c63bfd7cdd6ce89f5818949c765f100db745

Observation 5a226575-a1be-4601-b9d0-1b1f8ab7ca45 · outbound

This paper cites L.; and Girshick, R.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; and Girshick, R

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.578561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.726452Z digest=sha256:ec792d0a1faed6641e708bfd9f116ec8dc389ebc20cf259ea8692c97c767a145

Observation 81537453-421d-4dd5-add1-b1ca663a9126 · outbound

This paper cites VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.867176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.867176Z digest=sha256:b8e94dffea8e2b89bea6417ea59a50028a8c937d3ba25b3dd6b8c17d0835233e

Observation d91c9d61-57eb-42aa-9c96-71709166d49e · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.025621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.025621Z digest=sha256:f80b3b186a696f4dca50962d7d0624a00393b43b05f34685b17035534d83f6b1

Observation 7677c8b8-5d86-44cb-9938-e5aa4cda0546 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.559694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.245785Z digest=sha256:b9d14f3c2f26f7ae2a7ba673bbf7897184c5bed15d74c7e87a8e459b9de5cdf5

Observation 5d1fd752-2a2d-4fd6-905c-cd9935f0481b · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.540311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.384918Z digest=sha256:090fcd28c3dbbe039d5eb41118978a7c4b889a76db13f0f06f633b5bc3456268

Observation 6ebfd26c-e9d8-4278-9bdb-640d56553c47 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.511970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.511970Z digest=sha256:64b32a30d0af52b9a5afdec5aff831262d7f87c08ac30570c339224cb30041ff

Observation 17e824f4-ed54-4ee8-8f78-5cd733c5cd47 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.524376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.664409Z digest=sha256:ae600d8b547660c62e8cce2ed3932392b13591eb2a54f994bccb3930bbd72ab5

Observation db54f330-9e3e-4d4c-ac75-3c597c212482 · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.749022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.749022Z digest=sha256:8744107729afaa03b14db4702d9b5264a40a0ecf3b1a792f4909acfad8de0d16

Observation 62aba644-7176-4f22-b1ed-11b45250e50a · outbound

This paper cites Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.755039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.755039Z digest=sha256:c24f9cad7eb598d992fe75fc937928037c95412051875862bd7b8d9dbf748853

Observation bb9418fb-fd26-484e-91b6-02a8fea9c261 · outbound

This paper cites OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.759751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.759751Z digest=sha256:7104ac29ba796f1097dbf9963db3ac98da9c22468ac66f748e3a79bff461d4ec

Observation bf8a1780-e20e-43be-a311-3ad12ba35cb3 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.764865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.764865Z digest=sha256:6ad530aec8f074da523072884e83d1c209b8ac06ccda93335ecddb632e98140f

Observation 8e51ee67-2932-4e6c-a1ff-4b3e0843eb3e · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.770078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.770078Z digest=sha256:b26613d39ba8f35a4686e2b695daddcff47adb80885aa203313c6df35a6a56e0

Observation 37f66cc9-eb61-476a-bf83-4e07eab41068 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.777262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.777262Z digest=sha256:da436f330ca9105e5f5a54f16435043880b55ab364e96fd03ac22c39ac1af6a3

Observation 9a10bc43-c406-4344-ad18-d58920138353 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.505970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.788332Z digest=sha256:8ce98112c993fcef7d90b2d37a6134b69459c7b8ae5aa598d6b483e7e0f58b9f

Observation abc2a3de-42ec-41a4-9a5c-bffa94bcc04d · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.793516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.793516Z digest=sha256:e9c3bafa0da05813f017a67cce66f33f9cc56aab8f3d3920c2442025588cca89

Observation cb2382fa-fb81-4374-8239-61807de95740 · outbound

This paper cites J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.799125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.799125Z digest=sha256:2a6b624c5a10f7723e3a474b4718b2c49d7ade4000adaf1b43c9915886530bc1

Observation 19cbbeb2-51d4-4e30-8d45-ca82d8efd493 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.809548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.809548Z digest=sha256:21ac976c7e5d5bd4545f8f5fa17855f9dd010012f58a9810ecae7f0b35f9293e

Observation 5f0f0fda-46cf-4be4-b234-0ebd0af11645 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.815382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.815382Z digest=sha256:fdcafd5ed5e654646c5a60b44fa3230f0e56a818da0761ca80809ab884c6ef0a

Observation 5a27b7c3-fa34-427b-bced-52db97aefa1a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.474953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.821197Z digest=sha256:79b89b09e3e39f5aff1fbf7a07f806a013d3aaa3b37e1510f1896943d265c842

Observation 1b8f7891-dae9-4ca9-8a64-7008f05e4b0a · outbound

This paper cites T emp C ompass: Do Video LLM s Really Understand Videos?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping T emp C ompass: Do Video LLM s Really Understand Videos?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.453882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.826419Z digest=sha256:45fb6e2cbb7e5805204ee4c9452c87ece6cdf22579ce36f089611d4d43eb565a

Observation fd09c947-e957-4f7e-a9c0-a0ae21b9e44c · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.435710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.831264Z digest=sha256:ea15a8f09299f4553211db6a7e4c7dbcc0c8c075468b8dc5c4ea3002ef9c7fb4

Observation 7f2dcae5-0272-4da8-9aad-5be1c79680d3 · outbound

This paper cites and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.418861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.836116Z digest=sha256:56db2c0d59d79eb666cd016346828e7a10ac1fe95109dbbf204cecfef48b29c5

Observation c3b7d9a1-f192-4f65-9048-b2c86a063eeb · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T04:24:55.067902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.842447Z digest=sha256:5b1168cb937128ceaec6db08534f9a2ae2d45037e1bf849da6fa6e68e3a0fb66

Observation 8f5df2b8-86b4-45ec-be73-207a60f414fc · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.401075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.848233Z digest=sha256:de262a2ad8bd2af8d2d42cd04dbd770f1b35215cc489c24baf18aeca0ed76f73

Observation ace52e79-d713-41aa-9ddc-050d36c33230 · outbound

This paper cites International Conference on Learning Representations , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping International Conference on Learning Representations , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.381945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.853320Z digest=sha256:4b680562fe37d410af8d8e2c5d1b97d3339742750db8ec5ddacf24b13f4e3d84

Observation 15810289-8cee-4a45-9f3b-74ca14ececa6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.363864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.858091Z digest=sha256:6c90ed46a7ce78ff6b3b8896ef8db4496a95a29512eada8fc79ac8c977acc5eb

Observation 6db93d44-64d1-43b9-b14c-0093100b1754 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.863248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.863248Z digest=sha256:69f2dbbcc1df2a88b4f88fd683b8ed6dd017686d09fffa498bff525c42c992dc

Observation 4ccd3904-b946-4a70-a5d7-9cfbd058001f · outbound

This paper cites arXiv preprint arXiv:2512.05091 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2512.05091 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.868214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.868214Z digest=sha256:d9bc22ac2a55faf4a84fb097b7e1e456af9d3c659635a8efed7bc15b1c2ae129

Observation 679bf50f-250d-406c-b904-95b71e93cf43 · outbound

This paper cites arXiv preprint arXiv:2505.23359 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2505.23359 , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.873262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.873262Z digest=sha256:841289a85558563efe22409f877a2ed2910bd3a5f1c59ac6be207632373c717a

Observation b2341d0e-c233-400c-8f1b-1e587130ecf7 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.346770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.878051Z digest=sha256:7d8ba030ca8b1d41e7433a7726abb361b181ae3498d8ac4fe94ff19423933d9e

Observation fa2ed78d-82ba-4cda-bd07-c3d25ac51509 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.330599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.883261Z digest=sha256:e61b04ca13a98826df42b22d5821504699c637c6a24a9a7fc51fe45f8e3dc348

Observation 5788ee92-2d80-4450-8b1d-78067fbe5342 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.310306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.887607Z digest=sha256:e3233184aed49dbf51e97792c0cbc58cb765dcd61cf3ad6e604afe336598c2ce

Observation aa320b76-db6a-4cf0-82d2-21f43de77622 · outbound

This paper cites TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:24:55.795302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.891710Z digest=sha256:c00984cb183c43d657ec019275b7c0a6acf5341eb7336411e6c475cdf6e79bf3

Observation d5e61dbc-7926-43b7-8d8c-e253a4975770 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.895820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.895820Z digest=sha256:4dce5b2afad6480ea56cdf2638e6722de5d447febb6370e364a3368db8f5ff40

Observation 4dce49a0-883c-435e-8dcb-31d07b9cface · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.900531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.900531Z digest=sha256:88f81d51723bf3598ab45162ef3775b05994429be25ad882740c304811be0b69

Observation e964b1e2-40c0-4100-bca9-cc1c452ce07f · outbound

This paper cites 2026 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2026 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.904905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.904905Z digest=sha256:e33b5758d561e1e3a9845a522db9913097a772e68aed88e404e1067b3df7e0e7

Observation 5e58298e-02f4-4315-aaad-bb7d062ea2b8 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.909336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.909336Z digest=sha256:8772046fde4b278f97a6b891bfcfecb9966f5ca55451cd7d1b7746e2a9886474

Observation 22ba40ce-b387-4a60-892b-8dcfff417bf8 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.242939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.913877Z digest=sha256:b61252d29aed3cc37b58e252651dfb91e0ef83539ff176fc3e9d54ba3e33e178

Observation 4290319c-0200-434a-b1a4-ef745eae8798 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.223694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.918039Z digest=sha256:ec7ddff6a8912d480a5073c9cca66c56b89fb81391ae940b9934a3c7d9b71b0f

Observation fe6093ce-4615-4fca-a0e7-e9bed56f58aa · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.922152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.922152Z digest=sha256:11792b91096010688c51d5c4c8a2e112d6f2bb2cac8ec8be5f2b48a704307266

Observation 95aa17f3-0196-46c6-b638-0ca70215d7d3 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.926274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.926274Z digest=sha256:cb73b7735e16edc5859f7fca9b511eb0ae3bfe99405f9eae48e93e44618ebf98

Observation 42bea15d-cbf8-483c-b6a9-f9a4b53578d6 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.179831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.931250Z digest=sha256:4bd8481441464b2bd842a0cbab61b09440329ec19a41d166542e1010b3c57c46

Observation 15e4ae1a-fc2e-4373-9667-8bfd1a58c922 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.935746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.935746Z digest=sha256:741a31ae1f1a3011ba0cb4a522a8b5ad8d06fd337dc31c8ca23c8985ba9339b6

Observation 9aeaa774-301f-46f9-ab75-6440462d76a2 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.939825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.939825Z digest=sha256:f08abe68ee31f9d806a7be41ec0280ab3d07cb29bf8750c941e398138235da89

Observation a6797bd3-346a-4119-82f5-e3ecdddf8424 · outbound

This paper cites Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.133964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.944352Z digest=sha256:619249355a5f33cc482d39079df2e86dbd30121daf6b4e24efdbb19f7681c1a7

Observation 8e7c17d2-d4e7-448a-9acf-e130fe5d8cea · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.114458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.949338Z digest=sha256:3dedec2b7deb8c3da38f5143eb056cabb6d0d5e52b4ef2bc184d58c82b2e1c57

Observation 779460ca-fa69-4b33-b59a-c3a0e5d7bffa · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.955160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.955160Z digest=sha256:ca6a40e291d8c62d507cf3d86c7b038595b3629123f56f3438e0b94e04819527

Observation 2283cd70-da8b-488a-a33d-4d1eaa6eec6e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Advances in Neural Information Processing Systems , volume=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.960617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.960617Z digest=sha256:0f579613475694e199abf216600f4bbc974958dfb56286b29859d77f408b37a6

Observation bcea1e89-b1a5-452d-9e38-2ce30b87d4b6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.965931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.965931Z digest=sha256:5f30ecd3c97bfd3fa55696b6570cf071e8170ffb4bb55548951cf699598a2a07

Observation 95794f40-93bf-49ef-8cf6-3c44bd445f98 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.970881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.970881Z digest=sha256:6c051a4bd2c45bc44d952fae08832b41285d4042fe43b2dbe068191993edb56f

Observation cbe1aa48-295d-4d78-82f9-45ef25c57e8e · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.054314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.975540Z digest=sha256:db51810ebebf8ed9997420af902220c0bc91d69b25c66854277a8b56dbad9d49

Observation e059e240-3680-4a48-8baa-dbd830e98b72 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.981358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.981358Z digest=sha256:1d052725acf9f8a489f00ab3d61d4326ce2b247e2c32e19d1880ebd6d318b8a5

Observation 44b4a782-ef62-4217-b57a-179974c28124 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.985958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.985958Z digest=sha256:6503685556c04156cbdcb3443b8b301b26d822d99a2e8b8f6bff47ee41033ff0

Observation 4e648f65-6112-46d9-9f6a-e5484246083f · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.991413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.991413Z digest=sha256:44d5c2965d7385a98fe7ef82dacc42c22dad1bd8762600b71950997a1f39c81d

Observation 6472183e-2c3b-4b5c-9936-59eaf987834c · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.036757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.996689Z digest=sha256:cdac08885945557f8492f6bb57983956ccab80955c19acf350584ade1bb21139

Observation cd04825c-b05d-4ba5-9fad-add7deb39880 · outbound

This paper cites Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:55.001360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:55.001360Z digest=sha256:77b793dd73b6109c94775c366fd6f07fb31c6fe884401aa562423b5e4aa39229

Observation 529b2d0c-38a0-4bd4-8414-23b60bbb4008 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.018883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:24:55.008240Z digest=sha256:6ff248bc304f58d75f49282f8cb5cba20a871a56b846330621512542e85696a6

Pith citing papers

No inbound Pith citation observations are available.