Pith. sign in

Paper Citation Record · LEDGER

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

As of 8 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.06361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06361 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:55.008240Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7c84f27-de22-4830-958c-c31c17eba640 · outbound

This paper cites L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:24:56.000894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.144137Z digest=sha256:fcc75a3c4b6103c497d9ffee1a6d94ac87baf4625cb5846cc6e801c68a8c28c8

Observation 07944c51-e15a-4a6c-a568-eb97ff54b32d · outbound

This paper cites Qwen3-VL Technical Report.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.231517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.231517Z digest=sha256:a374141f7856aae6d28725befa7ca4535ec7cf9dc4151ababec600967f22bc8e

Observation 8c0f946c-1274-42ff-8e08-c1fa16448b78 · outbound

This paper cites MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.386597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.386597Z digest=sha256:10d2d6fb0613cafe92607ed6d1d6c52c214b91574f29eae3b9511405d1aac3ef

Observation 599d86f9-2b43-4dd5-94d9-a12d01fc3ef8 · outbound

This paper cites M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.691178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.488003Z digest=sha256:0013f5ab4ce414cc4dd1eb84c853ac790a4f43d579263baf33eb4def2651f945

Observation 671a9d24-71dd-4040-a285-22661e74e1c7 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.643870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.643870Z digest=sha256:0a5bbc6e850eba18200b906d5e8919e8a962a9521546dee35c5a24db9f447a06

Observation 0805b9db-5acd-4cf9-8d63-735f9a1322c1 · outbound

This paper cites P.; Li, W.; and Gong, S.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping P.; Li, W.; and Gong, S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.672317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.734869Z digest=sha256:5fe50021a5e0a67dc34d4df592f8a68f5d525d1dafd530e44cb54568ad0452bf

Observation 07139ee8-aa73-416a-984a-5340b765adc4 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.655945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.830178Z digest=sha256:2b2c9c4ba35b083c6d13c06f0ae3132ea56fd7bea4b86a185e1e6123849cb54e

Observation cfcddce4-ae72-4b94-b290-f290aef5a431 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.914333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.914333Z digest=sha256:d8b72e423ece2adc79b5b27910ae1a969c2208826985e645105e42be7c8d6f17

Observation 823347c9-3cd6-4b20-86b5-a49d03ba5008 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.639322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.012924Z digest=sha256:af29a5ffc7bfb3a9c57d414d8fa4f7d0e00dad7fa3bd7545bba7b2d3a4df0634

Observation 92ed64fa-c936-432c-a326-d779fce80713 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.620228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.214323Z digest=sha256:742ae6571d368b7f354175734df7fdb7bcc06243c95e49a71f9aa890cda4d24a

Observation 2f4fecf4-dedb-4797-aab3-cfc5ecd2ec38 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.296769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.296769Z digest=sha256:b81b319289e23f888c319215f94e93fc3e56b1153bc8612143f7713366fdf53d

Observation 30ee07bf-454e-4f99-823a-6742c436921c · outbound

This paper cites W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.602422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.418701Z digest=sha256:bc11b7cda5010a8689b0a83ac9452cae6db3824177dc871521adb71883af03d3

Observation 5a226575-a1be-4601-b9d0-1b1f8ab7ca45 · outbound

This paper cites L.; and Girshick, R.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; and Girshick, R

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.578561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.726452Z digest=sha256:13b0a10dc50d42ec76531a7f7aa55b203bcd4b458f4db599cc8b1ad3014315bd

Observation 81537453-421d-4dd5-add1-b1ca663a9126 · outbound

This paper cites VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.867176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.867176Z digest=sha256:b4b97bedb5145f34f7cf4645eaf33cf988e97103474baebf58bdc1c1a1832ba8

Observation d91c9d61-57eb-42aa-9c96-71709166d49e · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.025621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.025621Z digest=sha256:b46e1ae117411da13a5b40c5324a1959b36f77b72437c47b4ae1c1521e288d4d

Observation 7677c8b8-5d86-44cb-9938-e5aa4cda0546 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.559694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.245785Z digest=sha256:6caaff79b46b5878d914dcc33f2e3aec6de9bf38fc6e3e5ea83f36a3e6eaf0eb

Observation 5d1fd752-2a2d-4fd6-905c-cd9935f0481b · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.540311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.384918Z digest=sha256:9cc27a75867db92e4f18982389654268d7dcbcd11706088ce7f75b38906351ed

Observation 6ebfd26c-e9d8-4278-9bdb-640d56553c47 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.511970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.511970Z digest=sha256:dfc660c521fd34088f8f266049d4872a702076d89bb58096a367bdef1a19eb73

Observation 17e824f4-ed54-4ee8-8f78-5cd733c5cd47 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.524376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.664409Z digest=sha256:4cbb593e985b01fa4098004b9ca4d6b7e522e4cb0d979e3ddc49b33c83ebdaac

Observation db54f330-9e3e-4d4c-ac75-3c597c212482 · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.749022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.749022Z digest=sha256:3a9800623bf16d9d84f6da855727d830accac011e9fdf8b4e3597a9b03eec0c7

Observation 62aba644-7176-4f22-b1ed-11b45250e50a · outbound

This paper cites Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.755039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.755039Z digest=sha256:1e685261c0e3afef6fd0623d425397e346faea1a87e9d6a53043af0c5d7c901b

Observation bb9418fb-fd26-484e-91b6-02a8fea9c261 · outbound

This paper cites OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.759751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.759751Z digest=sha256:cce8c8d297d6de923c0929b0644ae17ddd95175468ca6d53f4621072d9e52634

Observation bf8a1780-e20e-43be-a311-3ad12ba35cb3 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.764865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.764865Z digest=sha256:34c1d1f14f0c8b5ee7df04d1d218a7f754f8399b25654071e9e9aa6ae9661fa7

Observation 8e51ee67-2932-4e6c-a1ff-4b3e0843eb3e · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.770078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.770078Z digest=sha256:a8189586b4f4a98dc74256d4c0657f831c25e017388b8fe5e9d6fbcc475f9883

Observation 37f66cc9-eb61-476a-bf83-4e07eab41068 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.777262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.777262Z digest=sha256:d2ec917b9f5a2cc7de69554ccb0f3abf8cee59faa60a784efdf88b39e64da844

Observation 9a10bc43-c406-4344-ad18-d58920138353 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.505970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.788332Z digest=sha256:137182c7a601aebcd502907db305024d6277f4eca4b67557c03a805a8f6b5023

Observation abc2a3de-42ec-41a4-9a5c-bffa94bcc04d · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.793516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.793516Z digest=sha256:0638fb0ea36fce02e37b07e908a8e0e99b163f93c906cc06959cff3db325a23a

Observation cb2382fa-fb81-4374-8239-61807de95740 · outbound

This paper cites J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.799125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.799125Z digest=sha256:f9c5665071e200abc1a88f50d70d63504f350688d23a698e0ce637c3779f88bb

Observation 19cbbeb2-51d4-4e30-8d45-ca82d8efd493 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.809548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.809548Z digest=sha256:da8b0e66d9a8d5a690835cedc2af4aa25d8b7b1d998bf32f32d94c6d28633e93

Observation 5f0f0fda-46cf-4be4-b234-0ebd0af11645 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.815382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.815382Z digest=sha256:04bda14c54e370dfbe6ca6920473da8fca68846ec592f0defbefdbb649ad0531

Observation 5a27b7c3-fa34-427b-bced-52db97aefa1a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.474953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.821197Z digest=sha256:f06ba340c37c8f9b0038042bfbc94c863d5d6c839d25639a85d181f6f16e2f77

Observation 1b8f7891-dae9-4ca9-8a64-7008f05e4b0a · outbound

This paper cites T emp C ompass: Do Video LLM s Really Understand Videos?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping T emp C ompass: Do Video LLM s Really Understand Videos?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.453882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.826419Z digest=sha256:55be274cd35e3de9b8e5d61c57469e97869ce626e134fb29b9ee8e77d95914f7

Observation fd09c947-e957-4f7e-a9c0-a0ae21b9e44c · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.435710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.831264Z digest=sha256:74db51992e93a26a347847aadaad6400279aa29242670f43b5c4a0bced17dcf9

Observation 7f2dcae5-0272-4da8-9aad-5be1c79680d3 · outbound

This paper cites and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.418861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.836116Z digest=sha256:650aeefccb2a27387cf84463fcee73b91787d8cc944d07cfaaf9f4513cb98268

Observation c3b7d9a1-f192-4f65-9048-b2c86a063eeb · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T04:24:55.067902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.842447Z digest=sha256:db72171d04a5e69c6e40e5dffea9f7b3110e32f8203c5b04656104108ab1ccca

Observation 8f5df2b8-86b4-45ec-be73-207a60f414fc · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.401075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.848233Z digest=sha256:414c852b3731ebadb067998ccab9a40342f9705bab209b72fc48ec7025de9ef3

Observation ace52e79-d713-41aa-9ddc-050d36c33230 · outbound

This paper cites International Conference on Learning Representations , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping International Conference on Learning Representations , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.381945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.853320Z digest=sha256:084b728327efd93cbba4d97a23568051811edbd2e50fd8718272f1bede6d313d

Observation 15810289-8cee-4a45-9f3b-74ca14ececa6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.363864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.858091Z digest=sha256:bc4e6dcc8e490e4d8f718d8ad2d5548501e5bb69fd65d1b855b9b2b3719f2282

Observation 6db93d44-64d1-43b9-b14c-0093100b1754 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.863248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.863248Z digest=sha256:a152ef05e4116daeb2c323d359e241ed3a11d30d6502f75ea2a969817a98a600

Observation 4ccd3904-b946-4a70-a5d7-9cfbd058001f · outbound

This paper cites arXiv preprint arXiv:2512.05091 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2512.05091 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.868214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.868214Z digest=sha256:44319df0a97d3bcff5c08b1dcd7e79c6af211d2d884f7fd213c819566e2da85d

Observation 679bf50f-250d-406c-b904-95b71e93cf43 · outbound

This paper cites arXiv preprint arXiv:2505.23359 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2505.23359 , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.873262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.873262Z digest=sha256:66f90701806d40f15c8f1148279480c8657eff644d7e25b6e10f8946e3e688fa

Observation b2341d0e-c233-400c-8f1b-1e587130ecf7 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.346770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.878051Z digest=sha256:c8d55ea7084b8dfa4d0b2a0ad55b3d478d9ab5dee17e8ce912a0f4668e253e75

Observation fa2ed78d-82ba-4cda-bd07-c3d25ac51509 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.330599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.883261Z digest=sha256:722d9574aeb91a58c5cb6260f922a58b09dd36aed192d2b43f771cb8d8fced9b

Observation 5788ee92-2d80-4450-8b1d-78067fbe5342 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.310306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.887607Z digest=sha256:dacf166633dfe985873bb842a6d2c15af67f931643ae5bb843a2a3191fd0ee7a

Observation aa320b76-db6a-4cf0-82d2-21f43de77622 · outbound

This paper cites TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:24:55.795302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.891710Z digest=sha256:f81382b1db3eafee02e4f2acbfad7ab511d06888f5f561cbba1ddb639640e6e1

Observation d5e61dbc-7926-43b7-8d8c-e253a4975770 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.895820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.895820Z digest=sha256:c258ee8d3603e921fbdc2519072c90b512160565222c1bf5b495ee2ab0fd2527

Observation 4dce49a0-883c-435e-8dcb-31d07b9cface · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.900531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.900531Z digest=sha256:a26510066c8a6542d2e34178083c408bfedd569ceb4fc9a7b9325d7b99bc29d9

Observation e964b1e2-40c0-4100-bca9-cc1c452ce07f · outbound

This paper cites 2026 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2026 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.904905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.904905Z digest=sha256:013ccb0857d9560f339d0f17dfdbe86923e160046f1531090152ec74060f8943

Observation 5e58298e-02f4-4315-aaad-bb7d062ea2b8 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.909336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.909336Z digest=sha256:5c1f22571116b370e3d9aba1d55932f11fe205813652b42383d660ff46fb64dd

Observation 22ba40ce-b387-4a60-892b-8dcfff417bf8 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.242939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.913877Z digest=sha256:f7b4ec975b019993c50771e6b3c487b3362639e49cecda624ff4726a9c6b5eae

Observation 4290319c-0200-434a-b1a4-ef745eae8798 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.223694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.918039Z digest=sha256:3c07c2c6456565df4eee470418ff0a7601ecb4a31da63d7a69a650b05c64b167

Observation fe6093ce-4615-4fca-a0e7-e9bed56f58aa · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.922152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.922152Z digest=sha256:9490446570e3587250478915411932055a89ddc182b68f768f8a9eb50d8e2504

Observation 95aa17f3-0196-46c6-b638-0ca70215d7d3 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.926274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.926274Z digest=sha256:a3a63df6127acc374988c4657ab1ddbd7906366ad9c60df550cc096026c375c8

Observation 42bea15d-cbf8-483c-b6a9-f9a4b53578d6 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.179831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.931250Z digest=sha256:8de36fb07ad741f4644bdfdf7985baff8da3d208f95b4cf904fb6c88fc2c3b2b

Observation 15e4ae1a-fc2e-4373-9667-8bfd1a58c922 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.935746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.935746Z digest=sha256:28246bacfc3e2c972bb596fd560005d91d3333d0ff2bf4adc87c5b2a36f9d745

Observation 9aeaa774-301f-46f9-ab75-6440462d76a2 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.939825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.939825Z digest=sha256:c5a476d66a6931a21d66f3f000914ac7e5ee0fe7e60cab20ab0e99d81ed10fe3

Observation a6797bd3-346a-4119-82f5-e3ecdddf8424 · outbound

This paper cites Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.133964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.944352Z digest=sha256:7c695304f087764ea9863fe19716686d0fe4a7832c6e9e6c5219179ed56eadcb

Observation 8e7c17d2-d4e7-448a-9acf-e130fe5d8cea · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.114458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.949338Z digest=sha256:ef4361ac4eb3d410c9f655c710a5595aebc6ea8b66600db0554604c7c2e4d771

Observation 779460ca-fa69-4b33-b59a-c3a0e5d7bffa · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.955160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.955160Z digest=sha256:268600d7607d419d94e1be982e64899a80191a857ec60b98e7e4dd187c7002ca

Observation 2283cd70-da8b-488a-a33d-4d1eaa6eec6e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Advances in Neural Information Processing Systems , volume=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.960617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.960617Z digest=sha256:23b4913729ba95b65e9c4583722c86b8efc6006d002b9864128c7b98ecb38211

Observation bcea1e89-b1a5-452d-9e38-2ce30b87d4b6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.965931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.965931Z digest=sha256:b39a4bf94988d2bb720a8ab34d425252de8f6e56915b08f2497acd720e470c4e

Observation 95794f40-93bf-49ef-8cf6-3c44bd445f98 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.970881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.970881Z digest=sha256:9ddef5e00addf7ed49fd2eaa601d8a741eeec32c7a1d6c0d15f514f77ccb526e

Observation cbe1aa48-295d-4d78-82f9-45ef25c57e8e · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.054314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.975540Z digest=sha256:3c9bc89e4e8d3a3cd875dd7296bbe0dbf651a83aa92542e9a9069844859b4a2c

Observation e059e240-3680-4a48-8baa-dbd830e98b72 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.981358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.981358Z digest=sha256:6f9bedfb3dcb7da5bd34cdbbf84af75ab53c7fb9694c436287da7c5e802af63f

Observation 44b4a782-ef62-4217-b57a-179974c28124 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.985958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.985958Z digest=sha256:c22f2e62a194ac36ca7de0b74474404ab7983b04c35761afb600c3f6123ee6ff

Observation 4e648f65-6112-46d9-9f6a-e5484246083f · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.991413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.991413Z digest=sha256:6aed30acf7c060df871370c7b57032edf5996abaab13d3bd1fe472129aa83cbb

Observation 6472183e-2c3b-4b5c-9936-59eaf987834c · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.036757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.996689Z digest=sha256:14bc58c2fcb8a1f748fad9e361ac65af588ab954e2d362b724be8947e22b4142

Observation cd04825c-b05d-4ba5-9fad-add7deb39880 · outbound

This paper cites Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:55.001360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:55.001360Z digest=sha256:6b641b7be9bef1b003f25097783e0aec1d73e93c52e5c1212d3f5c64c99abe2a

Observation 529b2d0c-38a0-4bd4-8414-23b60bbb4008 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.018883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T04:24:55.008240Z digest=sha256:64651330aa6c390bd548e8c09561677287f2efb8ccdacf03b4440c0ebcba3cbb

Pith citing papers

No inbound Pith citation observations are available.