Pith. sign in

Paper Citation Record · LEDGER

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

As of 8 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.06361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06361 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:55.008240Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7c84f27-de22-4830-958c-c31c17eba640 · outbound

This paper cites L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; Panda, S.; Meghwani, H.; Singh, J.; Dua, K.; Li, P.; Sheng, T.; Ravi, S.; and Roth, D

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:24:56.000894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.144137Z digest=sha256:22fe62d34ff20ca1d73d0d265b5a6cb1034ccdfb260bfd66cb186a1883a2140d

Observation 07944c51-e15a-4a6c-a568-eb97ff54b32d · outbound

This paper cites Qwen3-VL Technical Report.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.231517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.231517Z digest=sha256:f6aac41d7f90df383456cf44c9cbc0a5e5c369e36950885f67b510955494d4ba

Observation 8c0f946c-1274-42ff-8e08-c1fa16448b78 · outbound

This paper cites MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.386597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.386597Z digest=sha256:b0d48c1683297c56d1a4d2d1fb5e1559422a94a20eb1328f1a2719432f3c2345

Observation 599d86f9-2b43-4dd5-94d9-a12d01fc3ef8 · outbound

This paper cites M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping M.; Kota, T.; He, J.; Eyzaguirre, C.; Durante, Z.; Li, M.; Wu, J.; and Fei-Fei, L

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.691178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.488003Z digest=sha256:8d55004e63dc952cc2279ac5edd1f3710c8172b90e66c21dfc07d49994bd8f73

Observation 671a9d24-71dd-4040-a285-22661e74e1c7 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.643870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.643870Z digest=sha256:5bef9ec165040c2730af50c6535dfe1c3af947112e0d288a0bf62bf0a2277ed3

Observation 0805b9db-5acd-4cf9-8d63-735f9a1322c1 · outbound

This paper cites P.; Li, W.; and Gong, S.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping P.; Li, W.; and Gong, S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.672317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.734869Z digest=sha256:399e651d1ad1d48d7471c483747f34ed5c07d6845df52cd44139de71e4856d79

Observation 07139ee8-aa73-416a-984a-5340b765adc4 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.655945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:52.830178Z digest=sha256:5df9c66b04704c34cdb28808ad9428dbb22e82db0d899679fc6d557c580fe430

Observation cfcddce4-ae72-4b94-b290-f290aef5a431 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:52.914333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:52.914333Z digest=sha256:785672c20a57b640a94539532796c94c57d5b8941e056753b2c17db50a8a1897

Observation 823347c9-3cd6-4b20-86b5-a49d03ba5008 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.639322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.012924Z digest=sha256:c9bbb56bf9b610b352e6367bda3dbe5e210f1b8e247d37ceca511602d17b1d28

Observation 92ed64fa-c936-432c-a326-d779fce80713 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.620228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.214323Z digest=sha256:aaba2bdccfdf6fa1c5af37cb3af93514dd9f7f29c4bc04c9c1c0eab3324bea5f

Observation 2f4fecf4-dedb-4797-aab3-cfc5ecd2ec38 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.296769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.296769Z digest=sha256:ac8a48f77220ff4249aae0743974cb9d4f6f0a9278c7537bf80bedc7fd91c583

Observation 30ee07bf-454e-4f99-823a-6742c436921c · outbound

This paper cites W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping W.; Li, L.; Yang, Z.; Wang, L.; and Cheng, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.602422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.418701Z digest=sha256:22813d9e0ab5a54843e2245bd003280c5c8b0cb6451f47911bc0c8c8bb7fca0b

Observation 5a226575-a1be-4601-b9d0-1b1f8ab7ca45 · outbound

This paper cites L.; and Girshick, R.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping L.; and Girshick, R

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.578561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:53.726452Z digest=sha256:62092cf972401cbce735dfd3ac0a254a410fdc30355894963fee8864f51fd170

Observation 81537453-421d-4dd5-add1-b1ca663a9126 · outbound

This paper cites VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:53.867176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:53.867176Z digest=sha256:b6e23ffae4b1dd60e290cf73067f94b7fc20f8cc682553e9f5d3120e57fdc08d

Observation d91c9d61-57eb-42aa-9c96-71709166d49e · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.025621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.025621Z digest=sha256:1d2b61fcee85257ea65287af5175bdc51d5161a80431c2a1503a00cfb15b7cfb

Observation 7677c8b8-5d86-44cb-9938-e5aa4cda0546 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.559694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.245785Z digest=sha256:1a340d46d779b8e5ee71c2be0478c5b688ffcb04f66885125e324f76200b1f11

Observation 5d1fd752-2a2d-4fd6-905c-cd9935f0481b · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.540311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.384918Z digest=sha256:9bce0f3aa1915bbf9d7920bedcd2639c4dce3233da85a50c3a16017c6579ee3e

Observation 6ebfd26c-e9d8-4278-9bdb-640d56553c47 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.511970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.511970Z digest=sha256:6d98e3480c1b7d19fa830110073f94bdc57a21f5de1ada55a6d859c40435dd66

Observation 17e824f4-ed54-4ee8-8f78-5cd733c5cd47 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.524376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.664409Z digest=sha256:0f4bc9ac1c576c9699bb63ba049a71a1969f6fbb07ee5b98eac4f97141a31025

Observation db54f330-9e3e-4d4c-ac75-3c597c212482 · outbound

This paper cites The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.749022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.749022Z digest=sha256:2d87994eac92f93fc6e14272c20fbcc4a35c439b0958af2b22c9673d697cbf3c

Observation 62aba644-7176-4f22-b1ed-11b45250e50a · outbound

This paper cites Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.755039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.755039Z digest=sha256:4a9568a269d8cefe4bd97690a23664b74bb36bcfd7e70a4f473760e7184bc4f7

Observation bb9418fb-fd26-484e-91b6-02a8fea9c261 · outbound

This paper cites OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.759751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.759751Z digest=sha256:3e2c6fae97abe39a7b8c8de61a47ecddea94f88af85ecf28aea7e1ed2f2b394b

Observation bf8a1780-e20e-43be-a311-3ad12ba35cb3 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.764865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.764865Z digest=sha256:ed957404fc24a9e9142cf776f9361aa360751633987de9c436c8b05eb47626ed

Observation 8e51ee67-2932-4e6c-a1ff-4b3e0843eb3e · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.770078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.770078Z digest=sha256:3945f9f7979874e5fa9704a2af09feef2e20a30ddf2640fa50a4b65244b0b7b3

Observation 37f66cc9-eb61-476a-bf83-4e07eab41068 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.777262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.777262Z digest=sha256:0384a13cd31f516c7440913673018e733d4ef47460d0d6b56124b9c7d163bd81

Observation 9a10bc43-c406-4344-ad18-d58920138353 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.505970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.788332Z digest=sha256:6cb9549abb7782b9f0a2d387f3bf692f568d0583265fe47b6eccae8e870b2759

Observation abc2a3de-42ec-41a4-9a5c-bffa94bcc04d · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.793516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.793516Z digest=sha256:f2282066defde5ca93372d4c40ff2cb8a277472bb3f40599dbad9b06703c467c

Observation cb2382fa-fb81-4374-8239-61807de95740 · outbound

This paper cites J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping J.; Huang, Y.; Liu, Z.; Qu, P.; He, J.; Chen, J.; Yuan, Y.-J.; Han, J.; Xu, H.; Li, H.; Sachan, M.; and Liang, X

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.799125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.799125Z digest=sha256:7b7e3d7bebb33a4a06fcbc367a7cf9f240e7e20da7c08126189230c8c8783a5c

Observation 19cbbeb2-51d4-4e30-8d45-ca82d8efd493 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.809548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.809548Z digest=sha256:150246601d6eda5f790896d33588fd45d4ab95caa9fb0ac34e06ba36c8cff116

Observation 5f0f0fda-46cf-4be4-b234-0ebd0af11645 · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.815382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.815382Z digest=sha256:44873e38cd765c65460b22a2ba98ec0e40eb923f6898bb392cfb0b9833a078dc

Observation 5a27b7c3-fa34-427b-bced-52db97aefa1a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.474953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.821197Z digest=sha256:65e116ce0bf2e6250cf1ce9f6f48460e8c06550b7f181223a19bc099bcc1c90f

Observation 1b8f7891-dae9-4ca9-8a64-7008f05e4b0a · outbound

This paper cites T emp C ompass: Do Video LLM s Really Understand Videos?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping T emp C ompass: Do Video LLM s Really Understand Videos?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.453882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.826419Z digest=sha256:573383f09bf270ade438d22e227f86f83e6c15af38b7d64d61be1f16b226e37a

Observation fd09c947-e957-4f7e-a9c0-a0ae21b9e44c · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.435710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.831264Z digest=sha256:b257c8ceec7ccf8d6840b6dc336d98aa2439b9f856365d94fdbe2405f11baa6a

Observation 7f2dcae5-0272-4da8-9aad-5be1c79680d3 · outbound

This paper cites and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Fei-Fei, Li , booktitle =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.418861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.836116Z digest=sha256:50ced498b11426670d8b0f22e415bf531404275a8362b3374146c7d33e0ee80d

Observation c3b7d9a1-f192-4f65-9048-b2c86a063eeb · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding , volume =

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T04:24:55.067902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.842447Z digest=sha256:13e16c64e7176c9dffecd1d33c10b9f8595b942652ff6089ecea3f2e37e9aa38

Observation 8f5df2b8-86b4-45ec-be73-207a60f414fc · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.401075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.848233Z digest=sha256:d1c6556d7f8a8e38095ba049a81b4f8155a7745445d4549f707a6a47e7e4984c

Observation ace52e79-d713-41aa-9ddc-050d36c33230 · outbound

This paper cites International Conference on Learning Representations , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping International Conference on Learning Representations , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.381945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.853320Z digest=sha256:6974a282a01d70cecce49df64f5d40e0e1237b6faa77918026d84ea20fe070fd

Observation 15810289-8cee-4a45-9f3b-74ca14ececa6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.363864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.858091Z digest=sha256:602999a36c91dcfa1078c43ef2efb84d92b883c9a40e377c9376f8d95b93a8ca

Observation 6db93d44-64d1-43b9-b14c-0093100b1754 · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.863248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.863248Z digest=sha256:d0deb63af7074f0743759736c6545b97b39c0ea38a4b1fae2072367edb5756dc

Observation 4ccd3904-b946-4a70-a5d7-9cfbd058001f · outbound

This paper cites arXiv preprint arXiv:2512.05091 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2512.05091 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.868214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.868214Z digest=sha256:cbaaeda6b12836a3d2702682aecb10707405a9d0b6b196d62eae34a5e1169199

Observation 679bf50f-250d-406c-b904-95b71e93cf43 · outbound

This paper cites arXiv preprint arXiv:2505.23359 , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping arXiv preprint arXiv:2505.23359 , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.873262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.873262Z digest=sha256:cd2f57c09e77e6ab11c13e123f2562273b7c87ed8065694638b51a94adb70059

Observation b2341d0e-c233-400c-8f1b-1e587130ecf7 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.346770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.878051Z digest=sha256:f98333955de3b27606e589c26d37270d631cadcdc52af50a446dbb70ee71bc3d

Observation fa2ed78d-82ba-4cda-bd07-c3d25ac51509 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs , volume =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.330599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.883261Z digest=sha256:0f3d3120b034634f9c60b8947555d1458215db8da37adbd6ed05a1eb1677bc96

Observation 5788ee92-2d80-4450-8b1d-78067fbe5342 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.310306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.887607Z digest=sha256:05b57fd8c90c4c714ea13c0f816aebfc7a720d867ce08c177ab3f5ae81d8caa2

Observation aa320b76-db6a-4cf0-82d2-21f43de77622 · outbound

This paper cites TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:24:55.795302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.891710Z digest=sha256:3bfde370b59cc45fd09ddb402b9af49408ac4aea3abd4153885024a6b71a9023

Observation d5e61dbc-7926-43b7-8d8c-e253a4975770 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.895820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.895820Z digest=sha256:5507af1dfd1086719a10095ab2016fa388ac8c223d08d7bde419444c1d511801

Observation 4dce49a0-883c-435e-8dcb-31d07b9cface · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.900531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.900531Z digest=sha256:e82b41588fd10584f469cd90a93b975555c04c3a1377360b0a7396a082a39bdd

Observation e964b1e2-40c0-4100-bca9-cc1c452ce07f · outbound

This paper cites 2026 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2026 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.904905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.904905Z digest=sha256:b24591cc9b6c8a29b8d14ec3dc57103f7459a801076fc72a614246310e22bfd0

Observation 5e58298e-02f4-4315-aaad-bb7d062ea2b8 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.909336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.909336Z digest=sha256:1f0af6864180bcdbe5abd3c51922e769f1408004b782fcc1b8b30e84c1408ace

Observation 22ba40ce-b387-4a60-892b-8dcfff417bf8 · outbound

This paper cites Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.242939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.913877Z digest=sha256:96b6a69454033c36e49174556c9359e19cc2d2625ca9bd35e1d1657229d6b37c

Observation 4290319c-0200-434a-b1a4-ef745eae8798 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.223694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.918039Z digest=sha256:8c00c44a45edf5f3b283adf9fe9ed8d1cee433f911338d9edef17957e10c5d34

Observation fe6093ce-4615-4fca-a0e7-e9bed56f58aa · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.922152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.922152Z digest=sha256:d535a14b172e29c69211d766953ffc47d5bd3fd1ca8a874d170e5a3931d47cb2

Observation 95aa17f3-0196-46c6-b638-0ca70215d7d3 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.926274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.926274Z digest=sha256:cca493441b89bce5f1248a3fbd652a540df7c1bbcb61aa01c0706c5e3db8cbbb

Observation 42bea15d-cbf8-483c-b6a9-f9a4b53578d6 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.179831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.931250Z digest=sha256:109d1f84e633c95b0f143a7487b1f851dc16d4e2ea91973e6e3dfb7c11eb2702

Observation 15e4ae1a-fc2e-4373-9667-8bfd1a58c922 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.935746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.935746Z digest=sha256:0ca64b6cfbd7a888f76b24cdd93f4103183bb9f282abcf18dbf41e721714ceb7

Observation 9aeaa774-301f-46f9-ab75-6440462d76a2 · outbound

This paper cites 2024 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2024 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.939825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.939825Z digest=sha256:ecb224d4af71d6b63301b8a1e69c3adf4f49409c6d6d3b681449f7116c26a5a2

Observation a6797bd3-346a-4119-82f5-e3ecdddf8424 · outbound

This paper cites Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the 42nd International Conference on Machine Learning (ICML 2025) , year =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.133964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.944352Z digest=sha256:041b456fd99f12630b29b1ecb72535b735b59162abb368a54d153ba1e4f53efe

Observation 8e7c17d2-d4e7-448a-9acf-e130fe5d8cea · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.114458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.949338Z digest=sha256:f53dd0928a7083c8658963e72ae386b94f34253c1c096c27c9760ca97f177492

Observation 779460ca-fa69-4b33-b59a-c3a0e5d7bffa · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.955160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.955160Z digest=sha256:69bd17bdc7c36d14b15bc5c8a9443df45db16afd3bb253e5545a635031dfd6bb

Observation 2283cd70-da8b-488a-a33d-4d1eaa6eec6e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Advances in Neural Information Processing Systems , volume=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.960617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.960617Z digest=sha256:3b06052fd329e1fd1c2e5c171d2827311bee6841f0b9de8e264494613207d9c1

Observation bcea1e89-b1a5-452d-9e38-2ce30b87d4b6 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.965931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.965931Z digest=sha256:5a42e9cf34d533566b8d884127e51c563af31cd41dd891785c5817c0f0c6d540

Observation 95794f40-93bf-49ef-8cf6-3c44bd445f98 · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.970881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.970881Z digest=sha256:54e93f05491ca81fb92d9dbbbb0d457aa786c5d5c90aeadf1751e37118ae4fe1

Observation cbe1aa48-295d-4d78-82f9-45ef25c57e8e · outbound

This paper cites 2025 , eprint=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2025 , eprint=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.054314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.975540Z digest=sha256:5633ed9ddcb653d4c1334aeb3407446377d60a5caf5ae6716b58b241e836f291

Observation e059e240-3680-4a48-8baa-dbd830e98b72 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.981358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.981358Z digest=sha256:3f1129e23727ce19917cd7051a05838a8086f497c0d66366e866d5e5cb3dba75

Observation 44b4a782-ef62-4217-b57a-179974c28124 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.985958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.985958Z digest=sha256:380d8f0e94c9ad4bc03669d260f076aa58cd9a456b1b07b346983accb0747804

Observation 4e648f65-6112-46d9-9f6a-e5484246083f · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.991413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.991413Z digest=sha256:e5818b549c1bf05b76464065d65024bac71d124d69b03c6448d728f4516c6004

Observation 6472183e-2c3b-4b5c-9936-59eaf987834c · outbound

This paper cites an unresolved cited work.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:24:56.036757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:54.996689Z digest=sha256:17bb1e48243ffeac5c25e087f01b58e7d81ffe91a6d6e87f4e0a98ec78cc10e7

Observation cd04825c-b05d-4ba5-9fad-add7deb39880 · outbound

This paper cites Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:55.001360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:55.001360Z digest=sha256:5eac4effad5281782d82967f45cfd324b8fca97672f7ce18271d61df903da213

Observation 529b2d0c-38a0-4bd4-8414-23b60bbb4008 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:24:56.018883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:24:55.008240Z digest=sha256:315a8ed4d60f073d8258b0696a400e19f75ae44078ddeeb0522368bc1de056f2

Pith citing papers

No inbound Pith citation observations are available.