Pith. sign in

Paper Citation Record · LEDGER

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.15028.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15028 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:38.262186Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:12:09.932226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5df95d99-6457-4549-98e8-bc2296fb4eaf · outbound

This paper cites Vqa: Visual question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:48.049637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:30.939545Z digest=sha256:ea339f4d55f73427d6e17302fcc6db022d119abe33efb29f82a75aa586a5b706

Observation f91d00e8-c6f6-491f-b1a5-7bdc970597a8 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.035101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.035101Z digest=sha256:080eb3d895ec9e2f65647c11530a0de381321dbacb9a6c7589206e2433b54a4d

Observation 455bff05-8a2e-4bca-bcb1-2426ad781cbc · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Collecting highly paral- lel data for paraphrase evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.771747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:31.146121Z digest=sha256:7d74e420b25f16ac5cd3ffed0d8429e721fc84b963ca48fc8b24220fe9e52037

Observation 5d1d99ad-54ad-465f-83df-4725f87d8d78 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.309034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.309034Z digest=sha256:31b8140f12a0cc14a6c15a00f46ffd0f9209d27f57d3e0d83027537345b5e19c

Observation 7811dae3-8c3b-4a69-ac98-5634d2a7ae5b · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.423977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.423977Z digest=sha256:d7fde377aec9773f52672f178ed2257d0abf0dec3dc924a710f09bb3238946b6

Observation 419a10bf-2fc6-4e6e-a497-af875e6cbe22 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.569319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.569319Z digest=sha256:c71d9da35b652292e7590b6a2da23973ae8ce16279914293ca4adef9f7804d81

Observation 20cae88d-8699-4879-bab5-a606e491c550 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.692811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.692811Z digest=sha256:5c4d03330906c9c44911a47b42b1c70946ae33e42efbf660d9ff8d97c9e2c2a0

Observation 6601ce9b-8936-4466-a2e9-2e1660cd6698 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.570794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:31.798160Z digest=sha256:c1ffb3020d2bbb550bac3ddc71cad9ebd32a81d484fc87aa7abd8e978e5a8e8f

Observation c5bb15e9-d108-40c0-95af-fc0fd2632a22 · outbound

This paper cites Similarity and fea- tures of natural textures.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Similarity and fea- tures of natural textures

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.327426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:31.967725Z digest=sha256:9dc087edbe8cb7dd034b6739a81331d9259323aace23fb926d694836bb1f5baf

Observation fef48a7e-f210-40f0-a400-b5f99936a1a2 · outbound

This paper cites Natural adversarial examples.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Natural adversarial examples

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.012478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:32.140486Z digest=sha256:8e505a41a6be6c1228745ee9c7f53c3869a74e228cd38cb37f76e186f20ef83b

Observation db349333-1200-42ed-9571-7d04be0ba812 · outbound

This paper cites Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.740041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:32.242853Z digest=sha256:034df44844a656260ed72e3b1059b6d5df1df4837e85d17ce858a5e5a7f043cf

Observation 09e9c5cb-554a-4e44-a014-9f9206135d35 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.351024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.351024Z digest=sha256:4eb5402bef40d189d9ae328aee951c0b91473e02ea14935ce63d57754fa8810e

Observation ddf0b73b-170f-460c-a1b1-b11ec693ce2a · outbound

This paper cites Robust modeling in cognitive science.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robust modeling in cognitive science

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.500872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:32.475759Z digest=sha256:60f53d4cbb87138f104f4d78178d41f19da8c019566c782c297c89679fda357c

Observation 3a2e6533-99d6-4ee3-b0d6-3e5259281859 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TVQA: Localized, Compositional Video Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.606188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.606188Z digest=sha256:99e31067914336c64b4623425eede789f08a6b04824322cef371542ff20dc2f9

Observation 1ac2132a-7884-43c0-b223-c3ae4ba09978 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2023.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Mvbench: A comprehensive multi- modal video understanding benchmark, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.290809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:32.768765Z digest=sha256:9af6df09df3a6c8d39abacff87face150c807ca6a1c0a6e4a5d9571fbf15ef15

Observation a15c7e92-a2ef-4d72-87a3-cf47fee2a286 · outbound

This paper cites Grounded language-image pre-training.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Grounded language-image pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.871605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.871605Z digest=sha256:e3f219fcbf77ea4966bcef07fa5ca917e51f8d579c34b97faf0ff563a269eda5

Observation 7196bd3a-3cfa-4d87-8e80-86f386fdbdef · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TempCompass: Do Video LLMs Really Understand Videos?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.000709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.000709Z digest=sha256:b3adb75d2fdb4bd322be87ffc000edc8185f7e269d74204c1426927783c882f5

Observation ef22287e-6b45-409f-83b5-5da365626f32 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.180292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.180292Z digest=sha256:6de62139ea2d88c4d5673ee44552dbbc03b233e3992b88110ce6d901d98b24b1

Observation c23e29bd-49d4-4ee2-a256-716b3e5727d9 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.294475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.294475Z digest=sha256:9736d0ba7596a1a3fa68dc38d4125cb417c4afad526f2058789b1efe7e5778cf

Observation 37c6c68c-dfa0-458c-aa6f-d438fb959ff9 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.446078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.446078Z digest=sha256:f4a5bd39f9a9760aef76c2699cb323bde730efe0cbef385aff8099ade37a485a

Observation 765b1288-3e8a-46f3-a236-2e205096e505 · outbound

This paper cites Video detail caption, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video detail caption, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.002685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:33.588967Z digest=sha256:f0356a12426d16cb3132c34c6f3b79719d297bfa1435f8419584cacf4a3e84b5

Observation d89edf69-5883-4e24-b7c9-6462a3481db4 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.768516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:33.712835Z digest=sha256:6303c0d770f1729c3cd5e55a937fb39b2757b2988030d8caac9265294055da5e

Observation a8110d16-6f0f-469c-847f-5dddec64fb52 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.555971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:33.843680Z digest=sha256:e3b58f6233a63ed248446ca067c96e6e3aca4b3a8a2830b89c8cf428225b4d8f

Observation e2f1aa5c-f826-4ed0-bcbd-fdbc2fef3663 · outbound

This paper cites Identifying the perceptual dimensions of visual complexity of scenes.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Identifying the perceptual dimensions of visual complexity of scenes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.268685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:33.963318Z digest=sha256:ded8e541400eb2c6b64331895096b43204e8e67443d98291b60705f31c9b5cb5

Observation 48f1ad09-e144-43d9-8af3-4afa5c535bfe · outbound

This paper cites Hello gpt-4o.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hello gpt-4o

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.992721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:34.090197Z digest=sha256:716aa09fb5177f51f39e9028899c5665977585fccbec7fe9b3bbb3c50250a38a

Observation 8b509cc9-209d-490f-bf22-4e17d8a20f18 · outbound

This paper cites Robustness analysis of video- language models against visual and language perturbations.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robustness analysis of video- language models against visual and language perturbations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.690772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:34.215109Z digest=sha256:a1210ea8e8c8f874b5a63337c0ebb6745270738d1c6d9b2122b2a879f6d1e6af

Observation 8624ad6a-af68-4a15-b742-e3fedcfbcd4c · outbound

This paper cites Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.432636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:34.327788Z digest=sha256:d2ad95902f5825fef88aaa1dac5bb503cb7d8cefa139cadd8e6cc71ba78a62e3

Observation 3c072a10-ee74-403c-baa0-46613aa21e3a · outbound

This paper cites Complex narratives.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Complex narratives

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.211901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:34.436761Z digest=sha256:ec7cd69d3290ec48b2dbfe85f39f742188d1064aa5f70423053d412101e4a517

Observation 8b651d54-b4a8-49bb-baa6-c740cfbb7d51 · outbound

This paper cites A standardized set of 260 pictures: norms for name agreement, image agree- ment, familiarity, and visual complexity.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding A standardized set of 260 pictures: norms for name agreement, image agree- ment, familiarity, and visual complexity

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.994154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:34.545309Z digest=sha256:0cb287d13a0c2c73f84b334b0b67dc9ed0965492a6865ff5450ee4c899e082b8

Observation 83a81e53-1273-43de-940c-88950fadb301 · outbound

This paper cites Visual Agents as Fast and Slow Thinkers.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual Agents as Fast and Slow Thinkers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:34.682053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:34.682053Z digest=sha256:86f5e3fd3baea01fbf57b09947e825f00a4b6819196a630dd045c81c3b81defa

Observation 4be120cf-78ac-448b-a092-a75e3fc7285c · outbound

This paper cites Curious objects: How vi- sual complexity guides attention and engagement.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Curious objects: How vi- sual complexity guides attention and engagement

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.757236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:34.828072Z digest=sha256:68191636ec40255fea1ce829fb4f9d4ba5f75dd693f524c761c3bcc69b9795ce

Observation c6d3932e-6194-486a-946a-42cb352693bb · outbound

This paper cites Cognitive load during problem solving: Ef- fects on learning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Cognitive load during problem solving: Ef- fects on learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.510219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:34.956258Z digest=sha256:5b274a046dcebc3244dba81188433b9ab360223290339ed974aa530e0df206b9

Observation 80026ae7-918d-488a-a749-558296988b73 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:35.106376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:35.106376Z digest=sha256:820c1fe0995ad9378305099ab9703abbc63c2975542d4345cd795f74a38417e9

Observation 42b5846c-c080-4b1c-b886-c596ca59958e · outbound

This paper cites Qwen2.5-vl, 2025.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Qwen2.5-vl, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.298939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:35.204531Z digest=sha256:e7e99e8faad4366ed307c00abdec1a0c9ff4b9b9c6549e558bfd011257e9fe70

Observation fc779f9f-e966-485d-b1a5-df5446d54495 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.068789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:35.335517Z digest=sha256:485de2f3f036d99c9e2274bf76590645b9e1a0d1872d0e971135f5eceb6cb6d7

Observation 544bbc8a-a175-4be3-93a3-8ee3c3b255c1 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Star: A benchmark for situated reasoning in real-world videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.832074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:35.438170Z digest=sha256:f109b52e3c92438bdc5f7d2409bc4746b93682367b87cd6cfd33d32631cb72e6

Observation d1b49f1d-6949-4034-b8e3-6b30f8aca8ca · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.558730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:35.549580Z digest=sha256:fda1d9f19745ae0be41bfa842394c189e101af1403c3325778d46f3fae75d7fb

Observation b79f48c1-7e3f-451e-8075-04db7d05888d · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.321877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:35.677619Z digest=sha256:27dd545960d315e8948a98fca28999fc45b25190316c3784ce07744d6bbed419

Observation be4cf59e-3032-4c5d-bc8b-fb4d9ffe5508 · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding FunQA: Towards Surprising Video Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:35.898141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:35.898141Z digest=sha256:b23ab8f7dd5b476bb679fc95cfd6dccfc9963d3e16c2f8a99f9aa875385d9a0b

Observation 8615b8f0-4b17-4cfd-8b16-7733e1e6d649 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.030420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:36.010706Z digest=sha256:74583abc5968612a4130ca0b34c1ed2aadd137c277e607f1622bc4383aa12fd7

Observation e765a4ed-3fc3-4c43-a5a5-09eb60da7363 · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.727394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:36.149550Z digest=sha256:27c15e19c81763700f16e87fd9b9db3a1c3a361f3aef5132789bc5201b62dc91

Observation 698ac212-0b18-4618-aebe-1225e15ab204 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:36.257872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:36.257872Z digest=sha256:e0d9a3e3eecd2c54ca5a7969ab2a98ff0545d8963b573d9370cfaa080522b411

Observation b02e17e6-0383-4c89-b9c7-9ef1a00d10c9 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.517617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:36.361773Z digest=sha256:0d9ecc2cc4605f14c355d7d16c5567205362c0b523d6217ff07ced0270c1b0f3

Observation b49b54d2-abcb-4a86-aee5-8c7353405cd7 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.286050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:36.499719Z digest=sha256:b3cb28a49d32e133ed163dff854159901e3cee00ad7a74e34333a72cafcfd0f6

Observation 29c5a716-8008-4753-9afe-d0c154f4e367 · outbound

This paper cites Social-iq: A question answer- ing benchmark for artificial social intelligence.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Social-iq: A question answer- ing benchmark for artificial social intelligence

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.022509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:36.656128Z digest=sha256:64d83d959fc2c145470b1e3cef444f205d74a5645c8dd9c042baa8d87e121aa5

Observation 263fd12e-d9a0-49d4-8027-6217a5f54e3e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:36.756457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:36.756457Z digest=sha256:f76cc957f807d007ac2d748c6af6dfa5bbdf03a63066a6990121005fc9a9a7e3

Observation bbe8c10f-0b64-438a-b966-196f252f44a8 · outbound

This paper cites B- avibench: Towards evaluating the robustness of large vision- language model on black-box adversarial visual-instructions,.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding B- avibench: Towards evaluating the robustness of large vision- language model on black-box adversarial visual-instructions,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.773699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:36.911831Z digest=sha256:5f2668b95496963fa76c0bdb2e03bde5d4e5183f96bd41497677217180ece861

Observation 7d7965d5-2b29-4fac-8610-7ec7b3bbff7a · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.032657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.032657Z digest=sha256:3f534d35d5dfc0152f221edd5c8732bb709eb82fc584fda1b83df143637bb8e3

Observation fe22ba31-6ef0-422e-a89e-4b5188af112e · outbound

This paper cites Long Context Transfer from Language to Vision.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Long Context Transfer from Language to Vision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.124321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.124321Z digest=sha256:6dda6abf04e4d46b3628ceb0211611d420ad9c20175b56d3b1ba820a2595550c

Observation 0d4b1904-9a86-4848-a923-e83820d76d7e · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.466201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:37.232691Z digest=sha256:691261250eb22f53ea5649d7b5914f42af8428d1f2bb003de29a83e75fac74ed

Observation 8364d8e5-e494-420c-88ef-7445affc3200 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video instruction tuning with synthetic data, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.216826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:37.388074Z digest=sha256:099f5973382d3152f071daa0eaa453cefe6e16a5f23f4c2e7e44e4bbb3e5713b

Observation 91e9c0ea-54b4-49e1-91ac-05896924ac87 · outbound

This paper cites Worldqa: Multimodal world knowledge in videos through long-chain reasoning, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Worldqa: Multimodal world knowledge in videos through long-chain reasoning, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.995188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:37.519206Z digest=sha256:8a2de32a6f0a415056ecfcd5a625be0de5ccffb05ebd1447abf1a83586cba496

Observation 2571e13e-793f-4fc5-bb6d-3ebd87c5ccdd · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.637372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.637372Z digest=sha256:91ae0b4722baba855ebc4faee1d93ee01056f3754d8b8750b2a1092b5d42373a

Observation ddccc800-85a9-4126-838a-ac6ec14d2461 · outbound

This paper cites Hierarchical video content description and summarization using unified semantic and visual similarity.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hierarchical video content description and summarization using unified semantic and visual similarity

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.748619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:37.755434Z digest=sha256:a4d3d19038ff972323ab8c72fdc976108410bd3793d88ad23b92821509737710

Observation 12f08e3e-6f6d-4d24-8a18-9afa369badf2 · outbound

This paper cites In total, the annotation process cost 8227.32 human hours.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding In total, the annotation process cost 8227.32 human hours

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.494867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:37.878023Z digest=sha256:7a939a887bf12548c8d1f58f21047d3adff58ab9f07c4528b6833cb7421064b7

Observation 5e1f3085-8a05-494a-8fe5-7f5260b953d0 · outbound

This paper cites • Aparaphrased correct be the set of videos where the para- phrased open-ended question is answered correctly.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding • Aparaphrased correct be the set of videos where the para- phrased open-ended question is answered correctly

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.209624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:37.984186Z digest=sha256:df57519085c648a9e63ec81248d46b0dc291bc31ef2c4d1d15faceee1ac1e01e

Observation 4661c413-e575-40cf-8cb5-b6b4260aa7d2 · outbound

This paper cites 3 shows the prompt for evaluating open-ended an- swers.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding 3 shows the prompt for evaluating open-ended an- swers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.958699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:38.104212Z digest=sha256:31c6f4df08fdd65f997ce7827b190c8e46c2fc63bc53032a63d22f65b7a0a198

Observation f2b4833d-0cf6-4494-a4ee-ad5b033b6135 · outbound

This paper cites element” and “event.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding element” and “event

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.751950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:48:38.262186Z digest=sha256:a20bd0d84edcabbe6336c0be6e973c154f5edf93cbc905257ef7f3528861bdd7

Pith citing papers

Observation 029bc2f4-af1c-4e0d-9ee6-6c811843f413 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.932226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.932226Z digest=sha256:0011b3a8ce1afca40c76d6ca8e7b6cdf7144028bd9ceaa6ccc3fd84774ad8605