Pith. sign in

Paper Citation Record · LEDGER

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.15028.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15028 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:38.262186Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:12:09.932226Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5df95d99-6457-4549-98e8-bc2296fb4eaf · outbound

This paper cites Vqa: Visual question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:48.049637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:30.939545Z digest=sha256:ab25edb4d998336b45a5eebc846799679da212f7938b6b4d184b7ce87808ccf9

Observation f91d00e8-c6f6-491f-b1a5-7bdc970597a8 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.035101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.035101Z digest=sha256:b1f322d39261f1db4a57e99a3c44abb7904202663a68249470dd3a20ca297491

Observation 455bff05-8a2e-4bca-bcb1-2426ad781cbc · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Collecting highly paral- lel data for paraphrase evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.771747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:31.146121Z digest=sha256:4dc7d4b0bdcf4749aeacccfe9aba2aa6e5af16d1afb0aa6aad68195bc63322c2

Observation 5d1d99ad-54ad-465f-83df-4725f87d8d78 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.309034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.309034Z digest=sha256:dabb39a907cb4db24110084cd240955402e6f7045884dea6e0d54422c822c5de

Observation 7811dae3-8c3b-4a69-ac98-5634d2a7ae5b · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.423977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.423977Z digest=sha256:01df4bfe2ab359bb1b7042f06decc914f81b9c97931081432ce035d8da2b3a36

Observation 419a10bf-2fc6-4e6e-a497-af875e6cbe22 · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.569319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.569319Z digest=sha256:6234a0ca739b553547333b836ba0056b1507ced81404189ee2c04d3c480c171c

Observation 20cae88d-8699-4879-bab5-a606e491c550 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.692811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.692811Z digest=sha256:0bb7afeccd50780306627f33aeb72e84b9fe48ef620d3a6dda4ab8b5e1fe81cb

Observation 6601ce9b-8936-4466-a2e9-2e1660cd6698 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.570794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:31.798160Z digest=sha256:e1c3dfbcf7f55a0c7d50b43dbf18b22dffbd1b69ddc2c8671e25bc7f9347f98c

Observation c5bb15e9-d108-40c0-95af-fc0fd2632a22 · outbound

This paper cites Similarity and fea- tures of natural textures.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Similarity and fea- tures of natural textures

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.327426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:31.967725Z digest=sha256:5a7d230c09d1671cfb7a97fc3b1e7821a924abde88b734b4684fb36abc40d36f

Observation fef48a7e-f210-40f0-a400-b5f99936a1a2 · outbound

This paper cites Natural adversarial examples.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Natural adversarial examples

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:47.012478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:32.140486Z digest=sha256:02accd79e2e73724dc03113688380f63f1ef9b63286b7618e9270beec0e5b04b

Observation db349333-1200-42ed-9571-7d04be0ba812 · outbound

This paper cites Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.740041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:32.242853Z digest=sha256:f9aec4735d5697586e1247323b4115f8466d8180b07cd05e1cd079208858d457

Observation 09e9c5cb-554a-4e44-a014-9f9206135d35 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.351024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.351024Z digest=sha256:b0920050ce3ec67bc139cf5eb48037fadb3cc67b724106a7008a01d881ef2a32

Observation ddf0b73b-170f-460c-a1b1-b11ec693ce2a · outbound

This paper cites Robust modeling in cognitive science.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robust modeling in cognitive science

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.500872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:32.475759Z digest=sha256:a811676b391b73e3e5fb4a80b63c6f9463d16e675b8d9113c5f9a51c61427c48

Observation 3a2e6533-99d6-4ee3-b0d6-3e5259281859 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TVQA: Localized, Compositional Video Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.606188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.606188Z digest=sha256:bf0f45a391d93efe266e609f95366b7835a9cef59aaa1c9333b0816339e7ae6a

Observation 1ac2132a-7884-43c0-b223-c3ae4ba09978 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2023.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Mvbench: A comprehensive multi- modal video understanding benchmark, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.290809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:32.768765Z digest=sha256:d91b7cf128c94f3e53a903c16fcfe01caadb5146043a1bc882c17dc5c0b62011

Observation a15c7e92-a2ef-4d72-87a3-cf47fee2a286 · outbound

This paper cites Grounded language-image pre-training.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Grounded language-image pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:32.871605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:32.871605Z digest=sha256:e08a79d38180cb328ce829fe2a4ae9b83e83ef3301e1253282c5a7d898592c2a

Observation 7196bd3a-3cfa-4d87-8e80-86f386fdbdef · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TempCompass: Do Video LLMs Really Understand Videos?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.000709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.000709Z digest=sha256:5eadd1689f95ef18e8d92c507fec2d3d6c3bd5a019fdef67ed65f4501068a814

Observation ef22287e-6b45-409f-83b5-5da365626f32 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.180292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.180292Z digest=sha256:f6a227368444332c3574be7ae80f851133236a47035c9d7424b2bf0c2cacc31d

Observation c23e29bd-49d4-4ee2-a256-716b3e5727d9 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.294475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.294475Z digest=sha256:a4f93722f4371de15fc6185eb41829332e96bb88ec2ab420a55f7f44fca82d11

Observation 37c6c68c-dfa0-458c-aa6f-d438fb959ff9 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.446078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.446078Z digest=sha256:7fe474b9cd719ce1a801e6943513df8ea0ffc1331681a271fc6e6d8ef1054e34

Observation 765b1288-3e8a-46f3-a236-2e205096e505 · outbound

This paper cites Video detail caption, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video detail caption, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.002685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:33.588967Z digest=sha256:f331e39e0c10fcf63051d80bfaac9ef656a9edf5ff7f3003e24c3b22480b334f

Observation d89edf69-5883-4e24-b7c9-6462a3481db4 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.768516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:33.712835Z digest=sha256:78dd2d2b8f7df0430d32c579547187677a74ca09c23b0818289f39f100fa3cb7

Observation a8110d16-6f0f-469c-847f-5dddec64fb52 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.555971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:33.843680Z digest=sha256:c6ededda8f76b66ed7211c49189cf2cf997e03bea7d5c4bac33be38b789ccf95

Observation e2f1aa5c-f826-4ed0-bcbd-fdbc2fef3663 · outbound

This paper cites Identifying the perceptual dimensions of visual complexity of scenes.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Identifying the perceptual dimensions of visual complexity of scenes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.268685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:33.963318Z digest=sha256:ceb0ca21b579bcc5b2f0fae39dfe49ce86375e156ec5923e92752234050bdd46

Observation 48f1ad09-e144-43d9-8af3-4afa5c535bfe · outbound

This paper cites Hello gpt-4o.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hello gpt-4o

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.992721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:34.090197Z digest=sha256:dda15ffb63c7cfee9c116f06664ed294113789a54f38385a55c8f80c90751930

Observation 8b509cc9-209d-490f-bf22-4e17d8a20f18 · outbound

This paper cites Robustness analysis of video- language models against visual and language perturbations.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robustness analysis of video- language models against visual and language perturbations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.690772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:34.215109Z digest=sha256:25c739ab5b84cddae284c0ee75457830af3a22ad8cea986a89f2a4bf76b95565

Observation 8624ad6a-af68-4a15-b742-e3fedcfbcd4c · outbound

This paper cites Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.432636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:34.327788Z digest=sha256:85141da61dd44bce1b90bcac82eddaac1438a359fd8afed904333e6b1e34c657

Observation 3c072a10-ee74-403c-baa0-46613aa21e3a · outbound

This paper cites Complex narratives.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Complex narratives

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.211901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:34.436761Z digest=sha256:3bb8cce7409e3e5eb933e59c640f1c8d8e2ea25d7892bc9d119d635ae6c75ca0

Observation 8b651d54-b4a8-49bb-baa6-c740cfbb7d51 · outbound

This paper cites A standardized set of 260 pictures: norms for name agreement, image agree- ment, familiarity, and visual complexity.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding A standardized set of 260 pictures: norms for name agreement, image agree- ment, familiarity, and visual complexity

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.994154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:34.545309Z digest=sha256:133fd97fc240902438ea3092dd4c80a0e9271540d3dab895002eba21ca0b9b95

Observation 83a81e53-1273-43de-940c-88950fadb301 · outbound

This paper cites Visual Agents as Fast and Slow Thinkers.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual Agents as Fast and Slow Thinkers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:34.682053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:34.682053Z digest=sha256:44e1115b583e0956c5be60ac87c93bd64eab98b0d9ca50c7cc960dc771bbb4fe

Observation 4be120cf-78ac-448b-a092-a75e3fc7285c · outbound

This paper cites Curious objects: How vi- sual complexity guides attention and engagement.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Curious objects: How vi- sual complexity guides attention and engagement

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.757236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:34.828072Z digest=sha256:152fd150d9124609fc6dd97c944f7b6030605888b9b030fc219740eb1495a1b3

Observation c6d3932e-6194-486a-946a-42cb352693bb · outbound

This paper cites Cognitive load during problem solving: Ef- fects on learning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Cognitive load during problem solving: Ef- fects on learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.510219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:34.956258Z digest=sha256:7ff19a43442fd0a26aa1f589e8c47d32e8dfaba340aeff2ebde20d123c6bca1f

Observation 80026ae7-918d-488a-a749-558296988b73 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:35.106376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:35.106376Z digest=sha256:9e8f8f7d0014dd5c8d9db5df15145df433f80c9f42d17c800b45f517ed906b2b

Observation 42b5846c-c080-4b1c-b886-c596ca59958e · outbound

This paper cites Qwen2.5-vl, 2025.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Qwen2.5-vl, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.298939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:35.204531Z digest=sha256:6ede9bae6ad447cbc41afe45d82f691c167950b57cb9e0a11f40c0febfc269ee

Observation fc779f9f-e966-485d-b1a5-df5446d54495 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.068789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:35.335517Z digest=sha256:8fe393f072dc24a064907289d05cd456a40791619d0e19002335f1bbef7a1740

Observation 544bbc8a-a175-4be3-93a3-8ee3c3b255c1 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Star: A benchmark for situated reasoning in real-world videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.832074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:35.438170Z digest=sha256:de998dadd9a804223180c897e9ff9601a506dcdeb4535fd349027a978e725987

Observation d1b49f1d-6949-4034-b8e3-6b30f8aca8ca · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.558730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:35.549580Z digest=sha256:345d6e59dc720d5e936a84d06d8d4eda7bfe9911a7709589c5fca750ceb3406f

Observation b79f48c1-7e3f-451e-8075-04db7d05888d · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.321877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:35.677619Z digest=sha256:dd2d31ee18035c8fe39808f64133ab0a990e196db7c0f647f8142d43d8446ab0

Observation be4cf59e-3032-4c5d-bc8b-fb4d9ffe5508 · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding FunQA: Towards Surprising Video Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:35.898141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:35.898141Z digest=sha256:cd8ad98512a25bf651c3123eacca5474e5cb98d7cbb2666841df7d85ccde8916

Observation 8615b8f0-4b17-4cfd-8b16-7733e1e6d649 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.030420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:36.010706Z digest=sha256:b12509f8f8154a12099dff9814584369076b3cb4798ab3401d7b2c48abcb01f4

Observation e765a4ed-3fc3-4c43-a5a5-09eb60da7363 · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.727394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:36.149550Z digest=sha256:e66e8218760f8d9e5a59ec7a5f5c537714d1b6cb2806122c72bb3f03f9470012

Observation 698ac212-0b18-4618-aebe-1225e15ab204 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:36.257872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:36.257872Z digest=sha256:db0fe4f83846af0d94636b14eef512350addd9d9c561d094875cf9fc1f66e03a

Observation b02e17e6-0383-4c89-b9c7-9ef1a00d10c9 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.517617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:36.361773Z digest=sha256:1a340cc80bf1af88cf6222dac0fd4d6699286e704d2aa590ac599a89b37b7082

Observation b49b54d2-abcb-4a86-aee5-8c7353405cd7 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.286050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:36.499719Z digest=sha256:7bb2ebf893d64737f06a7cec0e54c894e2ad5b9c3eb2f0821c4b03eb8a2f5f15

Observation 29c5a716-8008-4753-9afe-d0c154f4e367 · outbound

This paper cites Social-iq: A question answer- ing benchmark for artificial social intelligence.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Social-iq: A question answer- ing benchmark for artificial social intelligence

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.022509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:36.656128Z digest=sha256:e750f69e3ea4b6cda17c03be7b12af14203c9a6a89e17f8078ccac65ea8db4a4

Observation 263fd12e-d9a0-49d4-8027-6217a5f54e3e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:36.756457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:36.756457Z digest=sha256:0245a86b693061b62df0214e892c26facfc10264662786047a0fdd219d00a595

Observation bbe8c10f-0b64-438a-b966-196f252f44a8 · outbound

This paper cites B- avibench: Towards evaluating the robustness of large vision- language model on black-box adversarial visual-instructions,.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding B- avibench: Towards evaluating the robustness of large vision- language model on black-box adversarial visual-instructions,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.773699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:36.911831Z digest=sha256:0e4f353fe495eec6c9e82a4cb001558243ab1992663fe35737de685b621aed9b

Observation 7d7965d5-2b29-4fac-8610-7ec7b3bbff7a · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.032657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.032657Z digest=sha256:2d12a53887284fd84115e8bce54c89a5192ea8c2eae9df45db5a9749e2f43eec

Observation fe22ba31-6ef0-422e-a89e-4b5188af112e · outbound

This paper cites Long Context Transfer from Language to Vision.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Long Context Transfer from Language to Vision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.124321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.124321Z digest=sha256:3e1f137b62207226520b1e78305bfe801df2bdf772320587188ddfe68b5b1db3

Observation 0d4b1904-9a86-4848-a923-e83820d76d7e · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Llava- next: A strong zero-shot video understanding model, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.466201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:37.232691Z digest=sha256:50119b35ba6fbd610ffcbe3b6fc010f845d7cd2388da498ccf0f0d3aeb3b0b44

Observation 8364d8e5-e494-420c-88ef-7445affc3200 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video instruction tuning with synthetic data, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.216826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:37.388074Z digest=sha256:6364d9d3c92e435a1b4ebad46df777fc2fb6ce8da410fad0ccd20030baa38ce6

Observation 91e9c0ea-54b4-49e1-91ac-05896924ac87 · outbound

This paper cites Worldqa: Multimodal world knowledge in videos through long-chain reasoning, 2024.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Worldqa: Multimodal world knowledge in videos through long-chain reasoning, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.995188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:37.519206Z digest=sha256:58a793565fb2a215efc4012779366bdbfd9ac3d68da2ebfb731909f394673450

Observation 2571e13e-793f-4fc5-bb6d-3ebd87c5ccdd · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:37.637372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:37.637372Z digest=sha256:4dde19a2281f60cfe16bba8680b50a999a992ab320fc07b24f4035b8a81e9ae6

Observation ddccc800-85a9-4126-838a-ac6ec14d2461 · outbound

This paper cites Hierarchical video content description and summarization using unified semantic and visual similarity.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hierarchical video content description and summarization using unified semantic and visual similarity

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.748619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:37.755434Z digest=sha256:d1dbe9183b4a70ba18e944ca620fceb1f7b10fed52fb07a5eb72c92bc15c5c37

Observation 12f08e3e-6f6d-4d24-8a18-9afa369badf2 · outbound

This paper cites In total, the annotation process cost 8227.32 human hours.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding In total, the annotation process cost 8227.32 human hours

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.494867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:37.878023Z digest=sha256:cf2ef06d1696d5a8a08c1e6c102521c88d957ca38e5a9cad33e0d776cd3d82a8

Observation 5e1f3085-8a05-494a-8fe5-7f5260b953d0 · outbound

This paper cites • Aparaphrased correct be the set of videos where the para- phrased open-ended question is answered correctly.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding • Aparaphrased correct be the set of videos where the para- phrased open-ended question is answered correctly

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.209624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:37.984186Z digest=sha256:13fec8a17055cf7ef1393b7f945e85d694019b783f3bdceee58406b25805b418

Observation 4661c413-e575-40cf-8cb5-b6b4260aa7d2 · outbound

This paper cites 3 shows the prompt for evaluating open-ended an- swers.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding 3 shows the prompt for evaluating open-ended an- swers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.958699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:38.104212Z digest=sha256:d31e4f3e14f75aba44a719f8b476ac8d58e30536f9552facecbad5381e324be4

Observation f2b4833d-0cf6-4494-a4ee-ad5b033b6135 · outbound

This paper cites element” and “event.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding element” and “event

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.751950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:38.262186Z digest=sha256:6a188e5508a15fa8842d6b25a32251232f10c64e321be187234ef2d78fdf9ceb

Pith citing papers

Observation 029bc2f4-af1c-4e0d-9ee6-6c811843f413 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.932226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.932226Z digest=sha256:c211ff218e26ded3097163c45dfad921926bc513a82d07c0a009644811792551