Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:38.262186Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2507.15028.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:38.262186Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:12:09.932226Z
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5df95d99-6457-4549-98e8-bc2296fb4eaf · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Vqa: Visual question answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f91d00e8-c6f6-491f-b1a5-7bdc970597a8 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 455bff05-8a2e-4bca-bcb1-2426ad781cbc · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Collecting highly paral- lel data for paraphrase evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d1d99ad-54ad-465f-83df-4725f87d8d78 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7811dae3-8c3b-4a69-ac98-5634d2a7ae5b · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419a10bf-2fc6-4e6e-a497-af875e6cbe22 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20cae88d-8699-4879-bab5-a606e491c550 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6601ce9b-8936-4466-a2e9-2e1660cd6698 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Agqa: A benchmark for compositional spatio-temporal reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5bb15e9-d108-40c0-95af-fc0fd2632a22 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Similarity and fea- tures of natural textures
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fef48a7e-f210-40f0-a400-b5f99936a1a2 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Natural adversarial examples
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db349333-1200-42ed-9571-7d04be0ba812 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09e9c5cb-554a-4e44-a014-9f9206135d35 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf0b73b-170f-460c-a1b1-b11ec693ce2a · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robust modeling in cognitive science
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a2e6533-99d6-4ee3-b0d6-3e5259281859 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TVQA: Localized, Compositional Video Question Answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac2132a-7884-43c0-b223-c3ae4ba09978 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Mvbench: A comprehensive multi- modal video understanding benchmark, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a15c7e92-a2ef-4d72-87a3-cf47fee2a286 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Grounded language-image pre-training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7196bd3a-3cfa-4d87-8e80-86f386fdbdef · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding TempCompass: Do Video LLMs Really Understand Videos?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef22287e-6b45-409f-83b5-5da365626f32 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23e29bd-49d4-4ee2-a256-716b3e5727d9 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c6c68c-dfa0-458c-aa6f-d438fb959ff9 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765b1288-3e8a-46f3-a236-2e205096e505 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video detail caption, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d89edf69-5883-4e24-b7c9-6462a3481db4 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8110d16-6f0f-469c-847f-5dddec64fb52 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2f1aa5c-f826-4ed0-bcbd-fdbc2fef3663 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Identifying the perceptual dimensions of visual complexity of scenes
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48f1ad09-e144-43d9-8af3-4afa5c535bfe · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hello gpt-4o
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b509cc9-209d-490f-bf22-4e17d8a20f18 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Robustness analysis of video- language models against visual and language perturbations
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8624ad6a-af68-4a15-b742-e3fedcfbcd4c · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c072a10-ee74-403c-baa0-46613aa21e3a · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Complex narratives
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b651d54-b4a8-49bb-baa6-c740cfbb7d51 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding A standardized set of 260 pictures: norms for name agreement, image agree- ment, familiarity, and visual complexity
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83a81e53-1273-43de-940c-88950fadb301 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Visual Agents as Fast and Slow Thinkers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be120cf-78ac-448b-a092-a75e3fc7285c · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Curious objects: How vi- sual complexity guides attention and engagement
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6d3932e-6194-486a-946a-42cb352693bb · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Cognitive load during problem solving: Ef- fects on learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80026ae7-918d-488a-a749-558296988b73 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b5846c-c080-4b1c-b886-c596ca59958e · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Qwen2.5-vl, 2025
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc779f9f-e966-485d-b1a5-df5446d54495 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Chain-of-thought prompting elicits reasoning in large language models, 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 544bbc8a-a175-4be3-93a3-8ee3c3b255c1 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Star: A benchmark for situated reasoning in real-world videos
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1b49f1d-6949-4034-b8e3-6b30f8aca8ca · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b79f48c1-7e3f-451e-8075-04db7d05888d · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Next-qa: Next phase of question-answering to explaining temporal actions
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be4cf59e-3032-4c5d-bc8b-fb4d9ffe5508 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding FunQA: Towards Surprising Video Comprehension
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8615b8f0-4b17-4cfd-8b16-7733e1e6d649 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e765a4ed-3fc3-4c43-a5a5-09eb60da7363 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 698ac212-0b18-4618-aebe-1225e15ab204 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02e17e6-0383-4c89-b9c7-9ef1a00d10c9 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b49b54d2-abcb-4a86-aee5-8c7353405cd7 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29c5a716-8008-4753-9afe-d0c154f4e367 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Social-iq: A question answer- ing benchmark for artificial social intelligence
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 263fd12e-d9a0-49d4-8027-6217a5f54e3e · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe8c10f-0b64-438a-b966-196f252f44a8 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding B- avibench: Towards evaluating the robustness of large vision- language model on black-box adversarial visual-instructions,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d7965d5-2b29-4fac-8610-7ec7b3bbff7a · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe22ba31-6ef0-422e-a89e-4b5188af112e · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Long Context Transfer from Language to Vision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4b1904-9a86-4848-a923-e83820d76d7e · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Llava- next: A strong zero-shot video understanding model, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8364d8e5-e494-420c-88ef-7445affc3200 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Video instruction tuning with synthetic data, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91e9c0ea-54b4-49e1-91ac-05896924ac87 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Worldqa: Multimodal world knowledge in videos through long-chain reasoning, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2571e13e-793f-4fc5-bb6d-3ebd87c5ccdd · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MLVU: Benchmarking Multi-task Long Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddccc800-85a9-4126-838a-ac6ec14d2461 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Hierarchical video content description and summarization using unified semantic and visual similarity
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12f08e3e-6f6d-4d24-8a18-9afa369badf2 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding In total, the annotation process cost 8227.32 human hours
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e1f3085-8a05-494a-8fe5-7f5260b953d0 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding • Aparaphrased correct be the set of videos where the para- phrased open-ended question is answered correctly
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4661c413-e575-40cf-8cb5-b6b4260aa7d2 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding 3 shows the prompt for evaluating open-ended an- swers
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2b4833d-0cf6-4494-a4ee-ad5b033b6135 · outbound
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding element” and “event
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 029bc2f4-af1c-4e0d-9ee6-6c811843f413 · inbound
Video Reasoning without Training Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.