Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:22:56.139701Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 11 inbound Pith citation observations for arXiv:2506.10857.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:22:56.139701Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:23:08.702909Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:49:17.495045Z
100 of 122 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 67bf88f4-239b-4a02-b712-e4ca6827f655 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e9ab553-3fcc-48c8-9150-e56adf8fecbb · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd844921-0989-4044-aa69-dbffdb445ed8 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos InternLM2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3c841a-826e-4063-b441-59d7aee0678b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db6db4af-9926-4f46-9e83-3b382a3516cf · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos TheoremQA: A Theorem-driven Question Answering dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9acc9133-3f9c-41c5-a370-c8ca38e496ea · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Autoeval-video: An automatic benchmark for assessing large vision language models in open-ended video question answering
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eef8657-de74-4207-a670-d3d40a65939f · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee2aabb-d15d-4471-8307-85557ae27857 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7009cb-616d-44a6-bd44-2b6e59c6e59e · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75212c6-a201-462d-8cf5-5bfabcb1e49c · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ddbae0-b8c6-41f9-935a-b57fe314946b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Deepl translate: The world’s most accurate trans- lator
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a950fbf6-7e2b-427b-a098-4b7e970c0926 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Gemini 2.0 flash thinking
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3f8b0f-833d-40b8-a89f-00cee31d231b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Mmbench-video: A long-form multi-shot benchmark for holistic video under- standing
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9675d94b-f65d-4fbb-a217-82b0059f9647 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 753a87c7-a443-46e3-9bce-90ba8d4e2f5c · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Sciknoweval: Evaluating multi- level scientific knowledge of large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0849dd15-06d7-43fd-adfe-969771bbac93 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2695ae24-3a76-4e1f-a39a-35dc60f54fbe · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos H2OVL-Mississippi Vision Language Models Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4859939-007a-444d-aa0b-c725f89ac507 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c58fbad-de91-4d02-b625-61d4030b97fa · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f68f3c2-48fb-4c88-9dc4-c9e4b0a985d6 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f4926b-9b3f-4e21-967e-b95936779a4a · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Measuring Massive Multitask Language Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b5364e-be60-48e2-8e6d-77a1ca0006f8 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495389a9-ed4f-40c5-a39a-4467210737a0 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0582dd-d1da-43cb-af60-a13f67595808 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Vbench: Comprehensive bench- mark suite for video generative models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251dc541-9fc2-4f9d-9e0f-43a43d65d45b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Olympicarena: Benchmark- ing multi-discipline cognitive reasoning for superintelligent ai
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb48fda0-189a-4e90-a294-989e03148d7a · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172d0c3b-4e7a-4db9-adac-bce723c074a9 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eae764e-27d8-42d7-a2e7-8a04c60c7397 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos TVQA: Localized, Compositional Video Question Answering
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd312b7a-59eb-4997-941c-a8f1f084af16 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Veu-bench: Towards comprehensive under- standing of video editing
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdeedada-c8a1-4732-8f62-8830251ce4c6 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76134e5-71df-41ac-80c0-b4b20c61c9f5 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464c5e5e-07f5-46f0-ae02-fcf6e2b9ff09 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a7f6a6-d097-497d-887f-4378e5386efb · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72944cf2-74c2-4dd2-91fa-073fcf1672fa · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d67a2a2-fdbe-4af4-ba1b-bf14d005537b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehen- sion
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 777b0caa-5f9e-4163-8fe5-9811e7a48d40 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3686116-b77e-4086-86e6-93a127e3afbf · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Vila: On pre-training for vi- sual language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a99a1ed-4c12-4cf3-b715-9a75287a794a · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos DeepSeek-V3 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4aacb62-809e-4ff1-8917-6c9b2db81e35 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos TempCompass: Do Video LLMs Really Understand Videos?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca08579a-f769-47a9-87cc-a7bb7a4e650c · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea1b2c88-4fb8-4786-841f-074d073c603c · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Llama-3.3-70b-instruct
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cafefa2f-9f66-49de-9955-28857541679a · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc8e262-74ec-44bc-ae1e-02b2b7c087ab · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4288717a-c077-4503-8d0f-6b943dc16086 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4dd8165-bd5e-4613-a574-b558d5fe57e6 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Plotqa: Reasoning over scientific plots
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1ab0a5-8d56-42b7-abc7-9ab6202c0e16 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e2f9bc-4a34-40b3-b521-86f6f1ad70c6 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Hello gpt4-o
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9514c253-91e1-4207-9685-bde03caf9b22 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Introducing openai o1
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f039a56a-1291-4aa6-8310-e8a13a0b95b8 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Openai o3-mini
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8356cf3-7c1c-489f-b74f-cd3836ef58e5 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Per- ception test: A diagnostic benchmark for multimodal video models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8456456e-a978-4287-883a-6bb39db100b2 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Robust speech recognition via large-scale weak supervision
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248a203e-e894-4ac2-9ce7-d528e407e952 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Direct preference optimization: Your language model is secretly a reward model
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bce9224-a4cb-41a3-853a-88f33fd1a58b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0403ef84-b2cc-4ba8-b7ea-606cd2b175b1 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Scienceqa: A novel resource for question answering on scholarly articles
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7567d2a-afe6-49ac-87b3-75bdf9b6cba0 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Proximal Policy Optimization Algorithms
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67660f1-c14d-4565-a888-d47027a2f72c · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Moviechat: From dense token to sparse memory for long video understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91aa0012-efe1-4a8b-ba32-15907154c07b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Scieval: A multi-level large language model evaluation benchmark for scientific re- search
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20738237-2470-4115-95df-50ad77de4008 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Movieqa: Understanding stories in movies through question- answering
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c0df03-ede1-4de7-8286-ca5772f806c0 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Claude Team
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61343ccd-0658-4382-8a41-57d7baedba20 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MiMo-VL Technical Report
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc8315a-a995-42eb-a74d-b37f89622c28 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Gemini: A Family of Highly Capable Multimodal Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223a5ac8-47d9-4218-9aed-c491e5c2d770 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Kimi-VL Technical Report
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 574deb97-d5bb-408c-a6d8-0403b5d74e9d · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Kwai Keye-VL Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ba791d-667a-4b4d-a2bf-a7ef871bb11d · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Qwq: Reflect deeply on the boundaries of the unknown
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc786e6-eb01-4356-9800-19c9bfa5dc5f · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Qwq-32b: Embracing the power of reinforce- ment learning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 998bdf44-fffa-4ad1-a2a1-5fda2aa5c176 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d08530-1010-4849-ad8f-3a52100a62b0 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Mea- suring multimodal mathematical reasoning with math-vision dataset
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a733ad3b-02c9-4ce9-ab71-cf3ee71ca713 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61e18b8-a4b4-41a3-9831-f60c6e691041 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c63bfad-d6e4-44f8-a6d1-0365e150803d · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos LVBench: An Extreme Long Video Understanding Benchmark
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7721589c-31dc-49cd-85bc-14090328f51d · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Internvideo2: Scaling foundation models for mul- timodal video understanding
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35576631-0d20-440c-93e9-b78340cdeae2 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c1bebf-3b71-4951-8ed2-20c1839bf3c5 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70a7195-8eeb-4046-a743-304badfbdd03 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Charxiv: Charting gaps in realistic chart understanding in multimodal llms
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c7b20b-7eaa-440d-a569-3db0431d1f06 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a8e578c-f3ed-4db6-b45f-5eba1d035eca · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d024e8-859c-4f80-b3c1-2e52fbef5d7b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a6e6e0-2b1a-4f35-8642-cb8ada73e42d · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9615c86-c7a2-4a34-872b-ee5f4d09608c · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c6a477-35c9-4891-a79d-16c5ff0e1c75 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Next-qa: Next phase of question-answering to explaining temporal actions
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f499c913-d804-40f1-819d-bf2eaee718d1 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e523dc1-8827-4b22-a657-8ce90271cccb · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Qwen2 Technical Report
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf03061-74b5-4672-9873-f22e3c60fa29 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Qwen2.5 Technical Report
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63efaa54-d18b-432f-8376-97fd9af0a511 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Vript: A video is worth thousands of words
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a02474-e2ec-4826-bc43-a1a96178e9c0 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c644f51-4323-42f6-9942-e752d92768b3 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3352a403-d713-4bc4-8714-334ec1debcbf · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 425c0e8d-d767-487e-9278-d7c6bfd58ae2 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aff777f-413d-4142-9836-e8741bc8053b · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636a398c-4161-4ff9-b1c1-fda5a06ae8a4 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780e01ce-6412-459e-9e7b-5403e03752ce · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Long Context Transfer from Language to Vision
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 737f9633-f141-463b-b19e-f15e39a533b1 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae769e7-ed42-4d67-b01c-f74bb405cab7 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MLVU: Benchmarking Multi-task Long Video Understanding
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759fdcf6-4d27-45f7-b53b-2c2bedb2cbd5 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Au- toshot: A short video dataset and state-of-the-art shot bound- ary detection
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b0cd36-2a4d-4ce0-b797-719a13978b05 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos • Multi-step process 1: A → B → C → D
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2abdced-fcf3-4473-80e9-510a730d7196 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos • Multi-step process 1: A → B → C → D
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054d7390-470b-43cd-8851-e4691ff87b5e · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos • Multi-step process: D → C → B → A
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bd4edb-813b-4535-9876-4ea20b8ee2d5 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos • Multi-step process: D → C → B → A
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee36e1ad-397c-49c7-8fa2-1440431e6943 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos • Multi-step process: A → C → D → B
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1081268f-4dd7-4e6c-a68f-7147abad3086 · outbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7cca12f-0ef3-48ae-8da9-24c49c949add · inbound
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f4061c-be93-4ade-b1a7-a088468f27ca · inbound
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca25f97a-b15e-4692-acc8-dac577a40f42 · inbound
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a3b076-d662-48e9-9632-ba373fc1cdd5 · inbound
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80897de9-8c46-48d2-a396-247283b3bc12 · inbound
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 606a425b-9f9f-4c25-bfa8-494342b57fef · inbound
Towards Temporal Compositional Reasoning in Long-Form Sports Videos VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78e151e-b2dd-460d-a4f0-a06fc516a95f · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8993290-6c4d-4334-9dc6-93555f599665 · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8827aa6c-7368-48d9-bb61-1f90ba55de1a · inbound
SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c116fb8-6802-400d-9428-98c7b1178a58 · inbound
StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97e2e6a3-bea5-46fb-bea7-8b7f1ff6b6a4 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
Reference 279
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.