Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:52:35.700887Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2506.04141.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:52:35.700887Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:49:50.117587Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T10:58:02.808545Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8a16a489-26ec-4e1d-855e-a27405a9f5aa · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608af6c8-5391-47ea-883a-49f3e032d898 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 696125e8-28f7-4289-bcdf-0cc6e607d783 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2837dcb6-1fc7-4490-a770-0340dcad625f · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Openai: Introducing openai o3 and o4-mini,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38117a36-c7c9-48b2-92fb-2433210e2447 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Research of intelligent home secu- rity surveillance system based on zigbee,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8da11fa4-00ac-49d7-8db0-bbc9bafe1b09 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c289e8f2-a3dc-4736-8ed1-02d23e992757 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MLVU: Benchmarking Multi-task Long Video Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f326eb43-54e8-478d-9c96-d0d02bd77ea7 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 815ece1e-e41f-4d65-99ae-0595086be7c3 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Heuristic and analytic processes in reasoning,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7a147bf-5864-4e56-b4ec-38a7aed68782 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos The clarion cognitive architecture: Extending cognitive modeling to social simulation,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b9c24f5-db67-4ecf-a81b-62c1009f6fa6 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Polanyi, Personal knowledge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4dda2206-841f-4a8c-b824-b2ec497eb91e · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Kahneman, Thinking, fast and slow
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 754fa109-f05c-449f-a839-53549b917c38 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae94a2c-6515-4e47-baf9-9371e5a6adfb · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Measuring multimodal mathematical reasoning with math-vision dataset,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edb1385a-aa8f-4b6a-a67f-e6175f0dace0 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c96661-efa9-4687-8115-466c2b575ef9 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf43c518-8b97-4078-bcae-45ad4cdfa3ac · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Chain-of- thought prompting elicits reasoning in large language models,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4d885aa-7179-4fe9-aecc-6e76ea3d0a9f · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos RankGen: Improving text generation with large ranking models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15f9f687-6393-4f1b-8cba-3fb5a58c9845 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos GPT-4 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dce1fa4e-302d-44ff-887f-749bc44c8912 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb9b44b8-3b68-46e1-ad0f-119a3b6d46e6 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Egoschema: A diagnostic benchmark for very long-form video language understanding,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b7b15e2-7b1c-43da-b22d-6946373d3fa3 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Perception test: A diagnostic benchmark for multimodal video models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dad58a08-81a7-4ecc-8a96-3cfa5e286a49 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Next-qa: Next phase of question-answering to explaining temporal actions,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e68601b-13b4-4b2b-b48c-1e5b8d02d133 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Video question answering via gradually refined attention over appearance and motion,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6009dc50-9be6-4b74-99c0-4c2b46d98e3a · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Msr-vtt: A large video description dataset for bridging video and language,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 265751a1-b2e7-4379-ad90-d3336e43c38c · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Mvbench: A comprehensive multi-modal video understanding benchmark,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba051163-bcca-49a4-83cb-e9ca4aaf9e11 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Mmbench-video: A long- form multi-shot benchmark for holistic video understanding,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e1d8bd8-16f3-4858-b9d9-64466ebf18c2 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos LVBench: An Extreme Long Video Understanding Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c07c50b-6b93-43ff-b0bc-7924086641cc · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Longvideobench: A benchmark for long-context interleaved video-language understanding,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bd4f8df-af75-426c-b290-f519c8a333c5 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a9f28e-6d68-490d-9c0a-b92f8113a146 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7477edea-e870-43d3-a06e-bbb798c429b6 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Measuring Mathematical Problem Solving With the MATH Dataset
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bf18601-602f-438b-9b73-a25056ed88bb · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6307fcba-3259-4113-84d5-d4ccd70da711 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Training Verifiers to Solve Math Word Problems
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b806217-6fdf-455b-a7f3-ea38eaebe4f7 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gpqa: A graduate-level google-proof q&a benchmark,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4896c574-0961-4e46-98b0-c87fd72cb4c2 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Mmlu-pro: A more robust and challenging multi-task language understanding bench- mark,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 223354ca-41cd-4172-a936-0abf93f001f4 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc10fc38-d2e6-452b-a61b-36f266a4b83d · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0552af4a-974a-47da-b9ed-a3540cec5b83 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 104ac23d-bed3-4f46-9dea-a9c5e2e4580a · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Lakoff and M
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d56947e7-039f-4d09-8949-cd08332cc74c · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Openai: Hello gpt-4o,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bc0ecf3-0ec5-4795-a9e5-1282987a84d9 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gpt-4o mini: advancing cost-efficient intelligence,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab86d69c-cbb5-4932-bd99-16234070d577 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Introducing gpt-4.1 in the api
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d61bee7-0ec2-430a-aad2-eeca0b8af9a6 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58c4fed6-ce9d-495a-ba14-8d44dec453b5 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gemini 2.5: Our most intelligent ai model,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a769eb1e-f719-45da-b3a6-30b59e601754 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Anthropic: Introducing claude 3.5 sonnet,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5c5ea53-26e1-4b1b-8981-707069f2b507 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Qwen2.5 Technical Report
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11546b1a-8c4d-4240-8fc3-7a2426ff067a · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gemma 3 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a93883e8-83a9-4a11-ac5c-a00642dd05e3 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38a5394-b012-47c9-bd6a-84ce83e0ae64 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3a22ae-1e91-4360-9657-788ee22deb90 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7b0446-b8e9-499f-8209-0ad8c96a9e49 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54be0b75-8e85-4d36-95de-86a8a0f15c25 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos CogVLM2: Visual Language Models for Image and Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56673960-7564-4da5-a3a2-2610c81da451 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos NVILA: Efficient Frontier Visual Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df71c6de-2a7d-4d99-9116-a1fef8e437d4 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Watch the video and answer the question and give a correct answer
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43116a44-f79e-4edc-bdf0-8526975beadd · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos High recognition interpretation?
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63ef1949-7e76-46d4-8d45-d312edc68e67 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 082f057b-d313-4e56-b086-1fb4e4699f0c · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 743b6930-a177-4b7a-8ec4-30b559a41c74 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 311c2a08-1cf8-4df3-b88d-180c1aff6595 · outbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos other frame desc
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebecd022-1635-4ac0-9adf-90481d7d2c4f · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fbfc4d0-759f-413e-a6ea-d4098a5fdc1f · inbound
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 891718a3-8afd-4116-abda-c9c01826bc0c · inbound
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc22239d-0da1-49c1-b065-5cfddb0ec259 · inbound
Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.