Pith. sign in

Paper Citation Record · LEDGER

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2506.04141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04141 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:52:35.700887Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:49:50.117587Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:58:02.808545Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a16a489-26ec-4e1d-855e-a27405a9f5aa · outbound

This paper cites OpenAI o1 System Card.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.535059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.535059Z digest=sha256:061313353197684a86142888556f0a4cf0d38dfdb38573808515e55bb66c23b3

Observation 608af6c8-5391-47ea-883a-49f3e032d898 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.539038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.539038Z digest=sha256:e57710da0e9f66c5ae79de4007f129a9ae8431a111b8a0ef2d6d6708abeb2d87

Observation 696125e8-28f7-4289-bcdf-0cc6e607d783 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.542436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.542436Z digest=sha256:03155af9f0c4d54718350581d0139ab06620b1c0594279abce8b889f9b9fc043

Observation 2837dcb6-1fc7-4490-a770-0340dcad625f · outbound

This paper cites Openai: Introducing openai o3 and o4-mini,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Openai: Introducing openai o3 and o4-mini,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.184407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.546254Z digest=sha256:9b19e8f958d2ccdffeb66cc755c6d95fa81d2e8749777d4570481723eaca9cd3

Observation 38117a36-c7c9-48b2-92fb-2433210e2447 · outbound

This paper cites Research of intelligent home secu- rity surveillance system based on zigbee,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Research of intelligent home secu- rity surveillance system based on zigbee,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.177611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.549199Z digest=sha256:d9ef742011f6e08a775849fd889de032eb462f10edca61cf3dcbe798186301c9

Observation 8da11fa4-00ac-49d7-8db0-bbc9bafe1b09 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.551878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.551878Z digest=sha256:c66aa8b4a294dd9c943494216b72a0a8522df6ffa784e1ec2af07cbc51bc1019

Observation c289e8f2-a3dc-4736-8ed1-02d23e992757 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MLVU: Benchmarking Multi-task Long Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.554510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.554510Z digest=sha256:e2ae00027843eb8b8f22e34069ff2c005f10596882a736f29f64ee60fbab032b

Observation f326eb43-54e8-478d-9c96-d0d02bd77ea7 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.557682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.557682Z digest=sha256:64413fa37861bc75ca6be6eb64c365f4bb68ca2dc1b5dcf2b42b14192186b50b

Observation 815ece1e-e41f-4d65-99ae-0595086be7c3 · outbound

This paper cites Heuristic and analytic processes in reasoning,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Heuristic and analytic processes in reasoning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.560026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.560026Z digest=sha256:f6e939f939d40a443c0b23589720a859728583fa5d977c0b3e4d9adba51cc2ab

Observation e7a147bf-5864-4e56-b4ec-38a7aed68782 · outbound

This paper cites The clarion cognitive architecture: Extending cognitive modeling to social simulation,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos The clarion cognitive architecture: Extending cognitive modeling to social simulation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.167076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.562565Z digest=sha256:e5305e46142b37b53cc9f5ab05898107f424062e91c58725d86cd9613fd2c110

Observation 5b9c24f5-db67-4ecf-a81b-62c1009f6fa6 · outbound

This paper cites Polanyi, Personal knowledge.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Polanyi, Personal knowledge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.160499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.565002Z digest=sha256:3e8ad3a4a9f262fd3be6b128d75bc727b0f8e005f85fde957f85b0879584b84b

Observation 4dda2206-841f-4a8c-b824-b2ec497eb91e · outbound

This paper cites Kahneman, Thinking, fast and slow.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Kahneman, Thinking, fast and slow

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.153382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.568284Z digest=sha256:735465fa71728ff964d8525f4f07b0e801da7987bedf153dd3f3f6bc6ab2ccd1

Observation 754fa109-f05c-449f-a839-53549b917c38 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.570857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.570857Z digest=sha256:0ab0a8ccab6034c75a77624347ad924bb2af49fa26713badd3de521e054ca097

Observation 3ae94a2c-6515-4e47-baf9-9371e5a6adfb · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Measuring multimodal mathematical reasoning with math-vision dataset,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.145456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.573771Z digest=sha256:f2292798a4086fd3fa93dd7f451fd4d85f2a040dbe36a36eac13b54f8677483c

Observation edb1385a-aa8f-4b6a-a67f-e6175f0dace0 · outbound

This paper cites HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.576700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.576700Z digest=sha256:a120457295c878b9e6d6dad51f4e5bed2f617a032fd9a7575ebca914adf7f22a

Observation 93c96661-efa9-4687-8115-466c2b575ef9 · outbound

This paper cites GPT-4o System Card.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.579503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.579503Z digest=sha256:bee6429c7f57199da837e5ccc14646ac20fccc4f48fa9ad071dc2f03d75f7d66

Observation cf43c518-8b97-4078-bcae-45ad4cdfa3ac · outbound

This paper cites Chain-of- thought prompting elicits reasoning in large language models,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Chain-of- thought prompting elicits reasoning in large language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.138576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.581603Z digest=sha256:4ea841ff77add0706c5e9bf9ea0080b9519b7fa4eed6a79dccac8658312458e4

Observation c4d885aa-7179-4fe9-aecc-6e76ea3d0a9f · outbound

This paper cites RankGen: Improving text generation with large ranking models,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos RankGen: Improving text generation with large ranking models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.131984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.584452Z digest=sha256:5c29aa4d6481ea95e5956c27e5d714af5c75ca030f946821a7e701a5100c15f7

Observation 15f9f687-6393-4f1b-8cba-3fb5a58c9845 · outbound

This paper cites GPT-4 Technical Report.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.588141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.588141Z digest=sha256:6353283589ff4469861baddb208a4088dc0ee930f9da9cf7ea7e86f12de0276e

Observation dce1fa4e-302d-44ff-887f-749bc44c8912 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.590859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.590859Z digest=sha256:5872af2bb6f9754d7dabd703b89ee2e66e33ad7370c9e7156bdb7b9a033a7d8f

Observation cb9b44b8-3b68-46e1-ad0f-119a3b6d46e6 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Egoschema: A diagnostic benchmark for very long-form video language understanding,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.123300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.593755Z digest=sha256:1542f410756ed1458665f73dc989e3686d33acffea4c968f243f92846cff413d

Observation 8b7b15e2-7b1c-43da-b22d-6946373d3fa3 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Perception test: A diagnostic benchmark for multimodal video models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.114394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.597811Z digest=sha256:b5583ac25f488996409901601473c82fff028a0e95f364e950d6d9165ef99a7b

Observation dad58a08-81a7-4ecc-8a96-3cfa5e286a49 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.107625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.600619Z digest=sha256:c9dbae5fdbc50d49facc081484eeb2e292ed0df21058a2d4a7d2b6d230aa80aa

Observation 1e68601b-13b4-4b2b-b48c-1e5b8d02d133 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Video question answering via gradually refined attention over appearance and motion,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.099703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.602947Z digest=sha256:ada45fd0a5d536038fd1a58b860c1eedff77b72e7a044f50d7ac247b2e3fea47

Observation 6009dc50-9be6-4b74-99c0-4c2b46d98e3a · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Msr-vtt: A large video description dataset for bridging video and language,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.093405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.605131Z digest=sha256:b66d9fe64afe5638f4e27106f348a9557ab97c7a32c4d9f64e2633b1a598624d

Observation 265751a1-b2e7-4379-ad90-d3336e43c38c · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.086512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.608093Z digest=sha256:08e1046cd83cf4396efcbb19b27d140012968a5026caae6acc5284cc01e8d642

Observation ba051163-bcca-49a4-83cb-e9ca4aaf9e11 · outbound

This paper cites Mmbench-video: A long- form multi-shot benchmark for holistic video understanding,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Mmbench-video: A long- form multi-shot benchmark for holistic video understanding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.078924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.610522Z digest=sha256:a04d6aad02fbb2b6d58c372ea5ceabdfe755f94af64ed4d89c6e9ab224a52fd5

Observation 8e1d8bd8-16f3-4858-b9d9-64466ebf18c2 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos LVBench: An Extreme Long Video Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.612961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.612961Z digest=sha256:bf02bf39c1792f99e91352fc2fe050fc8dfa87ba7dc0bc2a7c2e1523c9bf53e3

Observation 2c07c50b-6b93-43ff-b0bc-7924086641cc · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Longvideobench: A benchmark for long-context interleaved video-language understanding,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.069252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.615998Z digest=sha256:5e744a49f18c96d8f94ce2f0f5bbddf8f5e32ab1a20cb235144d7c86ac821819

Observation 2bd4f8df-af75-426c-b290-f519c8a333c5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.618681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.618681Z digest=sha256:4d5543f265c4293807fe9fc9cb17e90aa6c8af66f13da73214d2bb5c8847464d

Observation f4a9f28e-6d68-490d-9c0a-b92f8113a146 · outbound

This paper cites Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.621054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.621054Z digest=sha256:c40324cfdeb4c62bb17ce01c11ea8f8c08994a2303767ab06d18f328b520ff5d

Observation 7477edea-e870-43d3-a06e-bbb798c429b6 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Measuring Mathematical Problem Solving With the MATH Dataset

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.623579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.623579Z digest=sha256:01e453cebd0703513063f53240aac733e2d12579621273ad26b20f2957c650af

Observation 4bf18601-602f-438b-9b73-a25056ed88bb · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.626941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.626941Z digest=sha256:acb8b8e4bd8f8074d3e30a807a32472cb6ca3330049c26d14691e15b823089df

Observation 6307fcba-3259-4113-84d5-d4ccd70da711 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Training Verifiers to Solve Math Word Problems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.629166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.629166Z digest=sha256:b17a696ed32c3ddde2fa12345554ee4cc82ce057c957d13cc9f39798dfa5967f

Observation 5b806217-6fdf-455b-a7f3-ea38eaebe4f7 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gpqa: A graduate-level google-proof q&a benchmark,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.059415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.631579Z digest=sha256:0749e363617f075fc6467d711f10bd37a0c4543a27052e525efc27df5813be63

Observation 4896c574-0961-4e46-98b0-c87fd72cb4c2 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding bench- mark,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Mmlu-pro: A more robust and challenging multi-task language understanding bench- mark,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.052675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.634582Z digest=sha256:2e4e65aa3799a1ebe67e791609bd87fa7ea5a5728ff1952dc0ea2193f2f26272

Observation 223354ca-41cd-4172-a936-0abf93f001f4 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.637054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.637054Z digest=sha256:23853fddd5a24488ffff06bda8161ceabaf03a1ac031852ab0cfb4c3c53e1d1e

Observation dc10fc38-d2e6-452b-a61b-36f266a4b83d · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.639817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.639817Z digest=sha256:b0a302be031e2dba49dfc8cddebdae6a100edf62c08719d98242d3e08119f412

Observation 0552af4a-974a-47da-b9ed-a3540cec5b83 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.642571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.642571Z digest=sha256:f30292b325b8ae6fa387eabf5cedda9a7108309875786ec7cbc32619720aa2d7

Observation 104ac23d-bed3-4f46-9dea-a9c5e2e4580a · outbound

This paper cites Lakoff and M.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Lakoff and M

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.045708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.645908Z digest=sha256:88836c47d858bdf81e519ba33e9bffbfa2d2e5a7559e54b912e1ea59835da193

Observation d56947e7-039f-4d09-8949-cd08332cc74c · outbound

This paper cites Openai: Hello gpt-4o,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Openai: Hello gpt-4o,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.036018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.648604Z digest=sha256:6ddd2f520e8616d9dfbb5490527626478da711067743a8b3a82730b227be3329

Observation 2bc0ecf3-0ec5-4795-a9e5-1282987a84d9 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gpt-4o mini: advancing cost-efficient intelligence,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.027076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.651452Z digest=sha256:6301e8e37b0ffdccfa3151b322cca0a4fc191cdf83a8481d0d27dde5d1ff35f6

Observation ab86d69c-cbb5-4932-bd99-16234070d577 · outbound

This paper cites Introducing gpt-4.1 in the api.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Introducing gpt-4.1 in the api

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.015045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.654068Z digest=sha256:6ff406dacf17323319b0a6cb8c5cf3cbdbc06dbddda94bd27e71ef22a8c7025f

Observation 6d61bee7-0ec2-430a-aad2-eeca0b8af9a6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.656456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.656456Z digest=sha256:09dcc1135312e90d850bce2cd9c96aad924801453764a25cd247d2aabb18d4b6

Observation 58c4fed6-ce9d-495a-ba14-8d44dec453b5 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gemini 2.5: Our most intelligent ai model,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:36.006568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.658759Z digest=sha256:bdc1e7008d134e4704486fcd47c6358529b79d021bc0a3e4e23e7297804127ad

Observation a769eb1e-f719-45da-b3a6-30b59e601754 · outbound

This paper cites Anthropic: Introducing claude 3.5 sonnet,.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Anthropic: Introducing claude 3.5 sonnet,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:35.997341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.661183Z digest=sha256:6a7c54bdc11fad031280f9d7b74977b29ca90b675e1ca35ca03ecbf901ce9337

Observation d5c5ea53-26e1-4b1b-8981-707069f2b507 · outbound

This paper cites Qwen2.5 Technical Report.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Qwen2.5 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.663611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.663611Z digest=sha256:58e05c19712ad86a01b3b933b546b90b422c8bdd734dad83ea5ef9e0f9fd3975

Observation 11546b1a-8c4d-4240-8fc3-7a2426ff067a · outbound

This paper cites Gemma 3 Technical Report.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Gemma 3 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.666040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.666040Z digest=sha256:e233c49d233fed1eaa6d57f982595890b579acff50b99121d179367bc8609f04

Observation a93883e8-83a9-4a11-ac5c-a00642dd05e3 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.668975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.668975Z digest=sha256:881077504ef856a0b1cfb58b6620dc3b46335ef773b5b42251b34e3436540d8e

Observation d38a5394-b012-47c9-bd6a-84ce83e0ae64 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.671480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.671480Z digest=sha256:cac46d65b34362ab68dd476ca3b20d5ba04e377524109c12c65ce859a3c5084a

Observation de3a22ae-1e91-4360-9657-788ee22deb90 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.674042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.674042Z digest=sha256:28aec80a84a06bb76a205bc0c1f4e7e49ac4777f9b2d32d56b9c616773b0f76f

Observation fe7b0446-b8e9-499f-8209-0ad8c96a9e49 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.677159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.677159Z digest=sha256:c71e24dab8c73c179100175d549c6d8940b024cf96aa6d01e16c27afc255cd17

Observation 54be0b75-8e85-4d36-95de-86a8a0f15c25 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos CogVLM2: Visual Language Models for Image and Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.680171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.680171Z digest=sha256:2d2ee58cdad58d1a92fa198c19b39102ce031a42c2881b7dbd1ef88571f43637

Observation 56673960-7564-4da5-a3a2-2610c81da451 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos NVILA: Efficient Frontier Visual Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.682714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.682714Z digest=sha256:121a5a4b100bd753209281d7c62e53ad80b264c48865d02be038e8112e94eb38

Observation df71c6de-2a7d-4d99-9116-a1fef8e437d4 · outbound

This paper cites Watch the video and answer the question and give a correct answer.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Watch the video and answer the question and give a correct answer

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:35.988996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.686243Z digest=sha256:b327cbc1da8de1d5ac0d5f6adc00fc329d7d4aae112e7bcfd5136328bd8ec380

Observation 43116a44-f79e-4edc-bdf0-8526975beadd · outbound

This paper cites High recognition interpretation?.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos High recognition interpretation?

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:35.979480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.688938Z digest=sha256:649890c873ab496dd1308dfabf4674fedd1f8faff24473c644f3365fb4dd9ae1

Observation 63ef1949-7e76-46d4-8d45-d312edc68e67 · outbound

This paper cites an unresolved cited work.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:52:35.971576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.692710Z digest=sha256:988914fa8d5293379792e73544392944c3ec03dec18284adf63669fd2894a320

Observation 082f057b-d313-4e56-b086-1fb4e4699f0c · outbound

This paper cites an unresolved cited work.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:52:35.962889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.695310Z digest=sha256:522173f33c93c950586194c3c7021afc8fb82d8184d9a69d8dd3e25dd2881268

Observation 743b6930-a177-4b7a-8ec4-30b559a41c74 · outbound

This paper cites an unresolved cited work.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:52:35.954443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.698314Z digest=sha256:c2d208fb6f49dfaf5d89188f9c701862f1964545d5648fdc80245e7905b11223

Observation 311c2a08-1cf8-4df3-b88d-180c1aff6595 · outbound

This paper cites other frame desc.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos other frame desc

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:52:35.940971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:52:35.700887Z digest=sha256:07062b641bcd59e7459aca749489df797e7a23c964787ce8e4a46ab08056d186

Pith citing papers

Observation ebecd022-1635-4ac0-9adf-90481d7d2c4f · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:06.002557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:06.002557Z digest=sha256:12bd8dc678e6a4ee6886c14d9a8ec14879c42733bca8915922505c1202c4c722

Observation 8fbfc4d0-759f-413e-a6ea-d4098a5fdc1f · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:21:18.406523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:3367a1d3a3267cceb18d3eca4c6532f331a5983c0985f798daab7758b3f6bb67

Observation 891718a3-8afd-4116-abda-c9c01826bc0c · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:21:18.406523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:b00046f49653d2b6f279b5d354a7c1457b1843ab79bbfdb8341faad8af62a2cc

Observation fc22239d-0da1-49c1-b065-5cfddb0ec259 · inbound

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning cites this paper.

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T00:49:50.117587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:49:50.117587Z digest=sha256:0e7fd4e25f8c82cf246d0737ed4b536fa7e4c40e715ebe229c7e36dd3f290ff1