Pith. sign in

Paper Citation Record · LEDGER

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

As of 10 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 100 inbound Pith citation observations for arXiv:2501.13826.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13826 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T00:32:41.059558Z

measured 162 of 162 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 106 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:23.214406Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact26
  • verified fuzzy31
  • unresolved3
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation d1731bd0-fc5a-4283-9187-75745f8a10c7 · outbound

This paper cites Claude Team.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Claude Team

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.359065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0641464911587a55e0688022b9d4d8f593f91bad243113f798ad9d758c36a12a

Observation 0103abb2-85af-42fe-b43a-7fd3de990e89 · outbound

This paper cites A systematic classification of knowl- edge, reasoning, and context within the ARC dataset.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos A systematic classification of knowl- edge, reasoning, and context within the ARC dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.369890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:a38578b4937c4d4603970507e12db3981afdb5bd9ad8dffa056e2a85991a9bf5

Observation 8728683d-c1dd-4cae-8913-cb685f76a829 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.302744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0d11714a30b6ec63448bff3e1a7ec2e4d5287e399d02e0eb5ed21db3a2056ce1

Observation 68652cac-f1a8-4700-8de1-5f779442999b · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.152534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:830ca65b052c5b63dabcebd012313378d79174dde9aceb32da12fe7e996925d8

Observation adcbb1ca-eb72-4683-9220-a3db179f0bc8 · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.157897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:73f401a7a87dc0e64f09971a6e33d5bb0ff57ce6fd9e75112654415d611b531a

Observation 499f417f-6454-400c-b216-2e2203827930 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.163874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:35b7160647423d8bb0ba1cdabc7f85663ff3a6e4ea78c0c364f8f1e69371773c

Observation ec1e6ef3-fe23-4b98-b54c-6247ef07148a · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.214239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:2a3e9d1397c2ec882d34ed515d9e4e3fb0e7ddde5381f583ac925e1a51e79b6d

Observation b211095d-f3a6-450a-995b-ace5fa0509a1 · outbound

This paper cites Bloom’s taxonomy.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Bloom’s taxonomy

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.377401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:3cee8a621549ec2a6784461fa2de1300eeeb57695d729e515e4d4bd950e309f9

Observation 6a3db641-70ca-4bbe-8e76-45f13f9b81ef · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.145557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:ac1feb9a3592f66c37612b296712feddbd826bc3efee1ca5f5c853ae3380639f

Observation 634b36d2-6e64-41b8-8fa4-8a709e07d1f2 · outbound

This paper cites Knowit vqa: Answering knowledge-based questions about videos.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Knowit vqa: Answering knowledge-based questions about videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.380996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:ffdaa9d79f9081cf0d17bd18da843fcc985a528a1b459cc72707908efedab0d5

Observation e40648bb-8a65-4324-8efc-706e63b9e0b0 · outbound

This paper cites Exploring the video-based learning research: A review of the literature.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Exploring the video-based learning research: A review of the literature

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.385621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:f09d3e0d5b7c1ccf5221e8531d3430044411a4295abce5877aacad5b5c6e2ec3

Observation 5995dafd-53b9-4ea0-bb20-296df47c3427 · outbound

This paper cites Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.389614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:95abc976d6023d8ac02ad7d6e0daf3090ce28fe0d2ef69130ad9f6ec557f57cc

Observation a70dfd21-e07c-497d-9544-5d3012bedbcc · outbound

This paper cites Measuring massive multitask language understanding.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Measuring massive multitask language understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.393152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:ad6d42408f251ad8b4d31d77d52f12577868e164ebda4373f0b25b2f60bc4f95

Observation 56bea1ca-668e-4db5-862a-4ad84f6da898 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.232391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:fc68773f3dc0bbb1aeee8efa29f4c68a28820ea562e163969a3fd423a55c50db

Observation bc0b2524-7f91-4dc2-ba18-60771d57269b · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.237653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:7585aa401b9fcb2a78c2441bf1a0cb188fd7d7957d52d43e526298a0f76f8030

Observation c5f32c4e-ddd6-47aa-8090-8408a836f07f · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.254763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:ae5137cd44906ecbdeeb48dbb267157eec78ea3cdb5d9c1c8323db8e284cddb4

Observation 915aa26a-9de4-43ef-8e4d-801e92bafa3a · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.397179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:255eeb5c54f5003138ab4d58463c1c3102eb35b7d97efb092b4b588068e8975c

Observation 59ad81f0-88ca-41c0-8374-acbe6a98c6cb · outbound

This paper cites VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.123945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:b00bcaa052f43ed68c6912c00869a4ef499972aabfb7c8288ed1fb556aec679e

Observation 588643b7-9054-49ca-8bae-8de9a918f330 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos VILA: On Pre-training for Visual Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.131367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:c43deb7b2eb5bd0b52d4e8a1582d288209c4e145c2209107e7706280822aa672

Observation 3a2498ba-9a00-451f-a231-ac83e85b75f2 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos TempCompass: Do Video LLMs Really Understand Videos?

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:f15a7581319a21cae5f1da181155315bc7c2a966f19590d2161a7435054c11e8

Observation b0e2c3a6-1603-4f3d-ab28-3152b6d0830d · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.401915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:38f8f7fad563b06b5ae259f7c6b6d6ccdeb9203ea548e4903ccdaae26cbe5db8

Observation 7cdd6559-5e27-4bc0-aa65-028447b70a19 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.406667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:a4b699c72465fe68fbb487517cec92b2ef942c897114a8efb4e4abb8d60530a5

Observation 430b8443-435c-442c-8373-0ff0ab6d710d · outbound

This paper cites Llama 3.2: Revolutionizing Edge AI and Vision with Open, Customizable Models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Llama 3.2: Revolutionizing Edge AI and Vision with Open, Customizable Models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.410646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:cc4deff532b00f3fc70c14c09e7957710b7a4ef601b30026be67525358e8fda3

Observation 4083582e-e888-40b4-ae0b-4be05929bbf3 · outbound

This paper cites Position: Levels of AGI for operational- izing progress on the path to AGI.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Position: Levels of AGI for operational- izing progress on the path to AGI

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.414083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:d2da8297649fd978b58917e7613ce3d322f1ca04f2d8589827de8db4940f4b1d

Observation c06a4d2c-520e-4fbb-aff2-933b4f80fb63 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.171272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:612ac8639b868b7679c2b1b26f97de789db355a8267e1af1309f73e1ff4ec4a8

Observation 2eb31417-79ae-49f7-9ffc-2a8a39dcc7c6 · outbound

This paper cites Introducing openai o1.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Introducing openai o1

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.418401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:4b04bf7dedcbdeb6bc43fbb4a4ece58c882c0791261becfca61e33371c33101e

Observation 1ffafa23-a8d3-4939-848d-3f5d1006a3de · outbound

This paper cites Hello gpt4-o.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Hello gpt4-o

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.423009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:ab3dc51be597586a15a12a1f6ac7c5cd3ffc44e55d4bb07328ada1ab7857c789

Observation 9404d1a3-bec6-4ce9-9a44-35ce9b13fc2c · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Per- ception test: A diagnostic benchmark for multimodal video models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.427044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:4edad8d2ba62f17fb4c52f04ab1e1735963750380ba6df59ea1db635e2e0d4ae

Observation 07987b6a-79e5-4c70-adfd-fdbf8e344ddf · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Robust Speech Recognition via Large-Scale Weak Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.244010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:823496efe08ef57a6e7f2218331fcc4ef204d8c86181632e3a4d301116add4af

Observation eededdb2-bd5f-4664-9a1e-58e7b16b11a6 · outbound

This paper cites Video- based learning (vbl)—past, present and future: An overview of the research published from 2008 to 2019.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Video- based learning (vbl)—past, present and future: An overview of the research published from 2008 to 2019

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.431485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:58f40b6b0e6e8cc6f01509e309dcbdd47ac362193945868ba32dd33daffdfc02

Observation c87f1033-5839-4a62-b3a7-7730ec1d7bda · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.261120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:4983317d424bbc5182b05283a6ced5340cb15207edc112ad348cd98210c9a326

Observation ffd3e082-84bc-491c-8f1e-16c2db43c57c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.271297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0f3142dfb9a9a118940f13a4ba56122a0a7062aa9d4d734b5c3ce179983bc787

Observation 0f0ff1f7-be7f-42cb-bc7a-f9921acf35de · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.282133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:e484facc37bd52d82cf4d6daedd45416be31d6ce320b3ee687915bfb3f167496

Observation a65d6845-9f16-4523-aea4-6ced7e032767 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LVBench: An Extreme Long Video Understanding Benchmark

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.239348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:8f75cfce497a07c0773ba4d03ffc3725913cddc6bc4cb939fa92f37e400d0ab9

Observation e394b953-b258-4a47-aed4-17f5205795cb · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.436260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:b55b102d910c1b6ee2eda35ade5262b92d05b9e9d7abbdc1ca8a1821bb7bdfca

Observation 95cf7e11-30a1-4de7-bf0b-3c86ec5b9bdc · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.308387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:470685a8229998bfff5b0994bcbe583c6ff6dc97407e94d8895a05654dfbf664

Observation 73d0fb67-fa4f-4517-ba83-ea220dd15072 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:12.519784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:3931bb2b1aa0dee0205a55f9e4370786bd094c1f5a23217959d47f7515bb8f4e

Observation 9c7ea373-073a-4463-8849-0d687316e734 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining tem- poral actions.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Next-qa: Next phase of question-answering to explaining tem- poral actions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.440952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:8be20cf67133aca3c3196a473a394f94e0d03d1f44e64c9b6cec9dd81847b70c

Observation 09cfd761-6a95-423e-a405-10bc3d937db2 · outbound

This paper cites Funqa: Towards surprising video comprehension.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Funqa: Towards surprising video comprehension

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.450986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0d80effd075da4358726d8650493f27e077ac41015176d6541d6f04e5686ae24

Observation 78f3424c-fce0-4833-813d-99694014b486 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Video question answering via gradually refined attention over appearance and motion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.455373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:1fcb008ad587277ba5034233e2093970d8a82337664e9165aa5bb168c66935e4

Observation e6d6e0df-a896-43d1-804d-d341cc1e4878 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Msr-vtt: A large video description dataset for bridging video and language

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.459075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:780e186cfc8162614afe0d67e77fa135c3cc06d5b4bcb012c8592f897cb8bc34

Observation c4a27a04-9bb7-4f9a-bad5-ab798e268567 · outbound

This paper cites Tenenbaum.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Tenenbaum

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.462790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:8bee407a692f757b38faf8436ee357ccacc9fa5bf4234f5fbe1aba3857b3dafc

Observation 18ecb92a-f00a-4f39-8dbb-25d5a20f45ae · outbound

This paper cites The state of video-based learning: A review and future perspectives.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos The state of video-based learning: A review and future perspectives

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.466793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:94cecc768d97b6749afd078a25bd1865fef7e8e2a71949410187de0a48cd7255

Observation 5919365b-02e9-4e07-b197-b421e87e2b1a · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.470699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:aaa78e640561db0364bdf4c5d69a4175559e64c22d6d43d4fb9b5609d6d6e573

Observation 80b0e1f4-62f5-4e60-a75e-ed7a37b914cd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.474697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:95262e907e6cc5e4ed95e145086912b4287c80dfa4844bd3a35c93b8f7753289

Observation c4e8f526-35d7-4824-83e4-1d5975a085a1 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:51:48.633687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:d1974ead49553827a7db49db71b70afb0697ee3fa44f9f51ce4f3feecc5c4d70

Observation 8525a6ab-88cd-4869-90a4-a9bd1a23a7bd · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:22.513527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:495a51fd8364d5b3876704ea3131128d71714d00b8fc5b84f3b6365eb133ea9d

Observation c28d9e8e-ea83-486e-a407-c8508b62acbc · outbound

This paper cites Long Context Transfer from Language to Vision.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Long Context Transfer from Language to Vision

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.190008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:40a3c61c2bf7b80b9687566bad50d6810d18e213c9b21c6e74bd8496c1a8edca

Observation 0e5a4491-57c8-48bc-8d32-d175db17ac2d · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.197198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0473f97d6d56c17c2a76932056bc4db313223cec64406890fd902cb368781102

Observation d9882c54-9125-4896-870d-a8a50bc149c4 · outbound

This paper cites WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.207259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:176200342e5deab49c9e377fd71cbc286e2275625543e930f0a2b3a733b220e1

Observation b625a5af-0660-47b0-bf93-357a2637536e · outbound

This paper cites AGIEval: A human-centric benchmark for evalu- ating foundation models.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos AGIEval: A human-centric benchmark for evalu- ating foundation models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.479053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:cdf4a64b03290f5213016b5a11aa9d7980e747dbd464500f6c0206b163d7eabd

Observation 783b2a2c-8d75-4270-a132-88339aad9d51 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MLVU: Benchmarking Multi-task Long Video Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.650724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:b5a0d2d81081f66e75674ac784760afbb638dc5b34a8cba5ff5f1aadfe199e6e

Observation 6135b1d8-db26-465f-8541-3de4210e528d · outbound

This paper cites Towards au- tomatic learning of procedures from web instructional videos.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Towards au- tomatic learning of procedures from web instructional videos

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.312363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:6b80b6e336a65cf8759deeb5118bfe874d7ef97045f494b32d90184adea02048

Observation 76f97e8b-fff3-4f23-9609-cdf66b036536 · outbound

This paper cites Subjects categorized under six disciplines.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Subjects categorized under six disciplines

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.316163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:990f9b84d7e96ad908e3658e2e8c201eb540f1b07bbf633297d564ba5337fe12

Observation 1bdaac07-7ced-44f0-b9d6-6a24267829eb · outbound

This paper cites an unresolved cited work.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-14T00:32:41.319941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:9532cf8fcdd0a70b499ce890e23af19bd2f3321b2733ab26fb26b8701f098868

Observation ba349dcd-e51c-4ada-bcf2-16c57c3a6072 · outbound

This paper cites The ∆knowledge metric reveals a gap between human ex- perts and models, particularly in their ability to learn new information from videos.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos The ∆knowledge metric reveals a gap between human ex- perts and models, particularly in their ability to learn new information from videos

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.324113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:1009c777daadfb72e43b028414b7ffe00f00266587140bf744f6639952625d07

Observation 604b140f-ea30-4fba-9eed-e66aa98bab85 · outbound

This paper cites We introduce the prompt as shown in Fig.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos We introduce the prompt as shown in Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.328133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:705c33196c9fb983e153fc6a33882ae231c7a808be89de335ee232571527f571

Observation 43e47588-beb8-4e79-a037-a249a5b4c3b4 · outbound

This paper cites an unresolved cited work.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-14T00:32:41.331873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:0b8db97f032b208b51c9755abc78ec5d86911c265693f5e97151c2deee685e58

Observation dbbdfb37-9ea2-494a-8272-128dbdafa546 · outbound

This paper cites an unresolved cited work.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-14T00:32:41.336735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:18b6d5a426e5e1110e976abfc8941c46fb614ffb75f9f04948b59f6baca629e7

Observation eb692f5b-9469-4d6f-8f12-39cff4aae646 · outbound

This paper cites We begin by examining errors made by Claude-3.5-Sonnet [1] in the Adaptation track.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos We begin by examining errors made by Claude-3.5-Sonnet [1] in the Adaptation track

Reference 60

Resolution
malformed identifier
raw_fallback, observed 2026-05-14T00:32:41.345301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:9de9ed756a116e3e86be3380e6f7fd3c4ff16402c1d2df732c1f4457f5c37c6a

Observation a7a853d7-670f-44e2-a0fe-13c299240a80 · outbound

This paper cites 17 and Fig.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos 17 and Fig

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T00:32:41.355104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:ef82d3e97128a04ac925b07c929cec083b3a79f4fc25e8aa17f012887faf1294

Observation ef256c8d-5598-4cde-9142-ab530c034b35 · outbound

This paper cites reason".

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos reason"

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-05-14T00:32:41.364605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:6bb062e3b2a84559c4618b24ce713051668221b9968085cfcdbb7831797e1e53

Pith citing papers

Observation efd0b836-425d-4aa3-b5a7-ea57af6c3468 · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:25:19.057503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:a9dd343f4d5d3f432c0e8f97dac79d12e3e8bba6492ecd34afbcf2810ba121c7

Observation 9fc59b67-919d-457e-843e-70be6809261f · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:12b0063941e9880f0a822107277086bd3e76b2710e7e8b6f3c4f89ad9495bc49

Observation 8b90fa82-393e-438c-9a3a-e2da6583651c · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:1752c799755bfcc7d2380d500a91a42437161e6c35555c3e9ccbe85f8458144b

Observation 2bf9ee16-4a31-4cbb-a89b-bbbe64aa3707 · inbound

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models cites this paper.

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:23.214406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:23.214406Z digest=sha256:67693ed57e5d433bdacee948ab2b25cf669041d401812ad84b363901789a1bce

Observation bf0f1754-3b0e-4d04-9148-d9a3824fac12 · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:40:56.099484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:a00a8bf38183b082c9e20759c12b76a259c6b502af10068ad737bed1ba0dbfcc

Observation cf25da3c-0d03-4627-8fd6-25d567b1309b · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:26.488247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:26.488247Z digest=sha256:a9a2eb14fa206df9b0fc655a40315eb49a36ad399794ad925021d2318eb20d88

Observation 69b421ce-8c08-4917-abe0-bdfc3944b2de · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:02.095752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:02.095752Z digest=sha256:addef6ffa71fdf1c51537190de445bce0a5eccf844c7c95c8f757b7dd49edcc5

Observation 5a3b0e9d-49b5-48eb-889e-a442eaca222c · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.956199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.956199Z digest=sha256:40a650b1774f84601e9468a3bb4df93b4d35dec77c6b6cbe77aaa308494fc7ec

Observation d211c42b-33cd-43bf-a248-ebabe66b7dc9 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.095997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.095997Z digest=sha256:93693a1cf7a0535936fcff6832e485b34d462ed281f14bb121b124b7644f5204

Observation 9737389c-37da-4dc9-96bb-e87c317bf88e · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:10.911718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:10.911718Z digest=sha256:5ce09e834d22ec2afc24dc784f1ebbf10a30505b7e4716b0684be0acec062c4e

Observation a6155ea1-b6fa-4f25-a4e4-e83f48302a82 · inbound

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos cites this paper.

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:36.257644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:36.257644Z digest=sha256:ea49b7c2128cb255955484a3331340d7741cbfced7f9a45d786401d682951b46

Observation 83866043-2fe4-421c-baad-5889f5fab71d · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.873858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.873858Z digest=sha256:f69eacd290c89dd9a695848e58bd9be94915fdb2edc74451a6c3587ffbb934b5

Observation 3afe48cc-0bfa-4fef-b495-e4872d0f038f · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.116523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.116523Z digest=sha256:db10c55f24353b05dbe0b55b67cec95448e968d79f7dd582662aa1e0a58bc0e5

Observation 23b5364e-be60-48e2-8e6d-77a1ca0006f8 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.863069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.863069Z digest=sha256:4c08eaf1a4406a02f34a1bab5f4fe32653d295138553d989aa1dd8e0716e8072

Observation 7e2d0813-3723-406c-9f37-99613b9ef861 · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.556314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.556314Z digest=sha256:022a3bc7e4d7bf015ffc4b7a9459235cc952f14344844e7ce08bdb8e83e084ba

Observation eb3617e0-da3f-4432-a830-7d9fd7b958eb · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.298315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.298315Z digest=sha256:f90ef37097e10d1ab93e863d7e17145d5c686e231443c4877975de5e2ec69db9

Observation 6ec607c9-c843-4993-88d0-adafbf7c95ac · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:52:07.912263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:167b35f4286b71ad878621236ee33f3697695a5323262475f8618b3368b6ef9a

Observation acb10c26-4472-4002-b820-f8ac1d9702d1 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.303562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.303562Z digest=sha256:37afd2192627403dc0dd4f652f23d005626a56365b70496835478875dc7ca73a

Observation f6fbbd43-c890-4ca8-a9d3-fa6583e59791 · inbound

Position: Reasoning After Perception Means Reasoning Without Vision cites this paper.

Position: Reasoning After Perception Means Reasoning Without Vision Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:02.209054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:02.209054Z digest=sha256:47e0dad9de7df077fc0b7f26f592ff17d619ba83d6c090f262fec0bf3407d3ec

Observation cf07f444-28eb-4b9c-a345-6ac8b2ac5981 · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.338467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.338467Z digest=sha256:9b45a90463e26dd942feaa0cedbca6aca38479a660d195242c727dca2668b916

Observation dc331370-fca2-4f22-95be-38d86ae9cd9d · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:28.623907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:28.623907Z digest=sha256:963006f09a5e4a2100485c31fb51cd553dea1c164bb4a709501d057061da12ef

Observation d32cec37-b735-4122-b313-e80965f80cf5 · inbound

FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding cites this paper.

FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:06:31.824530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:06:31.824530Z digest=sha256:961cdd573e6db2c7c351b80116ded113a21c90678fb9eb18e3fc73ab79326556

Observation 11e48a79-c945-4f94-a423-e6c4edf28133 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:06.005845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:06.005845Z digest=sha256:9b4606570f5cecb3b981fabccb713724ea4923091e9fffadd99d374a48360151

Observation bac3de43-4733-4cfc-8917-fae8087cf7ea · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.659106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.659106Z digest=sha256:4e522d473dd65effcc023d85ab719f628763d4abf3512d7dc7875b1277fa01ec

Observation 58847d44-65cc-47b0-83c6-4c87767cd765 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:30.702457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:30.702457Z digest=sha256:c3d1df4a57707e4ddc539c20b48ecbba52d121194585f0f396eb336e24bfd768

Observation b57281fa-1794-4517-9b12-31f4fb6d763a · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:14.160705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:14.160705Z digest=sha256:c58a7082749a65db43e398924e6892d230ddd477db2758597ea172d199627fa7

Observation 43d1e47b-5fec-4a95-8a97-da2b40f971b2 · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:07.603207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:07.603207Z digest=sha256:d3e3d9b8ddfdb6cb9aca58457e818473ac58032537275719d0df1e38af08f7a2

Observation 81837d7b-8245-4805-bf7d-2fa9af697a9f · inbound

Cambrian-S: Towards Spatial Supersensing in Video cites this paper.

Cambrian-S: Towards Spatial Supersensing in Video Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:46:04.732709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T03:46:04.363500Z digest=sha256:b059a25ebfbcc102c986ba88460d4649ed2dd53ec353f77a2f9b0166583ff7ad

Observation 58b52126-6b6d-4af6-b7c0-a2b0dafc457f · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:25:22.678636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:cdddb61acf7273921362d0c3854494fbbf62db27aa2e97fa290b27424c6d5d2a

Observation 6d4f1304-964d-4c0f-8f18-bc54829ac23d · inbound

Boosting Reasoning in Large Multimodal Models via Activation Replay cites this paper.

Boosting Reasoning in Large Multimodal Models via Activation Replay Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:09:03.966128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:05:48.682057Z digest=sha256:4f29bced158bfa6f168bb545a2ca92a94888792d060e6175b90b4b4dfa1d7fe6

Observation e7c55f0d-303b-4a64-9a88-1f4fcf08e9cf · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:31:32.096789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:39dde3a76ab86fb14de87ba07f6ffdf3d540d61ceebd53688452383355b04632

Observation fb5536a7-0eee-4369-8936-f4e7a9f9c40a · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:11:26.633941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:79d038cbbd88e32578ac74b2a5840e0c3c6d370d8ec7a37211b0b44026a9b743

Observation bc9bb639-42bc-455a-93fc-696a0943f177 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.805940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.805940Z digest=sha256:f4e11e751d9d4a2ee75d3b680240302b20254ff393ee1c457b7d6a6a93c4fe21

Observation 46230b7d-a056-4d80-981a-a5af8cdc0ee5 · inbound

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation cites this paper.

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T11:02:10.235519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:02:10.235519Z digest=sha256:562cd5f017126da795f5432c0650a287fd4e71f5e1709d5b33e1c58bed5ecf05

Observation 1502abd8-4c8b-4009-9db7-290ceeb31924 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:cc0924c21bb8d85b18062691fb862532a0103a3f0947edfcdaee38a070aee3d6

Observation e30a6bc3-0ff6-4fab-9f6f-95773074f92d · inbound

Multimodal Fact-Level Attribution for Verifiable Reasoning cites this paper.

Multimodal Fact-Level Attribution for Verifiable Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:02:16.278065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:bb6012dc67478f80ed60707f8e76d0ca2a446b11b43e932f98dcb0a7c21d68ba

Observation 975ed005-70c3-454a-a2f0-0753e32adda6 · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:45:14.477505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:ddb2f6e8ef8fecf586cc9f1a118135632294b019c53678d4345e7a55e4115795

Observation c8c948a5-b8cd-42d7-b9cc-1ab14bd6d24d · inbound

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering cites this paper.

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T21:12:29.596207Z digest=sha256:76e80998d0a98120b3a18de4e7a4476945b74399a1fb50db7c5ee0e194703c07

Observation df48e1ea-5eec-4fa0-aefa-fdebe8c86817 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:764e5e181b697256ed5d5e0c7493e04590f899cc9feac13f35bbc9154ac206c5

Observation 95b9441b-58d2-4316-86fb-d2b139bd910c · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:4c35ef7ac95f1876bdd686178d69270f64a174d0b82936f80fdf5995b341af1c

Observation 99e6c9fe-8f53-4edf-a8c4-5fe02ae8b085 · inbound

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding cites this paper.

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:47:32.778695Z digest=sha256:482c1823a342c58f7fd4fd8591a146f6f265ecb2ee2b68dc173c23124fbe61cd

Observation 32406f91-108a-4100-9698-0db4ae057762 · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:9be62e617892a736e55d9fc1f8d26fce982e72a58788e7dbe75eea2d3d26cf3e

Observation f13f9e76-f634-456d-8e11-b5626ab11d3a · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:59:48.877783Z digest=sha256:0fd32b743d05721da817a9c0ae30c477c080962f0830768b0992638206577b80

Observation 8d4db040-b2d4-4437-8065-0f0eaf386051 · inbound

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time cites this paper.

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:32:20.293703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:32:20.293703Z digest=sha256:f322ae88e88615b05ae415638a36182f297a9d68569538e3f4d7e254d37d454a

Observation 5860d76e-5fd8-4000-bce0-48763bc471b7 · inbound

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing cites this paper.

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:56:12.097628Z digest=sha256:40589b37fb7bc9e93ab752b9b969191634683f679341b24d2a96d1b223d56b1f

Observation 662af7b6-fcdd-4eb7-80f5-c21a804fea49 · inbound

EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution cites this paper.

EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:10:18.639003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T21:07:34.673912Z digest=sha256:3788234ae12a515770249a937a73ca07c037a0d20972bc11313b6ce4318d7ac3

Observation 0e8368ba-52af-4f81-8084-cf5a7813dbe4 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:ed4751ac88980be72d5a8b1712108a325592e2f8080739f2b6905439947cce1f

Observation 30bcae99-70f8-48b1-86e5-4f95d26b9c51 · inbound

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ cites this paper.

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:25:40.470947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T09:08:06.592577Z digest=sha256:72585aa1471608312bba256462bfa53d45fa5136023f3727226ef5c4f7de3119

Observation 0fbe53cc-b066-405a-83d5-fcc41ea0d682 · inbound

FCMBench-Video: Benchmarking Document Video Intelligence cites this paper.

FCMBench-Video: Benchmarking Document Video Intelligence Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:14:19.186123Z digest=sha256:f11538355b2687e1270383cb9eca4de11c3d08a3ee913e87066bfb180cecde4f

Observation a3e841cf-fbc4-4ee8-a112-989bbaea4e71 · inbound

Valley3: Scaling Omni Foundation Models for E-commerce cites this paper.

Valley3: Scaling Omni Foundation Models for E-commerce Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T14:53:55.160230Z digest=sha256:8a7d2124c9c48172d59fff6d144be45cd3042677d912ca5ae3e034322019a277

Observation 399d77b8-1975-4024-b06f-daaecd98fc2f · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:9568b3932ed75d8278a12e757138bd92e70575f42417657742c2941ef9c15a0f

Observation 670b33bb-aab2-4f71-a3b2-3a34d47f9dd0 · inbound

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing cites this paper.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T01:30:19.531699Z digest=sha256:f3db8f71f99f41597ef658c73f62b920fe5dbadc4be3ccac0a0c8d3c942ad859

Observation ecd96d5f-e3e1-4aee-a47e-1c5169d31208 · inbound

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing cites this paper.

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:43:34.898639Z digest=sha256:c811b161353942a2f0b93a6db980f0aeb858ce738882921d5eac0bea624a0811

Observation b8fe4e0c-8b91-44a0-bdcc-1a174110770a · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:5528af777badde5fc3face6af98ba194af18d5560bab46d0d42e676f3157bb9f

Observation d65a0186-1bde-4011-b6ab-6b082174943f · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:653f8788dc0351281c06f444b7bcdda10d5f148d4bde85fe1b330f02f5ebe8ea

Observation 37032c69-546a-4511-a58a-9e247519c44f · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:e2b1ca1d5b0fb81c7406a120d61ffa54ec1f74b4bea52749ee9d78780331b16b

Observation 6902eb56-7ec2-4c8d-9aaf-51e804f7ec20 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:9ce1dccad2739dde22023d29ea2fa0fb7120e4be877e6db2ebfbecd110eae78b

Observation 04079913-6695-4578-ae1e-593894db49a6 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:10:23.815676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:f56a8c2559802b4940a7a51075ff5242f42e8647e1828cb67eb2db0f8b51f6d1

Observation 3bbf2dbb-27cc-47c6-be66-704261458a7b · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:a23c7ce7f6cdeeae729013cedf5fab568e2a76e8a76b2021f67c8f1a28e426eb

Observation fdbd658b-4ca4-4051-bf46-7245cbe26f63 · inbound

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding cites this paper.

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:37:46.090777Z digest=sha256:7e153f08094d7c6c4c4992f46ad1e282c5b17bf202f817b8da64b41c66eaf600

Observation 77e403aa-7e0a-4617-9a56-89e70375d6f7 · inbound

EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs cites this paper.

EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:44:26.818726Z digest=sha256:205a37eb859b91e2898bb1cb5c7cf648394a0406505b70011d6a1ff1e8cae976

Observation 09d2012b-c9ad-4fea-a7cf-e3810d48b0cd · inbound

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding cites this paper.

AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:47:53.873409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:43:29.123615Z digest=sha256:dde96245a0e1381a27b8062397f0c06427ae3bf0f53821fb368816888ad4da7f

Observation eaf66829-84aa-4c51-a8f6-51102c737900 · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.426162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:c740a02bcde1fa028da5baf03182369667a197c441764ec6cf5a6ec73842ab88

Observation 5d37020e-f4a6-4ed8-9879-7241181bc7dd · inbound

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions cites this paper.

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:53:38.979624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T18:49:18.815456Z digest=sha256:44ddfcefba734ae742650add84242c3f6e6cf32e5360a6580c8bdcb0273f266c

Observation 11c768d2-96b1-4b5c-9b61-77d97b6e8cbf · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:48:23.420375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:d196019412423d61aceab83fd79243a464150d44cf33884a86187952dd49ee77

Observation 86dc8a94-eebf-4a07-91d8-daf4fbc58408 · inbound

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency cites this paper.

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:38:14.809166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:33:39.121508Z digest=sha256:7981b6b754c484bc14bfd7a0889db13084d27480d66cbc98c2875c01c53b27fa

Observation 35f983cf-b3fd-4ab5-9d4c-ed5e032f54a5 · inbound

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models cites this paper.

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:21:12.867000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T07:19:30.508843Z digest=sha256:3bc2eeb3488ffad8438a619a632018f50317ba35a9256572708bd051cfb5edb3

Observation f9f54f77-1808-4579-9437-5cf3ad7adc8b · inbound

Cambrian-P: Pose-Grounded Video Understanding cites this paper.

Cambrian-P: Pose-Grounded Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:51:09.016264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T05:47:13.493461Z digest=sha256:3f30e68c361d029403ba1b84cf0f269a72ca17946bbc73b0d3be30b86769ef62

Observation 02c75af7-b6cb-4797-8f26-a77bd34fa8d1 · inbound

Cambrian-P: Pose-Grounded Video Understanding cites this paper.

Cambrian-P: Pose-Grounded Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:35.418105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:35.418105Z digest=sha256:5bf849d65fea4fb7bd97283208ccc63545b5aa08b945568a215db88ba6b09c48

Observation ebcfad3e-b98f-46bf-872c-884bcdf6d89b · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:55:24.884952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:58bc7d9a791db43a426f2df28fc606906a57163d5b8f36419ec1c23cb65f1f64

Observation 0d39a4c6-621c-4053-b4df-56760c7414c9 · inbound

MetaphorVU: Towards Metaphorical Video Understanding cites this paper.

MetaphorVU: Towards Metaphorical Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.349973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:43:20.101830Z digest=sha256:02ae74b13b617be9e123286d44d853de8026df20cbcca25289ed8844e630d880

Observation 242c2a57-807e-4572-a409-cb13252dbdfd · inbound

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events cites this paper.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.859293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:b05115e2dd01cbcbc8e1937cd87faab217244182c4066d9618ab8cebc58f56dc

Observation 020b3d45-544e-4925-82be-d3faa3709af9 · inbound

Benchmarking Visual State Tracking in Multimodal Video Understanding cites this paper.

Benchmarking Visual State Tracking in Multimodal Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:27.939496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:43:06.811228Z digest=sha256:9056cb4f7e3b8b1a786646869865d3aec19a6f642ab34b8c3fe58f692fe7b674

Observation 3d47b554-d500-42a7-8e76-8ac6d3715538 · inbound

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding cites this paper.

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T07:16:45.155441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T07:00:21.192082Z digest=sha256:c83961071a221e9879f593e6293198eb05226429ed62c0126590fef68a7d4d2a

Observation 87186e59-cf0d-40a7-8c35-802d3cf697f9 · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:26:57.207629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:2a7ffd9c94f410d34bb463f516f5c1be9cf7ad456419d62f68b4eb3745010760

Observation 26d3b2c8-06c6-43cd-af18-ba2ad3dee28f · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 251

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:27:15.841155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:bcb184165c795a7f0baa3cb4d4e377066a35ce4a7014dcfc62a9ebb951e420e3

Observation 23dd6458-9750-4df2-98a5-95f316ea0833 · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:37:25.403851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:ec4e25ee7575bf9cbede0d7f268480c6cfda1752fb11cf07e3774b98c10e45f8

Observation e382ad80-428b-4bd1-a2a9-d79a528c69ce · inbound

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding cites this paper.

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:37:24.829550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:40:56.851297Z digest=sha256:6b62cc88b0da116f7e3549bc25b5e231e3cab01911ad36e27a2cfb55568c33f1

Observation b8a80bee-0064-452e-82b0-0978aaafad45 · inbound

Harnessing Streaming Video in the Wild cites this paper.

Harnessing Streaming Video in the Wild Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:37:25.747362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T18:47:55.910417Z digest=sha256:00916b1fb09230860d50f4f6053c3e3a6a7c5e6e7cc279745e28a9ef0bb2579c

Observation 40e94dc2-374e-4ba9-bc44-28a3d2e0ae44 · inbound

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? cites this paper.

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:57:30.606118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:52:22.811857Z digest=sha256:9127dc065d5bf3450f334a878fa4497dd05219e565cd4f0af95cb1086e2be53d

Observation 60d05fc2-1132-4415-b50f-0555ceb720df · inbound

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding cites this paper.

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:52:27.181806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:51:38.268304Z digest=sha256:e18fb7d8c8c60316b4a319ab7614e1a64333abcf0a259b2d560315fdce50c36e

Observation 8f22d88c-4c28-489c-ad91-d4cc8f183447 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:27:37.091762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:45a934ffd3ce96c81d01f5537d08d62ed3ac0bd3a425a01993c64570eaa1c6eb

Observation 477122f4-1492-47ec-b288-ce6ac919f074 · inbound

AVIS: Adaptive Test-Time Scaling for Vision-Language Models cites this paper.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.846147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:1f9093d0a4b93ad96471ab945cb18e973e272a760554bbbd929785b0af45e9a7

Observation 26f4f85a-2d54-4a07-a1e8-9e3a4eb3b666 · inbound

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models cites this paper.

MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T10:27:56.142403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:02:58.341050Z digest=sha256:c1592c4a2055651847e8f259fc88e78789264a3c37d9c241174a96ebaace4515

Observation 3afcec70-e702-4724-bc73-5f78ec9efa3b · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 154

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.923227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:be065ca0f94f9ff7221a9748907dae7d8826b8e61a2f920797aee9c3ecd18bcb

Observation f0e9a8d7-384a-4f12-b05f-22c7a349164b · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:58:57.741250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:1ed6f7d49f8a51352f6697db987619b0fd0660a68b0127a6a9b857f8e133479d

Observation a7df9cba-fea1-4bf8-b833-80e21872a515 · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:29:31.397416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:e64faa9c47fb49ea41a304d8732803b134130245f002cff27986c113b9efb7ee

Observation 096ab594-a4d1-4d3c-81cd-e78aea2f8b8d · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:40:00.862729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:65ca55f9a097c94296f559a1e499f763885da6db25506fc2b160e87af61f32c9

Observation 86253e11-0555-40ca-a8c1-1c3749753fe3 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:39:50.746585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:ced9e43297bb7e5eddcf600cb10c39c205cec68b9e1dbd3f557b9dcb8fd867a8

Observation 75bcbb7a-7456-4d0d-b18a-f5a9cfb67c53 · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:54:21.330174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:44641f5181d4ea8252f14a99fe9bf43c3a2a1786c9b9657f1c54e43d2976d264

Observation 9a1e924e-edb1-4bec-b26b-a1526b84a879 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:07:17.510059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:d7278390c606bc7ad6a89a598be966be252a67d36b326dc7f9826110d5757431

Observation 7a5f67c0-d40c-47cc-9019-72f1c7e0bf0d · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:27:03.311221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T14:19:04.890074Z digest=sha256:9e93f307d0480f10ba89059c59fa31d7380b5a0710b0670154a2c1694d6e9752

Observation d9c714e0-1507-4d88-956b-847ee6678d29 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T09:13:38.446621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:13:38.446621Z digest=sha256:7bd3819bbf6b7237861d5272fe24a14bd1ebc35409d4ea6886838f4033f2a085

Observation f704073d-3c4e-4b54-8fc9-4a3843aa86e1 · inbound

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection cites this paper.

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:20.550240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:20.550240Z digest=sha256:1afe2d850def6965a49154613f3f1be50e18919df4748cc2164f0e50fc25bc48

Observation 56ae4074-eea7-41a1-9e60-87eb220447d9 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:47bc83b11d4a1dfeeedc3f9e4757031961fd5e27a0de6d1322d5a367bf3f4c05

Observation 79b7c0e5-ac12-46d6-bdf4-d8a4afaae848 · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:01876d35ca339b46e75c2208b9143508c75e51fc36fbfb1324c89b61077c21be

Observation 87446d1e-82fd-4a2b-a961-4bb2e12e55a9 · inbound

S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval cites this paper.

S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T07:42:08.072354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:42:08.072354Z digest=sha256:4a88baa586049b0316c18f7572388b3abfdce55925823bd9b05c8a92f61a877b

Observation 403c2c3f-1b87-4da6-a15a-2e3900ab221a · inbound

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning cites this paper.

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T06:04:16.637378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:04:16.637378Z digest=sha256:dcd14b5d0427f260d6d1553c9f6e5ebf161c9a195b7b9cfc29472288f45deb1b

Observation 3b511d71-d1ed-44d5-8104-b1faa0a5c943 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:31e75344903e75899c74e1726f2c90bd9982c228ce26dc39e15c3630e3827745

Observation 1a0e508e-08bc-41c5-a574-c1a893454cb8 · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:47.002048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:47.002048Z digest=sha256:11f41bf028b10e394a51f41e8eff1b25d33d1db30a7f81e4fb6668b5796186eb