Pith. sign in

Paper Citation Record · LEDGER

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2406.14515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14515 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:57:23.556854Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:12:46.636950Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 38f029b6-d9a7-4cfe-a8a0-5b6ca001afd8 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.405489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:e69e8d996fdb4654a8766a7326ae49ccab78777a18493880e9c389d3d7d074c1

Observation b8e1403e-d56f-470d-9a35-365fd639fabf · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.609435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:1d66c28514635d18e578f42bcf2d3b3dd020365e98501ca82b1be327a76c31e4

Observation 0bbc87ab-4cc4-4e82-851a-e02b4159ce21 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.987120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:b66c9cbbe22649157e86f38a835d60890648497dab6c031fcb961099b9c30fef

Observation 44f54430-2512-4e8a-a767-9125548171aa · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:05:29.189875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:c168b630404843633e3e8dcca6d75d4d04a72a9872c2f5dc9354475ed8c04048

Observation be82f068-0bf9-414b-8437-76fa6d796041 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.438229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:f55c98a3d47a0d3a28174a10fc26df5b04cdf1b2761cfb62d7e264126ea7b62d

Observation bcd3f53a-9953-4561-8386-c85bc699baec · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.710199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:3899fb7176536064bca1d5a101e883448940e12a1991dddc5449fb399753b521

Observation ec1e6ef3-fe23-4b98-b54c-6247ef07148a · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.214239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:dc2e38dc3077e5695bab81eff005d61a14911de9604aa90bf71c8b378779e561

Observation fd4d1967-f873-4294-9409-954ecfc5c1bf · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.306282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:d364937d1f561cd206b62c45d76aae719bd178e21eefb587011eebba562a55d0

Observation ccd52aed-5402-4615-9018-e7e59f80c64b · inbound

A Benchmark for Crime Surveillance Video Analysis with Large Models cites this paper.

A Benchmark for Crime Surveillance Video Analysis with Large Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.556854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.556854Z digest=sha256:3b016f5bbf64eaf093bd9dfd10910343c6f4d1a07f1a792f1b7317d125f8c7dc

Observation 3129f62a-b5cf-49d4-8aaa-2cf1e10f6c44 · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.026892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:ef873183b042e19614297fc93076724cecb9c4bd91a607f78627e90cc54a9995

Observation a6b6f22f-c849-40be-aee7-4c23dfdc372a · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.272679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:dd94b272ca2606d382b208a5cb09263113cc8d57df52f8142abc44edc918f52b

Observation 2b773435-f45d-44d3-a2a5-28ef34363b23 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.920605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.920605Z digest=sha256:e8d7f15627320518f9de2a08d4961becb7fecd28e57782520a035914e96bf3bf

Observation e8a93478-79eb-43dc-9b5d-4b9eed8e11e0 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.863411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.863411Z digest=sha256:75fa6d1189fa998e7c204dbb230fd9449d643bc5a72c092102dd3b4923a6b2a0

Observation 07a3c23f-0848-4e87-9a93-13c1267286ba · inbound

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding cites this paper.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.950597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.950597Z digest=sha256:8e4c00066d6668c2e337697a5f391fde71f00c6b2eaf2b5a18b2737de65727ea

Observation 2b04f4a6-1cdd-49c6-a73d-34ebafa03d10 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.983556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.983556Z digest=sha256:796473884583f0c8538e900dd484be5a94335aa568f53c9316d3b5c72ed71fea

Observation 419a10bf-2fc6-4e6e-a497-af875e6cbe22 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.569319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.569319Z digest=sha256:c71d9da35b652292e7590b6a2da23973ae8ce16279914293ca4adef9f7804d81

Observation 37ab95b4-f970-4028-bb65-c12a90eb0e61 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.080794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:2db7d035360a0ba7095cd09ec167150d340f94abf8cf7a505d430c17adc7681d

Observation 7943e501-1c3d-4caa-adb1-bb1002222058 · inbound

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models cites this paper.

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:37:20.820402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:37:20.820402Z digest=sha256:8657e50bfa7801e4cca42f95f6f52e8f4960af98d24c98b836d55134fbce1936

Observation 59fb7299-bb2e-4356-abf1-6aca0a378afb · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.638751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:f2442eb7165087caa16eb91b8521c0e637fef4c4ca3cde6f870eb7c7caa7baf3

Observation 08fc3bc3-eab0-45af-a13f-60e86cd4ff4e · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.251209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.251209Z digest=sha256:1f2101a8e9d8098349c6e13c74b659336ef03bf4ab16b2aad4d7f8edc0151c03