Pith. sign in

Paper Citation Record · LEDGER

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2406.14515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14515 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.815801Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:12:46.636950Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 38f029b6-d9a7-4cfe-a8a0-5b6ca001afd8 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.405489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:f3462380af17faa56d28f2352c751366f7fe4ac9069fb9a7e087784b91849c78

Observation b8e1403e-d56f-470d-9a35-365fd639fabf · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.609435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:09cf66df7c3795cfb86edd86becd5c1e2e1086de359522094b78742fe2bf08ba

Observation d7487486-8b20-42e7-96d7-c58ac5d1533d · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.446997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.446997Z digest=sha256:8d8d8050351e42137a2d74e63b41d5f7e6bf824780b6c6bdc74b410ce7203cf7

Observation 06089e05-5ea8-49b4-8d4e-daea623a9032 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.069516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.069516Z digest=sha256:2b073a3a2f703d046173fc62b434b534313c627d0c53f22a515c9913e0352a35

Observation 4429b2ad-bd1d-413d-a6a4-054da3e47464 · inbound

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability cites this paper.

TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:55.973925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:55.973925Z digest=sha256:9b5e8f7373e4b398318b4db3f5b5c2b92640594142089de627f30d531cdf3a2f

Observation 4550844a-d677-4d83-8c5f-d04588f1f544 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.027352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.027352Z digest=sha256:d8069c8eb18d7a59b8f42f20eca62fe474a32a2300b7f83b7ac6374cda4f276f

Observation 0bbc87ab-4cc4-4e82-851a-e02b4159ce21 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.987120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:e38ac46efc2e759c88adbf6cbeacb212d0270207ce4e6893e34a22f2b5828c23

Observation 95f27273-2fd1-4cab-9b94-5590a755724b · inbound

Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity cites this paper.

Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:00:58.980137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:00:58.980137Z digest=sha256:0aa07de401a92c5014ea367c131e1d2af867efc4a859d42b97f22551f06e218a

Observation 0475e5cf-6afb-480f-b64b-86e9bec5031a · inbound

Neptune: The Long Orbit to Benchmarking Long Video Understanding cites this paper.

Neptune: The Long Orbit to Benchmarking Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:28.853040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:58:28.853040Z digest=sha256:a66d28c749dbeff6c38f9805a1b90fc8a834a7730c1e34ea1fdc555d82101b1b

Observation c530d3e2-5ced-4dde-9fcb-81f40a943c60 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.112524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.112524Z digest=sha256:6fd1c5242550e3d24e1e193ef921009a1394ccd8ddf77d57b79c0dcb1a5080ee

Observation 216eef99-e433-4e15-ba2d-d43f10802772 · inbound

VCA: Video Curious Agent for Long Video Understanding cites this paper.

VCA: Video Curious Agent for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.116123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.116123Z digest=sha256:2074f322c5ed5181533762a9786256b17e49a47f542f79cbfa5bfd71e0516ec9

Observation c4427d20-f72d-465b-945a-e268ff6e010c · inbound

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding cites this paper.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:50.968851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:50.968851Z digest=sha256:05c183ec072c30e4b14a6a3341d6f28299e37e017074dbc7699ca0b7e18dcdb1

Observation 44f54430-2512-4e8a-a767-9125548171aa · inbound

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks cites this paper.

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:05:29.189875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T07:05:08.716223Z digest=sha256:50f32aefc752f81226454e8046b9dbac0fe41063d5b03024e763e7b473f39127

Observation be82f068-0bf9-414b-8437-76fa6d796041 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.438229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:cc70aaa66aa3e5e18c3aece74bb4bcb6ccb4b4f3e4289a9e303b754b2b243041

Observation 59ef57de-f721-4e37-b3cd-80215076d60d · inbound

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark cites this paper.

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:25:47.030222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:25:47.030222Z digest=sha256:faa29c23e56440d17bcafc34f528638c03169b374f0071a70d54992a24764642

Observation bcd3f53a-9953-4561-8386-c85bc699baec · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.710199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:bfd62b757e302711c4d0091f1beea46f057886bef981396bb3722cd58b2e266a

Observation ec1e6ef3-fe23-4b98-b54c-6247ef07148a · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.214239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:4de80cedc0cf1f20ded43730decfafa439b2d0f56389f30175023405eae64c58

Observation fd4d1967-f873-4294-9409-954ecfc5c1bf · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.306282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:e94d64e089bbd7226c3cdb49e32b1eb9be55ede9e695767b5c28dd894f7bd107

Observation ccd52aed-5402-4615-9018-e7e59f80c64b · inbound

A Benchmark for Crime Surveillance Video Analysis with Large Models cites this paper.

A Benchmark for Crime Surveillance Video Analysis with Large Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T21:57:23.556854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:57:23.556854Z digest=sha256:bbb6e1414e95f1ec79cddec9022e8e831ebfdadbcf96d63d629c9357e4620575

Observation 3129f62a-b5cf-49d4-8aaa-2cf1e10f6c44 · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.026892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:d0657f1fd67929ab79341bad2e5ec69a7212251c0f92a1179a404c64f8f978bf

Observation a6b6f22f-c849-40be-aee7-4c23dfdc372a · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.272679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b6e8102fb555cd1af6789ad8a3c4e7c9c7870f66d0034e419afb376234053b4b

Observation bab3da26-6402-45d4-aecf-bf17ed42d31b · inbound

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding cites this paper.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.815801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.815801Z digest=sha256:a7b11e444034e8c9652d3db4f396c235c988459054fe849ffaa360e07ae55511

Observation ade2840b-5cb6-4e3e-9f89-dc08e6d45abd · inbound

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark cites this paper.

Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:08:38.374553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:08:38.374553Z digest=sha256:08c7a9f74e183d0ad3a84a49eb36b3803ff2011443892c8bc69a292a00c89d00

Observation a386e1c8-8df6-49f4-a812-084e235ae6a5 · inbound

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension cites this paper.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.473696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.473696Z digest=sha256:dc3753028d7bc422d217f990c8692783899eb696b6f5428ee5371dd88a964cce

Observation b69cae93-61b1-4404-b6c9-e5551be0bc90 · inbound

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding cites this paper.

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:09:15.035116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:09:15.035116Z digest=sha256:e369ef6e84b31134479b212e6788432793954d2da54764c5d5464ef5ccdc0370

Observation c5bb7880-9522-4bdf-bebb-2110cf0c0af6 · inbound

R^3-VQA: "Read the Room" by Video Social Reasoning cites this paper.

R^3-VQA: "Read the Room" by Video Social Reasoning MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:41:34.108189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:41:34.108189Z digest=sha256:36c5793761ef3ca7f43fcc4c4f669938cb9d548bd5a2e02f81973e37a6b83c9d

Observation b05d6f78-5e5f-4c83-9fe6-b3e44bc82297 · inbound

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models cites this paper.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.389906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.389906Z digest=sha256:f450496c18fd9f6d10e5cfa43d8d2f106cda6784b7d9cba0c36fa0ea7beef330

Observation 2b773435-f45d-44d3-a2a5-28ef34363b23 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:44.920605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:44.920605Z digest=sha256:3b9ee8e616ea722c28415e8916d1405744ccea2a84af1f3aa0dbdb50a26c6d96

Observation e8a93478-79eb-43dc-9b5d-4b9eed8e11e0 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.863411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.863411Z digest=sha256:210e629463bf95997c1aad490c030662d2c17e7a055e9d9f68f940df956e96fe

Observation 07a3c23f-0848-4e87-9a93-13c1267286ba · inbound

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding cites this paper.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.950597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.950597Z digest=sha256:0afa7ee09fc4ed0ea0a160c7c0293783f41b4c1b78a96bdedf9aaba56cc514a5

Observation 2b04f4a6-1cdd-49c6-a73d-34ebafa03d10 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.983556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.983556Z digest=sha256:d3d8276866e971b3fd5124027c4f828e95b16d8033c1bd8b7bd56c8e7ef23736

Observation 419a10bf-2fc6-4e6e-a497-af875e6cbe22 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.569319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.569319Z digest=sha256:6234a0ca739b553547333b836ba0056b1507ced81404189ee2c04d3c480c171c

Observation 37ab95b4-f970-4028-bb65-c12a90eb0e61 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.080794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:0b5193e36934d49aeb33aefa3987fe03f93d6a555bc193f5fe024780118acf97

Observation 7943e501-1c3d-4caa-adb1-bb1002222058 · inbound

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models cites this paper.

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:37:20.820402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:37:20.820402Z digest=sha256:c245f12a2c31c0bc7d3274c94fb7408f2f76bfa9ddbffcc9173d25845ad4df60

Observation 59fb7299-bb2e-4356-abf1-6aca0a378afb · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.638751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:19b76cf025bcafb2de0f5280d1cbaf40e8194c2fbbef12c2cb50194e54355b37

Observation 08fc3bc3-eab0-45af-a13f-60e86cd4ff4e · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.251209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.251209Z digest=sha256:2e1a9b63d68ecb79e8590bee0583b1cd8e01b459a56d81ceb1d638f353fba195