Pith. sign in

Paper Citation Record · LEDGER

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

As of 11 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 50 inbound Pith citation observations for arXiv:2311.17005.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.17005 v4

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:22:34.954228Z

measured 150 of 150 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:41:47.557078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:29:41.281374Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact29
  • verified fuzzy64
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47b0421e-9227-4d52-b309-3a812e2031e0 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.111518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ad0f999803bff94d38307d75fa58e8789b8b927c70054ee92a51b67c2239f4ab

Observation 90f2ffa5-3409-475a-a4eb-0056df666e3f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.082058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:e26c91e4b9793dea287bc5329ffd450bc247eaca00d2d639b81a8eee8252372a

Observation a781def4-bb44-416f-9b1c-4d89c54d6663 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.287880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:4f70b01156fb6360599e31fadb52ed009423477eb8d54ab8ed5bc87e5ca19672

Observation 995854ba-d09c-4e4e-9e25-c06869744d6a · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.290504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:eef98901fc4eaf942a8b5d5329961b3fa3ae7865e68c1089db38284d341a8e57

Observation 52cfa3b3-a378-41d7-8322-8ffb0a02d404 · outbound

This paper cites Language models are few-shot learners.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.293232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f42e28bb1a639498b59a5d95f94cc0a9340445a7191c84a609ab040a9f71c016

Observation c3e1ca57-0421-4234-88e3-88fda803c644 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.296457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0915127e69169ea07669978125cf4072dacb2bfe78c8faec769e2d37bba0b96c

Observation 415a9afd-4921-4a85-96bb-8488c299c38e · outbound

This paper cites Chen and William B.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Chen and William B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.299565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:16a478b06dc04d157b66ed2bd41bbbca79fcfa6018de4f731427c9697b9af05e

Observation bf002db6-6836-4ce4-97d2-0bdbc78a7c41 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.086538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:34989438cd969e84b01b08b0c7e8a25cfdf99a69bec3b88ae30d35a460af980d

Observation 5b8838c0-e6d1-4221-b067-239d76fb8ecd · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.041933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:67d56fff17cea03d0a26f3303f991af2729244ff5732707d5f61b171732147dc

Observation 0613568e-bef9-4bb2-9dde-031db678094c · outbound

This paper cites Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.302838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:96f6b2e1cd7fde5250e174135b79729292e32ac6b20721f7cf7ef681e316da26

Observation b185b69e-4a16-4aa6-a4e2-e8a43e68198a · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.305479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3b944a292f03476e180e88e0e30409f317ebb3cfde2c5da1314241a2ddedb142

Observation 4359a529-2860-4cb3-9520-561c8dbb8953 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.308195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0e7016c248e63029653d6b7504d875caf8635be932af7f70d06e38ef3a607598

Observation c48d2239-c31a-40c2-971a-a003b75eba68 · outbound

This paper cites Doell, and Jason J.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Doell, and Jason J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.310868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2711693470047a0d4966928e1840259765c8671ec4a2b37637bf47092f6aab29

Observation 0e34e3db-000d-468f-8351-88cef33497d4 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Imagenet: A large-scale hierarchical im- age database

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.313831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3409166cc4ff656d9fdf8df21becff4261b3c812145f12d4e7d75be3da283ab0

Observation 5ad86f32-7955-4db5-89af-1ec1fe8cfd93 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.143046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:eaefd20c7887973705473632888046188f8aeecf39d94b9cf90ec21931cd3824

Observation 7d4ecb7f-b028-4a26-a347-e016f26f4a55 · outbound

This paper cites Xia, Mehdi S.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Xia, Mehdi S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.316545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:d517027866b0501fc31a9ac026cfcacb3f01f576abfc9a36607fabe9bc1fd425

Observation 490064e9-e5ef-4873-b3e8-33713a887665 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.053895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f1fb3364173d7e53d9b4c8dc66741856df60963d50617822eab3cb322d53ee85

Observation 067cc1fb-e7f2-489f-a50a-ad1df38a4057 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.059332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:29e949b8b01dfd206556c368d446b08e06ac1fcb4d711c33cea30b48253797e9

Observation a851186a-6e04-4636-8230-9472ca2a0547 · outbound

This paper cites Mist : Multi-modal iterative spatial-temporal transformer for long-form video question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Mist : Multi-modal iterative spatial-temporal transformer for long-form video question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.319781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:eb25b710e0f296110ef7cd6d79f50582b99b77471ed6f5a6aa24a2f17c36a438

Observation 3d21238a-6de4-4637-87ba-c1b3689434c6 · outbound

This paper cites Gao, Chen Sun, Zhenheng Yang, and Ramakant Nevatia.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Gao, Chen Sun, Zhenheng Yang, and Ramakant Nevatia

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.322872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ad8208f3f37888a16a4f5f30826380eb9207c068b1f0f5e924e80799d353e64f

Observation 3f44ad43-9dd2-4b8c-8bd0-78657d63d7b0 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.127729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:b374789e17e22bbcfb964f050c21aeee365740576e75ffbc4b115bf7edec3992

Observation a7a0662b-e956-418c-a4ac-b0879b911ced · outbound

This paper cites something something.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark something something

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.325601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:12f1bf78059e6f11d939e1a571ab82e4835e26766af87796fd4ba12f21c37bf2

Observation 6c708d15-c26e-48a9-8c76-ea2581ac6a17 · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Making the v in vqa matter: El- evating the role of image understanding in visual question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.328577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:324ad4ff63b2af88178ceb80d6a129bfee410451fc2650b9c23b8ea1cc497784

Observation 3b2ce860-54ea-43a8-ab4b-40115579d15f · outbound

This paper cites Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Z.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Z

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.332073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:76188819825a9abd292f72e58b0ad00c99d31fd6ac0956b1ad2d187e55616ee1

Observation 066cf6d1-4f64-4f50-bc53-d3990f8098b4 · outbound

This paper cites Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.336385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:df8f3fc91e113c424ba24912cdf21832cc9c1477cc6f20120f8a4b0ac15071a1

Observation 70db8531-0909-409d-bbf5-548bb0ad9086 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Language Is Not All You Need: Aligning Perception with Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.063626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:28115315c55b51fe77bc7c449bab716724f0813b87f18a0936e70365bdf27972

Observation ef3c492b-4165-4354-892d-41ee5ca3db79 · outbound

This paper cites Hudson and Christopher D.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Hudson and Christopher D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.340231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:abc458b9f75ba2a235a83f6cc85143f9cf30bb39b68a09045682d4416f14c2f2

Observation d0946226-b6b3-4a07-91db-745b0bc46d52 · outbound

This paper cites Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gun- hee Kim.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gun- hee Kim

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.344812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:988d5f1965a7051b792960137a095109364630483eeb61f21009ba987af0d739

Observation cd232c85-9449-46ce-a234-e511dc54e46c · outbound

This paper cites Mistral 7B.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Mistral 7B

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.103462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6d1f93a5b276f3fc536ab4abef4e5bab0b158916afb316f0c6e47ce66ab47811

Observation b7ee2ce7-d33a-4b24-83f9-95e86febe9e3 · outbound

This paper cites Lawrence Zitnick, and Ross B.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Lawrence Zitnick, and Ross B

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.348342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:110c774e5d03379011dab54af2ec4eb95d25862e6523a4fe2ce799b87468f523

Observation c6867e38-db18-4f71-8bd6-8a4ee3a4dd8f · outbound

This paper cites The Kinetics Human Action Video Dataset.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark The Kinetics Human Action Video Dataset

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.131390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:1d0dada1f0be5fa18269f93147fba66541cd4e60bd9ca477f443ed1ee84c3e89

Observation ae08a2f6-519e-420e-8d3b-b0d991e68af2 · outbound

This paper cites Beyond the nav-graph: Vision-and- language navigation in continuous environments.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Beyond the nav-graph: Vision-and- language navigation in continuous environments

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.352190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:bb8371c26b87c130103576adc0c8a32e5080e5d98998a8f66f2ac84f01a4ea35

Observation bcbeb098-4575-4909-bd39-833c1a54f38a · outbound

This paper cites A hierarchical approach for generating descriptive image paragraphs.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A hierarchical approach for generating descriptive image paragraphs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.355843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:b8bf704707a688865d9f2596abb3202e3715d2274112e67b189e75bd5643c4f0

Observation a695e460-0da8-4aae-a0f8-ac4b3f3eb42d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.359280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3fb8ec1d766427fe57c5f5620330fb1ac6afb2cd59ceda67c835dc2143b7cc9f

Observation ca74c5b5-1186-4cce-a794-1f3d61e0311c · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.362332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:5189e62ebee5ea0ae8499626bac9c59cd842bb648687b280585409dda898c980

Observation 690b9d0a-9d42-4bac-af25-04d8b6a2d0af · outbound

This paper cites Moreno, and Jes ´us Lov´on-Melgarejo.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Moreno, and Jes ´us Lov´on-Melgarejo

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.365833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f7c2831b85b04c66aa00e2010b94368025e7a68051ac19a93bff09d3bd9c28bf

Observation 30eeb6ab-3cc1-4e62-a226-dd766552f830 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.068109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:a5ee4db7ec457c517486a30eebd524ab10982359f4446ae22943485ddb6c0b4f

Observation baca8caa-4445-42e7-8a03-44098fc3dba0 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.072447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f53863d007e1f91e0c06055e82d09c440e2a1b6d25b7b230209da062a7973440

Observation 14800046-ca1c-40ce-8666-66ef67006f86 · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.368773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c6e1e8e81050b479ee22e382a18b4e36712ff80183144e5ff35d0fffeb0cae2a

Observation fd1694ba-3425-4fe8-b8d6-7a3d352f70be · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Inten- tqa: Context-aware video intent reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.371437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6b5c5eb04c5970bae3d8f2d75e39c30878ceed94d013ada18bc866e6c9903534

Observation eb5d0769-acbf-4ce3-9e46-8105f8c30d1f · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.091018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:fafe0a1e04379b3019e31374a54c8130fddce79f8fba50ca6c3f94f48c671965

Observation 884e9606-2235-40c7-b760-7076483c38d2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VideoChat: Chat-Centric Video Understanding

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.099555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3166837eefd6a7b730dd75d3619950b28683abb5cc4e38c53c8d3598540d22ea

Observation 5c253939-cc34-4738-ba97-ff8401600cc5 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unmasked teacher: Towards training-efficient video foundation models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.374276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:1400b52c655cd03677b303ec42f287bb86fb71b58734a7779a3b909380f5d05e

Observation ffad8371-cabe-4fd4-b398-09404bcf7c7a · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.115909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:7c6ed9ad133c3cc87957e9bf36b5ee695f49ac0685f2b2c9f537a9285e727fa6

Observation 15f2eb10-6ccb-41ef-bc89-8b39ce753c70 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Evaluating Object Hallucination in Large Vision-Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.119842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:781c58235c13614da3ba83e3aadfcac34a495b0badda8257d8e7dc3058998deb

Observation ac434c46-3150-4eed-9857-dcdd0b704875 · outbound

This paper cites Microsoft coco: Common objects in context.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Microsoft coco: Common objects in context

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.376770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0fd9e2ece16e46757d0b6f21b42c03407454c682680b07c6d5509150051f3b3c

Observation 60236d7c-dc73-4f1e-8d74-fd71425a3de0 · outbound

This paper cites Visual instruction tuning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.379131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:df76961d721df52c342b0b24eee97cac08be85b713bc55f97ca96b87484697d8

Observation ff17c2e8-d1ca-4029-946a-8a09807addf5 · outbound

This paper cites Ntu rgb+d 120: A large-scale benchmark for 3d human activity understand- ing.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ntu rgb+d 120: A large-scale benchmark for 3d human activity understand- ing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.381831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ba0c7df735820267a86726bfe364ff9ffcf54d2c15ebc003180e3ff795133ab9

Observation 5d779c7c-2c89-45b0-aa56-dc03a5e0ff35 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MMBench: Is Your Multi-modal Model an All-around Player?

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.147202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:7316881b9b7d8072d903dc9c16e74707b77672fd65d9719f2d39c7517b711500

Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.014045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c46e30aea933926708c8b2b3e5eb319d3323de5f07d14da7aa81121704e8059e

Observation 706d5a7a-6d15-4eed-a84f-052294720fe6 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.030804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:efa1d7d4540c2d48c386536b16899ca95cc8b4e5b5b558945ca49d3b3dad3fff

Observation 64afd474-cd84-470f-a8c7-086f8848c8d4 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.036418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:d5fce640c675542f0d4750d470ef9d5858aa2ad312e6fc186955f7749e0d708e

Observation 608c3e73-431a-4faf-a295-280d692164e7 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.384389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:5dffd6ae32f489f4b55d0989fd7752d02ab0abb8d1a1100f8159a648ea26f259

Observation c5f0ace5-ded8-4df3-9488-63f134d96c0e · outbound

This paper cites Manmatha, and C.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Manmatha, and C

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.387189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0c1382fb0d3a858cc3b7a57ae0d47dd875d8185c347d54ea0f6110710dd85cc2

Observation 8d385301-dd3a-4c13-a487-c90ce82e7b19 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ocr-vqa: Visual question answering by reading text in images

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.171192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ef09ea9c3537d77ee6abf343a87b1973e7eed0dbae9b60821aed01cfcd7b977b

Observation 90ea304d-a46d-4428-ad42-b05a0c1866ab · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video de- scriptions.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Spoken moments: Learning joint audio-visual representations from video de- scriptions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.175006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9a99b745b4b3e86aae08fc224dcca3ae0b3cd636421caeb190cc83b6a512e43b

Observation 0faf0da3-d353-42d0-a1a7-57b4ccf3ecfe · outbound

This paper cites Brown, Quanfu Fan, Dan Gutfreund, Carl V ondrick, and Aude Oliva.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Brown, Quanfu Fan, Dan Gutfreund, Carl V ondrick, and Aude Oliva

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.178454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9f545bfa06a66fbcde82f782c0f76a88e60cc155aab7ad01c9e68d3b29442518

Observation 82c7baf6-c260-47c2-9721-59e293ce147e · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.181408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:51b014c34b0f887ba36ec3ac700411d15855707123afb4552664a77029a2d1c3

Observation f9097022-6d10-46cc-94b3-4f05fff240a0 · outbound

This paper cites Gpt-4v(ision) system card.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Gpt-4v(ision) system card

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.184444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:199a40e17d49307f9fd62bb542f5e28c1531f38edb98ddb8889152e8d04fc8b2

Observation b73ffd91-2fd3-46d4-99b8-b47f03135d92 · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Im2text: Describing images using 1 million captioned pho- tographs

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.187488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6b9bfeab6215930256012831c638cbee05cca81b1678f87720afe6ac2b103db0

Observation 6499d076-800b-459b-8d4d-479ba6c6d609 · outbound

This paper cites Koster, Junlin Zhang, Stephanie, Winkler, Yusuf Aytar, Si- mon Osindero, Dima Damen, Andrew Zisserman, and Jo˜ao Carreira.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Koster, Junlin Zhang, Stephanie, Winkler, Yusuf Aytar, Si- mon Osindero, Dima Damen, Andrew Zisserman, and Jo˜ao Carreira

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.190584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:a9ab46b52ca5ce2da0139b5f7e19bdb398c75b2ef87bc08f101fdfeff49e1f76

Observation a39c1eec-c583-4a72-9465-f206b833dd48 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.194198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:66e95c224e3d19c45981cc9e031cdc05fd1ee5e8cbbefc216f764d34265270df

Observation d000b56b-1921-41be-9edb-3d2764562e6a · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.197586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6d5170526578bfb10c2588a1d917c5687d303ad02b95a348d5b83060a15f204d

Observation de097274-2a17-4332-8f47-9759e0228f82 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A-okvqa: A benchmark for visual question answering using world knowledge

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.200731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:75dca7b96f198a84a8fae4a1ed893e6a8c149e8cdc50e5d68d1f3ca91d695d0f

Observation e6902d31-2215-4dc4-ac90-278be090f162 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.204226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:34c889b93f32a4732a138c784f735a4ddf35b37e45353a8db933f76c3950c7ba

Observation 837d356b-1914-4a8e-aef3-bf846d1da07f · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Textcaps: a dataset for image caption- ing with reading comprehension

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.207541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9542234b1a7d90321ef66edba963e345f5d834d4ebb462c60eef83ca87cfdf1d

Observation a7e08426-b268-4fe6-ac78-bbcd4b051f4c · outbound

This paper cites Towards vqa models that can read.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Towards vqa models that can read

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.210930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0d78342f0e11546a4c77003e4c44b324360968aee667c9881231187195f9d735

Observation 0c9fbff2-1c1d-46ed-8d07-0a09f3a29daf · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.139025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2771c9235eddd6fa896dc9ee1c2c2fe1acd6d2a57fecfb275f0cdc1a4828f315

Observation 3c66ae51-ea23-4424-a939-b9fd94be8923 · outbound

This paper cites Vi- sualmrc: Machine reading comprehension on document im- ages.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Vi- sualmrc: Machine reading comprehension on document im- ages

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.213687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3e553ec545f74bbbfaece07d0081e6fc54c00cd44dc7ca2bda2a33a971d588f1

Observation ab0672ca-c800-4850-9011-4c96415d878e · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Internlm: A multilingual language model with progressively enhanced capabilities

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.216815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:68b9a40aa091752b90db1c3adafd534474dff2732265366188abba721ffa1bc3

Observation 8b001f02-1935-40b6-b1e9-07b8d5b49dc0 · outbound

This paper cites Vicuna: An open-source chatbot impress- ing gpt-4 with 90% chatgpt quality.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Vicuna: An open-source chatbot impress- ing gpt-4 with 90% chatgpt quality

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.220502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c4dd9fc4b7364c70f47b7d20285683cedac4e2238540fe0103090c505e53e9e2

Observation d5660cc0-b6fb-4c52-a400-f2744745d926 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.021031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:efebdf083c4930ae348c48dd629df91c5b6cf39bde3dabf10cc3bdcaa5f352e7

Observation d68bc025-cdea-4ebb-a0f5-38e38cf1201c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:22:35.025946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:42dafe4f61787ba93e3c8ed74d278d64771ec1a04096d0ffa48b6028f7f1a104

Observation 6e895c53-b0ff-4461-8984-0ecf93c6461a · outbound

This paper cites All in one: Exploring unified video-language pre-training.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark All in one: Exploring unified video-language pre-training

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.224621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f666d13c640e4020e7a289b6135100d5223b5367d61be00b256e815f61f9e3db

Observation e73b2057-7119-4098-b954-52adf08ff6bc · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recogni- tion.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Temporal segment networks: Towards good practices for deep action recogni- tion

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.228698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:57dbdfde718ba7c6544f66e32c049a596fbeb068e3d3db092a557484c1281fcc

Observation 9d3a4a30-1b12-4082-99cf-9ce870ff4ee8 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Videomae v2: Scaling video masked autoencoders with dual masking

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.231827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:17e06a8093ad18eb2052ee77c5bf6e7187ccbdc8247227ac798cdaa57169b5f1

Observation 107448f9-83d7-444d-b0ef-af0b68ae0758 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.048363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:dc09adf421a127b178cd6ed9be941ed924ab3c53244faba5e38d23fcb47130b6

Observation 4ebd95a4-819d-48a7-8632-f753c3b93a26 · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.235879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6be6cc51344b6116e577191382b7cda22545a014eff05c79f06a991c4549a28d

Observation d921c129-0c8c-43f4-81b4-a572d9209444 · outbound

This paper cites Pax- ion: Patching action knowledge in video-language founda- tion models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Pax- ion: Patching action knowledge in video-language founda- tion models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.238971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:eb043b8d4fb595e08aaf907652adc388cbdc3587e19e875fde8dd06722380468

Observation 02f8d3ea-133d-4dc4-a69a-d21891a158ad · outbound

This paper cites Dai, and Quoc V.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Dai, and Quoc V

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.241492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:dad02c12a6179a9dc51aa629ed627465e20183921b308ec55554eed5b9d0a313

Observation 2dabd89e-dac9-49a6-87b8-96e0ed1551bf · outbound

This paper cites Chi, Quoc V Le, and Denny Zhou.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Chi, Quoc V Le, and Denny Zhou

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.244201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:1c058a9e8a2e7966a56170cfe29803076624780158a6bc56042288b26089a68a

Observation 2b684e64-4a4d-45f2-b4df-d9b2da82d432 · outbound

This paper cites Tenen- baum, and Chuang Gan.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Tenen- baum, and Chuang Gan

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.247221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:84c4882059363f274e14bd84cedf6ce637330d5fcab475381de89b9454ee7501

Observation 1b30f322-f650-49d0-a240-162c68223b5d · outbound

This paper cites A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.077258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6fcfebdc027deb841da1f994933a08f74cc94224c4f3c6327ffef71c2cd0cc52

Observation 7cd696e1-d891-48f1-bdbd-1e220a9dd4d9 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Next-qa: Next phase of question-answering to explaining temporal actions

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.250280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f5f77d91a0ed5eb827a7d4a736e04e4a8622d0a0227921f34b9360f7bc4b1903

Observation 7e7b6dab-6a20-4f87-b281-d19dcd8f2150 · outbound

This paper cites Video as conditional graph hierarchy for multi-granular question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video as conditional graph hierarchy for multi-granular question answering

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.253042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:57641afdb0c7ee0a0d287c73966a31dfb9a4d20a1233b52b6049744ae94f4766

Observation dccfa900-c152-417b-a495-35ffd433cc0a · outbound

This paper cites Video graph transformer for video question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video graph transformer for video question answering

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.255929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c678636a15e680627b4e6cf6db6e57390539d76b285c8107fa6773d4269ad2a5

Observation b724392e-2dc2-40f1-99df-59cc4c16f0bd · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark FunQA: Towards Surprising Video Comprehension

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.095848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c5d42f4f7b7bbbe6910746a3e5d0d771be50a0091a01a9ae9514bde2889aacaa

Observation 0b2e3657-ec4f-4b5a-a812-8c5fb40b800d · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.258556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9384007cc5e79dc01ee76c12f1e333f5c423523ececdb7ebeee27d22bb44a7c3

Observation aa04807c-dc3d-48f3-9d7e-940b3781dbb8 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Msr-vtt: A large video description dataset for bridging video and language

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.261498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:4d50ba2633abcf6715597cc715419ad4fe251bc8988e666a5f0becfa8736f93c

Observation 8315d76d-f51f-4b95-b08e-8b1953fce6aa · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.107932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:51896c74b0a1b74197a3d5f16a8ae5caae60937cadd5df7e23998fbd1ad89f55

Observation b7637b85-bb7b-4170-97b1-1bfcc189b629 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Just ask: Learning to answer questions from millions of narrated videos

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.284437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:8d20d1ec7725e9745950999aca91eb37793fec5b5ecf5a69cf197c2fc202fd73

Observation 2a7ea0bb-f1e3-46f7-a52d-746ac6284090 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Zero-shot video question answering via frozen bidirectional language models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.264671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:af398161e72f948c2d7033f020755bb9e470b52a6bcb73bf20b2f3f00b569397

Observation b2b566d4-c165-49bc-b3dd-c8336ae6afa0 · outbound

This paper cites Hitea: Hierarchical temporal- aware video-language pre-training.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Hitea: Hierarchical temporal- aware video-language pre-training

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.267646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:a76511333427066d42a5145c86519395f0434a356bc0e368a261c6085d7953c1

Observation 54661be6-1b54-4002-b669-da2f5e91851d · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.123593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:5ef358121d7c1c9d953b0710426a3c0a891e9c2b54e397cedd7bd371721d9f22

Observation 6b79f101-d438-4515-b2c7-c0540d3c4f1c · outbound

This paper cites Tenenbaum.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Tenenbaum

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.270161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ce2e419e0797174035d1a7482f5c885a4f38015dfa8c1afa11175ad6099eb6fc

Observation 88e9485e-5ce6-4160-b357-820ea91ae442 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Self-chained image-language model for video localization and question answering

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.272772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:4e00c847e83f83eb50375ce6e0c27d6875dc380c7cf33206ca07c9166470664e

Observation 0c0e51cf-69ea-43a2-a439-32182c12aad7 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.135251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:57cface33dd1af1baefe8ad75564d9b112b26ac041fc4bed9bc20aa68bafe67f

Observation f6d84558-72db-43b3-9251-f5187fdcd617 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.275559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:8e4cbb39585dd17e7be6859b924b29beeefbb8d62e5f5c254e6220590e6b216b

Observation 5fd0b994-56bd-42e3-ac4c-52e295b56df8 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.278661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:50966bfb441af865d4a4f51208b420376e628853bebb22cfc047c8a6b8e7bc4d

Observation 86529275-f921-4e96-940d-fc3d046343d9 · outbound

This paper cites Zhang, Yuxiao Dong, and Jie Tang.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Zhang, Yuxiao Dong, and Jie Tang

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.281440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0f0782f52d96957459fdde7620ba418aa40c22bf515bc2085153f5c99e18773b

Pith citing papers

Observation 2e4f28de-4321-4ddf-9b94-5d5d7ce4b5bf · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:e8f918190ef615f7c00c4a816257b75edd7ba3af5171a760e68a32af6add7371

Observation 78825fca-ed50-493d-b9f5-23a9dfe39e04 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:968d8d9d97d174fd20bea755f086d91cffc57fb71e94e590a8710d61b43ba698

Observation 24bf758a-fda9-4c5e-a094-156e14d055f9 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:0ea68f40f7e9ae33ff47c6de1560379efa5a17c917222480e478bf43755eb391

Observation 54cf1153-ec5d-4c64-bf86-2e4f89809966 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.185743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:6d8927fe3bc8ca71f3765a156077443ea5a4962d78cf1ac2655e0320c5dd0f23

Observation 29263d61-c4f2-49bf-956f-2bab10c7f3fc · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:47a295bad4543bbd6fd85a58f5d11fa913765762f15e2f1a1c81fb8c316037d3

Observation 3fa8ee51-1ac0-4c2d-b14f-feace5327b12 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 223

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.427467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:d0ebd42749db81fed6620277ffe03f9c29b6e963b6035a8cddb3a8921449f38e

Observation 05de5a87-1bcb-4955-bec0-463ccbc8413c · inbound

Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding cites this paper.

Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:47.557078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:41:47.557078Z digest=sha256:efbcbb33d14ea5048bd113b6b6e23c09f58d3f22fdbee349b086c003acdfee85

Observation 38f105f4-5c9a-4fc0-bb55-2fb2def1c153 · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:40.056441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:40.056441Z digest=sha256:ace6e0279abe21d2816d3ddb45693ce3c98d41ba7c8b99dc461ab0d4929d4098

Observation f405959f-b157-46c2-a7ad-577c02bc2f9a · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.405872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.405872Z digest=sha256:05b767719185ff4727763da89d60b64ec3dc796dcce69adb84d6ef4883aa1c84

Observation 118b31ef-01dd-4a1f-88c9-4979b156b259 · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.377380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.377380Z digest=sha256:6829cad84fe33cc0032da8634ce54791f51586b809475e421c65981a5173f2e7

Observation 9545123b-93dc-455a-a7da-6026fec7ccc5 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.460300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.460300Z digest=sha256:cb88ecd242f0d77574f3a1526cb4ed741dcf7f25562510e7c204fa1701ee7d40

Observation c3c770c7-f02e-41f2-97e0-60799e2bb24a · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:58dba8be9aa07b69984796df09433c330c5ba37e12b25f192839c69836b1270f

Observation 2806aa15-9fd4-43c2-a6cc-a52db8a1468e · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.664685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.664685Z digest=sha256:0a04f310c6561b51ec0717fc363781a7ccbe876456e79dbf91af246698a037f0

Observation 0c98bf56-fdcf-4791-9623-3cec2e25a9f9 · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:b1cc0b9242d952730f4b84e18c4debed153acb5271e1c283fdc98d23e387c7cf

Observation ded1134d-8f6e-40d4-bcd3-355b0f706f3f · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:51.806239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:51.806239Z digest=sha256:3e04b2b00fa4f6447fac90c798ecde726050baff9429f0d24d9e67b597f3658f

Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.662954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.662954Z digest=sha256:772648f9debc397f4c2c863f7666939fa6ba029647c7b61f99a0b2d866ee240b

Observation c4787365-ae0f-474c-8cbe-961b343e0487 · inbound

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior cites this paper.

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:55.854748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:55.854748Z digest=sha256:c55e83d9a0b7c2857106036b9b3e24f162927c3cfac80288d5f6b74b2bca948f

Observation 0486619b-47e8-490b-ba3d-a4906dbf350b · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:01.214250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:2c605d6f4a8bf80ccdac5350764f23f55d3db1bf393b7590e12a1bb5d9c8983d

Observation fe55c291-66a1-444b-b22a-99fd953b47ed · inbound

Promptception: How Sensitive Are Large Multimodal Models to Prompts? cites this paper.

Promptception: How Sensitive Are Large Multimodal Models to Prompts? MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:34:18.342680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:34:18.342680Z digest=sha256:00099a20617da8c0bb364a25d536d0032058285f3814139cacf51c58c752bd26

Observation dec76cd9-6cc1-4bdc-b315-c7a48ee6782e · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:8239e7e0c5cbf3ecb08875ed5f03f3c31f1f5caf09df6e296a97d0fbf9649234

Observation 30f66e08-de3e-44e7-a9d2-d0ad4ff761ee · inbound

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models cites this paper.

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T02:37:22.472754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:37:22.472754Z digest=sha256:5e6a3bbb6b1d213e35e222ac935baf053f2848afc29699f5cbc1d3c429cbcbdb

Observation 1b4f11f8-a14f-41fe-aba9-d21691f4ad69 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:634d19198c11a08551c5bc261589c87eb26f9d63a540e0895e5d55af4579951b

Observation f3b5ea87-3001-41a2-b6e4-a5b1c3405525 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:f0ddef466ff0b65569ecd5bfef89639f30019036708310f4824300be31ded6aa

Observation 5b4fec27-de6a-47d4-9f6a-aaa57272fab5 · inbound

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation cites this paper.

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T21:14:42.021240Z digest=sha256:51517ef122852036e0c42421a47e66d15a94461e48e02b7099ae2fa544f3a149

Observation 15d9a4bb-9f80-4457-a5fc-f25e92342bdf · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:3712252bc5982bc2c5cda92483bfc3b5c92125574223cae3f2adda66a78ff6eb

Observation c03de85c-b9aa-4593-9672-80643601eb04 · inbound

QoS-QoE Translation with Large Language Model cites this paper.

QoS-QoE Translation with Large Language Model MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:01:57.703684Z digest=sha256:022674e082f6491549b660be9b7285d36c5ee186066ff36fc1851de244ab27c3

Observation ea0e2140-881b-4e1c-aa8b-ffc97f098bfc · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:963447feb527414162b2c924ca94ecf67cc768e95ec95dbed9d15c89eebec422

Observation f3286376-a16a-49f5-a67c-e20e624b13ea · inbound

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ cites this paper.

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:25:40.462773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T09:08:06.592577Z digest=sha256:cff0cd980b59b3199ba22cf2eefb8cdf24607269779c434b969c2f0312c50803

Observation 6be03a38-c61e-4b96-81c1-35ca1a0d723f · inbound

FCMBench-Video: Benchmarking Document Video Intelligence cites this paper.

FCMBench-Video: Benchmarking Document Video Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T17:14:19.186123Z digest=sha256:b820cf2ea8bd5ebe0af2dba4fc28a0311efc84dfbf184bdf5024ce4db58ad3cc

Observation 641ea536-c5c4-46f2-b3ad-6ea54775e99c · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:77c22cf162e8995d50e487887888e3c35af83067245179960da0562494c7614b

Observation bc72911b-dfb6-453a-91d8-f0bdd709f19f · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:a7567a0d0fa575408c64d06ab00467b2bf366a35603a54b5b0768a5709a98116

Observation 5809c465-a775-45a4-afe9-99d7cd7493e7 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:679985a936dc0dc7325219bda81c28cbe25122b4b06233918be3c374d756155a

Observation eedde43e-b378-467f-8e9d-1b4d0196a6e8 · inbound

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs cites this paper.

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:22.278869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T05:53:05.450946Z digest=sha256:e3c10664d5e79ff2a8a43ca5fe33eea83589676ad3b282dd345f682379252afd

Observation a260fcbd-c642-4343-975e-eeb7418c5157 · inbound

AffectVerse: Emotional World Models for Multimodal Affective Computing cites this paper.

AffectVerse: Emotional World Models for Multimodal Affective Computing MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:48:05.786338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T06:46:33.612905Z digest=sha256:2378d1e2c87bfb3724edcebb88ccac2035f2e24650d9c8265446e664db215e0b

Observation a521a322-6f9e-4180-8998-a1645aca73f5 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:21:20.626071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:5fd76440d6ade8c62ef4e0794cb8434a97cb85c8b9b45ec6f0c08e17df3e8f68

Observation a2f87306-3810-4368-a92c-4b59cea0e442 · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:12:46.646665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:fbb57c5ac40d5349d7a6670d26812344075cf558a119cd7dc1a82f28f16c5a46

Observation 1e2857d9-8d63-4ab1-b226-c19bca24d5db · inbound

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR cites this paper.

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 113

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.117632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T11:50:00.954670Z digest=sha256:0af7831f3e9bb5b339269fc04cea0d213437794a33f31f6dc0253e5581b409b8

Observation 3716f920-fa42-4ad1-966c-a3105a42bbe2 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:46:56.826694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:f3d5bb2b4a3da30cd31f15577913c617fe34b143da3426cefd7d718218bb03d3

Observation fdcd9de7-c746-48ba-b93d-fb1a1da61a28 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 152

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.970891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:9d644cc24f865e921b2fcfd3ffcebe8323fc33576edb4e5acb15c45ba1d2ace0

Observation ac9d4f6c-8f74-4f01-9182-ba005ea3cb3a · inbound

NEST: Narrative Event Structures in Time for Long Video Understanding cites this paper.

NEST: Narrative Event Structures in Time for Long Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 281

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:31.178505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T17:57:55.366051Z digest=sha256:e02a255856863876d7b38eeee31c2056fb9ec6a28a8dbab693aaf25d92a2f2bd

Observation 61179361-398d-4af4-9299-cb6a3ca4c218 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 153

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:37.664149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:bbc25dd9186f43c50db2c7fa757d032857371b4e93a2128bfc8a04260ca50351

Observation 5b512191-e716-4f60-a9e1-5304ad24fb17 · inbound

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR cites this paper.

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:29:41.282601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T11:43:17.464276Z digest=sha256:786b69c430c5659e495a54adf10735c500bd73ec40a6791dac7d94edb6a7a36a

Observation a3398f72-12fc-4496-a0a0-183f4d26b4dc · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:03:32.148979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:59f543fec939b3a073f2a4d345f97391469b978627c24c6a714ad9f2d03f3087

Observation 995be05f-5f99-49d1-801c-27032a7c919d · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:54:22.384015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:542f2ec53b8106fc0ccfa52647eb7186dc54c670b985d30226f54bf9c960e3c9

Observation 51f9cb7e-7a5c-4fb8-9a69-5c661e7ff162 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:2d2302f2e99ea37fadb804fb5f5452972ea3f583f400c569d9d813a5253ce20b

Observation e3f76df0-2a30-4fb4-9156-65494a8acfce · inbound

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning cites this paper.

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T02:42:46.866557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:42:46.866557Z digest=sha256:0611710893c90046e28f5d37676efa343d2807d7364013ee5abd5147e7335be7

Observation ff6fa96a-a337-43e2-b5ac-823cdd111a8b · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.590927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.590927Z digest=sha256:e312dbdddea3924b51b2fd1eb6b6eacf3862c699b78216b21c221e432893e182

Observation a617dda4-f763-4c75-84bd-f36af7609cbb · inbound

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models cites this paper.

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:16:34.442248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:16:34.442248Z digest=sha256:566f7c9fdf9fd4546ccbe692fbb6c8c044c49667cde82e682cc234fd1592ead8

Observation d91c9d61-57eb-42aa-9c96-71709166d49e · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.025621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.025621Z digest=sha256:7233e3256903980f814fc329aaf3fe0552e9d6c48cbfe1e82ccb5ceab22757e0

Observation bd18a110-790c-4eee-bddb-ff6638e71253 · inbound

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence cites this paper.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.635521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.635521Z digest=sha256:1a133fe1fa735daa6439a03a6735d27767e3438a79bcc539ef9c6f7aa0d55071