Pith. sign in

Paper Citation Record · LEDGER

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

As of 23 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 59 inbound Pith citation observations for arXiv:2311.17005.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.17005 v4

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:22:34.954228Z

measured 159 of 159 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:00:03.498647Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:29:41.281374Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact29
  • verified fuzzy64
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47b0421e-9227-4d52-b309-3a812e2031e0 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.111518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:eea464377ebb2a3b52e9d245096fbc4ab5f156d5b205e9051d4cf608e24bd410

Observation 90f2ffa5-3409-475a-a4eb-0056df666e3f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.082058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9abce6dd690482697caf303d8dad5d081ec85e3de8ca4d7af6c9d66534577db0

Observation a781def4-bb44-416f-9b1c-4d89c54d6663 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.287880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:00c1f7bb63d05af217c4abef91530b1321513d8e23212dd3e5a2fbf9f80a196d

Observation 995854ba-d09c-4e4e-9e25-c06869744d6a · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.290504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:28bc546511103d711c3e0bbed05c671560e83ad793bb37f8592463e306134d71

Observation 52cfa3b3-a378-41d7-8322-8ffb0a02d404 · outbound

This paper cites Language models are few-shot learners.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.293232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:b63d8183ea0f005e5ae5d3d51a9f39a899b45d013568216b18bf5ae5d5b76f24

Observation c3e1ca57-0421-4234-88e3-88fda803c644 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.296457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:966735966cc9c974a741602da87e2ae6c5d2e4cca44faaebe9b72850ad3a643c

Observation 415a9afd-4921-4a85-96bb-8488c299c38e · outbound

This paper cites Chen and William B.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Chen and William B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.299565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:1b9ed6daaa88098b98a7616c73f47f2611e36a593b1d1eb3ebb0686b21874957

Observation bf002db6-6836-4ce4-97d2-0bdbc78a7c41 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.086538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:4296c3deb661cb027a49baadb8e6c83a11c3f9df2fe794d4ecb53e1cd80c4d4a

Observation 5b8838c0-e6d1-4221-b067-239d76fb8ecd · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.041933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:487088975fbc5f07ce790bb2793ce42133f6bffa8eb5900ccff7c9abd6385630

Observation 0613568e-bef9-4bb2-9dde-031db678094c · outbound

This paper cites Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.302838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:fec2b06b51b02a4ab0c7c69226908fe4a0a552bd019cb3c378d81894abbac8d2

Observation b185b69e-4a16-4aa6-a4e2-e8a43e68198a · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.305479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ca5c099bdb75fa70c4fe2b34d2cd7b47082968e1207874008f503f83e8622adc

Observation 4359a529-2860-4cb3-9520-561c8dbb8953 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.308195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c8c206ad828714a4891b8309a0a6e07eddaa29a5719a13e257d1c90b21d8b3ec

Observation c48d2239-c31a-40c2-971a-a003b75eba68 · outbound

This paper cites Doell, and Jason J.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Doell, and Jason J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.310868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:56d636627f0b7c01bcdbfd892fa047f7fbe9e0c293f314c9e880c40a180f11a5

Observation 0e34e3db-000d-468f-8351-88cef33497d4 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Imagenet: A large-scale hierarchical im- age database

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.313831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:7dfb1be800f16535b21146896869cd004dd9e08c5b9031f1ebcc573229168432

Observation 5ad86f32-7955-4db5-89af-1ec1fe8cfd93 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.143046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:206c6f6429994859688446d246324905f064c5f3a29c51a023ce7a1440733185

Observation 7d4ecb7f-b028-4a26-a347-e016f26f4a55 · outbound

This paper cites Xia, Mehdi S.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Xia, Mehdi S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.316545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6b02ab62f061543cc66c9c3e14f417c81bd4bcf66de5385bc59eab2dd001864d

Observation 490064e9-e5ef-4873-b3e8-33713a887665 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.053895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:d2ca9d9f1c6369bf8a78ce217334915a97b38592591a74df8aafddd9acb2fc19

Observation 067cc1fb-e7f2-489f-a50a-ad1df38a4057 · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.059332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:e00e704166d78866804577b8e221671cf76a5a46af516ff1daef4db714928908

Observation a851186a-6e04-4636-8230-9472ca2a0547 · outbound

This paper cites Mist : Multi-modal iterative spatial-temporal transformer for long-form video question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Mist : Multi-modal iterative spatial-temporal transformer for long-form video question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.319781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:5e40db3e9da422c3e48ab69d3ddf0ce51bd18aa7249788ef386b3e22017c38e9

Observation 3d21238a-6de4-4637-87ba-c1b3689434c6 · outbound

This paper cites Gao, Chen Sun, Zhenheng Yang, and Ramakant Nevatia.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Gao, Chen Sun, Zhenheng Yang, and Ramakant Nevatia

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.322872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:1034e19fcc5e2e535f80c0642962e381da2f7ac41e7f1c4b80833dc982bab56c

Observation 3f44ad43-9dd2-4b8c-8bd0-78657d63d7b0 · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.127729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3e167baa2c86779899f8d722aad266d324d5c8ccf015330f0764ac5021b333d8

Observation a7a0662b-e956-418c-a4ac-b0879b911ced · outbound

This paper cites something something.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark something something

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.325601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:f7e3c90630eb55ba9e22050659b1dff74c5cf83325c8876cc8be64c317aa7855

Observation 6c708d15-c26e-48a9-8c76-ea2581ac6a17 · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Making the v in vqa matter: El- evating the role of image understanding in visual question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.328577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:b6eff3bdbfb50f8a8f5c52f1d700c24488b78b7986350ecc0ba08bce2c00ebb2

Observation 3b2ce860-54ea-43a8-ab4b-40115579d15f · outbound

This paper cites Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Z.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Z

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.332073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ead134a7d0291122abeab981f8c4bf028f81585517c95117ddd0d0cd4aef9b60

Observation 066cf6d1-4f64-4f50-bc53-d3990f8098b4 · outbound

This paper cites Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.336385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:d3c473ba0d5285ed4cb6053b8031178e77c2c34bf2bd45130ea3c5672b8683ba

Observation 70db8531-0909-409d-bbf5-548bb0ad9086 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Language Is Not All You Need: Aligning Perception with Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.063626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:424945954e8e52fdf6133b28fe08830ffb727ff95351f5627dfa2e223a955385

Observation ef3c492b-4165-4354-892d-41ee5ca3db79 · outbound

This paper cites Hudson and Christopher D.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Hudson and Christopher D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.340231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:86e6176f856a036e6d16ad9fefff5b6ae5294c80e4cf8a812775b8f193c7b117

Observation d0946226-b6b3-4a07-91db-745b0bc46d52 · outbound

This paper cites Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gun- hee Kim.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gun- hee Kim

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.344812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:40c2bf5c96540d6055d22e81d4aa8edf3a7b3f94f1d9e2cc19bbc5a5f2e2b62d

Observation cd232c85-9449-46ce-a234-e511dc54e46c · outbound

This paper cites Mistral 7B.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Mistral 7B

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.103462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2ff583be69cb6af5b86e5e9ea3812148670fd42f054de9d0c0a7353b9352b1d3

Observation b7ee2ce7-d33a-4b24-83f9-95e86febe9e3 · outbound

This paper cites Lawrence Zitnick, and Ross B.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Lawrence Zitnick, and Ross B

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.348342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9ac19fca1e399372edb8b0a7faf25df2df8e482ef41786953b41a85857d02318

Observation c6867e38-db18-4f71-8bd6-8a4ee3a4dd8f · outbound

This paper cites The Kinetics Human Action Video Dataset.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark The Kinetics Human Action Video Dataset

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.131390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:787df5b6bc366a716328bdd20a146c66afb9f94b3c38102bcc8b4edd8084ec3d

Observation ae08a2f6-519e-420e-8d3b-b0d991e68af2 · outbound

This paper cites Beyond the nav-graph: Vision-and- language navigation in continuous environments.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Beyond the nav-graph: Vision-and- language navigation in continuous environments

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.352190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:67278e38a49801a72bfa8ac8e6fcc5f9894e58a7c0631ed633ff4d5f6e2cb6f1

Observation bcbeb098-4575-4909-bd39-833c1a54f38a · outbound

This paper cites A hierarchical approach for generating descriptive image paragraphs.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A hierarchical approach for generating descriptive image paragraphs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.355843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9e1509f2245a763648d3ee7533149ac897da2768551f53c59591ccb8504b9976

Observation a695e460-0da8-4aae-a0f8-ac4b3f3eb42d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.359280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:7efb40388ffa779eb4b3973324fd6511e6d4376c9d08186ec5bc7b119a123f0d

Observation ca74c5b5-1186-4cce-a794-1f3d61e0311c · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.362332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:54de73a93e9c77d9c773ad586803024d0bd4ef847f1b5ac239b1c9aa02ffb5ab

Observation 690b9d0a-9d42-4bac-af25-04d8b6a2d0af · outbound

This paper cites Moreno, and Jes ´us Lov´on-Melgarejo.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Moreno, and Jes ´us Lov´on-Melgarejo

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.365833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:cce2fd82ad0c418ac06176330715571f64c9c74bddf90080e9ab021e5bd51aba

Observation 30eeb6ab-3cc1-4e62-a226-dd766552f830 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.068109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:1a916ba7eccaa046b14fee556d9d920638fd377763a1327dbc822c85322fd125

Observation baca8caa-4445-42e7-8a03-44098fc3dba0 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.072447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:af995e76614bfcfa74ce3a6dc86a78aaf958dd98272fa520c1d80ae21e4404e4

Observation 14800046-ca1c-40ce-8666-66ef67006f86 · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.368773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2b18c9b8c8d45a7384568f5e326cf163d74a1fff762fe29c42769988267fbed5

Observation fd1694ba-3425-4fe8-b8d6-7a3d352f70be · outbound

This paper cites Inten- tqa: Context-aware video intent reasoning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Inten- tqa: Context-aware video intent reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.371437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:cfc56e2b87c5f0dd69e933d03620f95bcd628c252abd25855f24d94103033089

Observation eb5d0769-acbf-4ce3-9e46-8105f8c30d1f · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.091018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:523fe03ab774c68dc992a6c352558f89f87ed030c241f4cf90abbf084872b518

Observation 884e9606-2235-40c7-b760-7076483c38d2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VideoChat: Chat-Centric Video Understanding

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.099555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c007d22f07e587c76de446e1bca2843ad73bbc399d0cc72311c1c74ed8b0d402

Observation 5c253939-cc34-4738-ba97-ff8401600cc5 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unmasked teacher: Towards training-efficient video foundation models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.374276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:a1411b4507d3737e04c7490501c87229191189dea7aeb65c99122742a4c83d6f

Observation ffad8371-cabe-4fd4-b398-09404bcf7c7a · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.115909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:b3a29767c29d866c684e2e2e3a84b256d822b7a8dd107622899867dd649917ff

Observation 15f2eb10-6ccb-41ef-bc89-8b39ce753c70 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Evaluating Object Hallucination in Large Vision-Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.119842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c4fce60f74e539ab46e9d7179879c05986f0d7945ed80e981860384a5010bc12

Observation ac434c46-3150-4eed-9857-dcdd0b704875 · outbound

This paper cites Microsoft coco: Common objects in context.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Microsoft coco: Common objects in context

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.376770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c7631e52ad1ce6a7c4740b57dd5204c32e52af03acfc069743ee9592ed3ae6ef

Observation 60236d7c-dc73-4f1e-8d74-fd71425a3de0 · outbound

This paper cites Visual instruction tuning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Visual instruction tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.379131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0ac0de4bef451a071935c789d3b516b23442c8ebf701344c984328f720d850e3

Observation ff17c2e8-d1ca-4029-946a-8a09807addf5 · outbound

This paper cites Ntu rgb+d 120: A large-scale benchmark for 3d human activity understand- ing.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ntu rgb+d 120: A large-scale benchmark for 3d human activity understand- ing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.381831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:53053d853e30d7ed09edb014bca08bc19e2c3a47408da896867564d123ad7312

Observation 5d779c7c-2c89-45b0-aa56-dc03a5e0ff35 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MMBench: Is Your Multi-modal Model an All-around Player?

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.147202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:7505274d0564a32bebd5ffc3eed538df3064e4bb57a30747aca3f67b817018de

Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.014045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ffd31ed43ca0bab23d797f2afe7f624f5c6ce70125c3ca216bc7d57de958ba4b

Observation 706d5a7a-6d15-4eed-a84f-052294720fe6 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.030804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ef003a3cea4f0a3dbf3bb626584d977c2b7ba9a606857313677257aff9c7749e

Observation 64afd474-cd84-470f-a8c7-086f8848c8d4 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.036418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:481de02f9cd2b6ebf919d193714e75905c882b5588e26864bffc42d629a70513

Observation 608c3e73-431a-4faf-a295-280d692164e7 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.384389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:d8bc1ba1e81a2a2c07ffe9230263fd72a22c9b4699c1ef3a37e81accdb08f334

Observation c5f0ace5-ded8-4df3-9488-63f134d96c0e · outbound

This paper cites Manmatha, and C.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Manmatha, and C

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.387189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:327aa6cadf6c0b840a1b341a6fadcffc187d884dcfc80c9ecd8933de41c48fe0

Observation 8d385301-dd3a-4c13-a487-c90ce82e7b19 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ocr-vqa: Visual question answering by reading text in images

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.171192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:e266348d39cb7489be667cfc76262143f6ec002d5f18ca608928b474a14d155a

Observation 90ea304d-a46d-4428-ad42-b05a0c1866ab · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video de- scriptions.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Spoken moments: Learning joint audio-visual representations from video de- scriptions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.175006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:069b38b78cdecae9a56f956610c29ff11ab2d2e99db5fddee8a051e021451a52

Observation 0faf0da3-d353-42d0-a1a7-57b4ccf3ecfe · outbound

This paper cites Brown, Quanfu Fan, Dan Gutfreund, Carl V ondrick, and Aude Oliva.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Brown, Quanfu Fan, Dan Gutfreund, Carl V ondrick, and Aude Oliva

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.178454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:79d7617f265b90ea6858fb00b415be8c95a50af9198538cb5b49453f81e7e7ad

Observation 82c7baf6-c260-47c2-9721-59e293ce147e · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.181408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0d17452464eeb3fdff6a67076e11374a61f3d0f49888d401523cc612db79e475

Observation f9097022-6d10-46cc-94b3-4f05fff240a0 · outbound

This paper cites Gpt-4v(ision) system card.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Gpt-4v(ision) system card

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.184444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:33630af00fed2af0c1029099c29235303a1d4bf272ddd1a9de8dab3fd605fbc1

Observation b73ffd91-2fd3-46d4-99b8-b47f03135d92 · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Im2text: Describing images using 1 million captioned pho- tographs

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.187488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c190d8096651f5fe56f2f66b54a73606b834808e6ff8464b8a7f8191e85e210d

Observation 6499d076-800b-459b-8d4d-479ba6c6d609 · outbound

This paper cites Koster, Junlin Zhang, Stephanie, Winkler, Yusuf Aytar, Si- mon Osindero, Dima Damen, Andrew Zisserman, and Jo˜ao Carreira.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Koster, Junlin Zhang, Stephanie, Winkler, Yusuf Aytar, Si- mon Osindero, Dima Damen, Andrew Zisserman, and Jo˜ao Carreira

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.190584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:25f7f026964c2124dc53a2e5a50d827f20f4ce3551e7fb7852e7fa2fec79f354

Observation a39c1eec-c583-4a72-9465-f206b833dd48 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.194198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ccfc15a45b963fc8e010e802dedb9d6e9dad7ee2e453183b6a394e787c84df85

Observation d000b56b-1921-41be-9edb-3d2764562e6a · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.197586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2fbf76ce378e728f8d787fc97df19c918d9622143db9142dbe8dea3eb00a5a4b

Observation de097274-2a17-4332-8f47-9759e0228f82 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A-okvqa: A benchmark for visual question answering using world knowledge

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.200731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c09d947c4d0e87b2564545fe18c85d8e3fbfe34a74278bde3e5bf3f065f8f6e2

Observation e6902d31-2215-4dc4-ac90-278be090f162 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.204226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:62538b7e3d88ed2a407e5b24d6b719dbabc757418bfb21639915ced227c36e2f

Observation 837d356b-1914-4a8e-aef3-bf846d1da07f · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Textcaps: a dataset for image caption- ing with reading comprehension

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.207541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:424dc37330a0e73696cb30995f383f0ac87db0c9d251e14d3eb7fda1fabb75ba

Observation a7e08426-b268-4fe6-ac78-bbcd4b051f4c · outbound

This paper cites Towards vqa models that can read.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Towards vqa models that can read

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.210930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:dbe7767ca3cf290c04dd72bb92c7a474cd8d8dd42dd16292145de896ddc1f112

Observation 0c9fbff2-1c1d-46ed-8d07-0a09f3a29daf · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.139025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:fa885d80a7bc532c130705ce741754d819bda585a81dd6f679ff2912261758bc

Observation 3c66ae51-ea23-4424-a939-b9fd94be8923 · outbound

This paper cites Vi- sualmrc: Machine reading comprehension on document im- ages.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Vi- sualmrc: Machine reading comprehension on document im- ages

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.213687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0342395e3073ce842a803f2119ef9a26dfcc339251bf99f4d6cb6a7cdc1a70ec

Observation ab0672ca-c800-4850-9011-4c96415d878e · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Internlm: A multilingual language model with progressively enhanced capabilities

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.216815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:d074879d16e78799428a74d83904ae7be6a8000f429ab02b5ddcb86c9746f0ca

Observation 8b001f02-1935-40b6-b1e9-07b8d5b49dc0 · outbound

This paper cites Vicuna: An open-source chatbot impress- ing gpt-4 with 90% chatgpt quality.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Vicuna: An open-source chatbot impress- ing gpt-4 with 90% chatgpt quality

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.220502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3b1bb2eb85748817ddf1b028dd994a94dba3ce70c3829e5a08b00b80ce65a5d3

Observation d5660cc0-b6fb-4c52-a400-f2744745d926 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.021031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:02c2a326f5dc3222366cec8a541723814b8611f086a5cce018bc3b073701da10

Observation d68bc025-cdea-4ebb-a0f5-38e38cf1201c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:22:35.025946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:bb93a61d8d5e7dd7d37e4605c5e2f08744ab3edcfe4ed11d77969c923e8b3e94

Observation 6e895c53-b0ff-4461-8984-0ecf93c6461a · outbound

This paper cites All in one: Exploring unified video-language pre-training.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark All in one: Exploring unified video-language pre-training

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.224621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:23cfdc384a625532b35283ff1c40b4821d0896f16ed69965b858b1562e16a583

Observation e73b2057-7119-4098-b954-52adf08ff6bc · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recogni- tion.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Temporal segment networks: Towards good practices for deep action recogni- tion

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.228698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:61c2058382dcd893a6ae1061bdbde570714eb241bf210b32122288ae1765c704

Observation 9d3a4a30-1b12-4082-99cf-9ce870ff4ee8 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Videomae v2: Scaling video masked autoencoders with dual masking

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.231827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:988e63eacc510337b61575fb4c920183f1ebcbcfc3335f760db0575da321cdad

Observation 107448f9-83d7-444d-b0ef-af0b68ae0758 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.048363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:a790fb16ec94a0899d3ffbf18fa186cfbfce9849806b397dc1e9931e92b3d7da

Observation 4ebd95a4-819d-48a7-8632-f753c3b93a26 · outbound

This paper cites an unresolved cited work.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:22:35.235879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:0a2713241c57ce3a7bfd93c4847180f20e3bb5257a6d432b4992169068bfb49e

Observation d921c129-0c8c-43f4-81b4-a572d9209444 · outbound

This paper cites Pax- ion: Patching action knowledge in video-language founda- tion models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Pax- ion: Patching action knowledge in video-language founda- tion models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.238971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2eee9af69700730ee6b63ec5ea96be23c704764ea1d8d452d610a6f20e50433f

Observation 02f8d3ea-133d-4dc4-a69a-d21891a158ad · outbound

This paper cites Dai, and Quoc V.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Dai, and Quoc V

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.241492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:5bb16ebc1b198f4b10f48f9070377ab39d024f30c2d1cbc6235f8fc44115a769

Observation 2dabd89e-dac9-49a6-87b8-96e0ed1551bf · outbound

This paper cites Chi, Quoc V Le, and Denny Zhou.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Chi, Quoc V Le, and Denny Zhou

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.244201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:409fb07883e0d19231e0971ef7ab42a2c0b3270a13d8b934713537b2c933dbfa

Observation 2b684e64-4a4d-45f2-b4df-d9b2da82d432 · outbound

This paper cites Tenen- baum, and Chuang Gan.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Tenen- baum, and Chuang Gan

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.247221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6d6e0c0c00ed159f5b6c04240e96694de7356fe2b8bde9f8adf7d346f5dc2496

Observation 1b30f322-f650-49d0-a240-162c68223b5d · outbound

This paper cites A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.077258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:910589f515a065a409270f0f1ced57dcd40cecc4af0b3f9eca5f82e4241b01bc

Observation 7cd696e1-d891-48f1-bdbd-1e220a9dd4d9 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Next-qa: Next phase of question-answering to explaining temporal actions

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.250280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:6839eb6c41e9ec3722d82fc3c263c08c4467ed1158db3d707e939d4eff213620

Observation 7e7b6dab-6a20-4f87-b281-d19dcd8f2150 · outbound

This paper cites Video as conditional graph hierarchy for multi-granular question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video as conditional graph hierarchy for multi-granular question answering

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.253042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:9b61c82930190e6546bb81811c0721c5a41f8f290ac25aeb8a21b1bbd7937492

Observation dccfa900-c152-417b-a495-35ffd433cc0a · outbound

This paper cites Video graph transformer for video question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video graph transformer for video question answering

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.255929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2c56ef3574ba9e51ab01a8ff3e6d8c9a44aabed1da5a6f9df62e24071a9bd91e

Observation b724392e-2dc2-40f1-99df-59cc4c16f0bd · outbound

This paper cites FunQA: Towards Surprising Video Comprehension.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark FunQA: Towards Surprising Video Comprehension

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.095848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:ffae3a164b24b375b87501c244e38d71785252de102c76048b5031e0f3722045

Observation 0b2e3657-ec4f-4b5a-a812-8c5fb40b800d · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.258556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:c6697190280338269253d53c3ba031e6dd4dbab1f1647fa04fd4a3f4dcde8e70

Observation aa04807c-dc3d-48f3-9d7e-940b3781dbb8 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Msr-vtt: A large video description dataset for bridging video and language

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.261498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:4065513d710b4398b1baa0ffac35dde0e09b7f6a97e84f4ca5482ff9d319949f

Observation 8315d76d-f51f-4b95-b08e-8b1953fce6aa · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.107932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:55234dbe391b0783374ff12ccc454f87e7fbde92061a6d06828447b5758653ad

Observation b7637b85-bb7b-4170-97b1-1bfcc189b629 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Just ask: Learning to answer questions from millions of narrated videos

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.284437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:96522edcfe2f49dbcf3669bc85c90f1565966b8160a92d20a15ee5e76a9dc54a

Observation 2a7ea0bb-f1e3-46f7-a52d-746ac6284090 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Zero-shot video question answering via frozen bidirectional language models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.264671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:3df8f0f5a80783d6ac86b277350d31e2b7ce15d349337773d005b29a916eff1a

Observation b2b566d4-c165-49bc-b3dd-c8336ae6afa0 · outbound

This paper cites Hitea: Hierarchical temporal- aware video-language pre-training.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Hitea: Hierarchical temporal- aware video-language pre-training

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.267646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:4ff1c9cec2e3d73f38867cdbdee16d4c3aabbdc62a1dd67e1df5c4503ca55d7a

Observation 54661be6-1b54-4002-b669-da2f5e91851d · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.123593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:a26197fc08c56afd5d2b05a46ab7cb8870dc244d59917c86a60fb960230a7ba1

Observation 6b79f101-d438-4515-b2c7-c0540d3c4f1c · outbound

This paper cites Tenenbaum.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Tenenbaum

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.270161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:7142e661c656645e777b86a101c33d21ff9e37a4d348dd102b9d8e7cc8b5c121

Observation 88e9485e-5ce6-4160-b357-820ea91ae442 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Self-chained image-language model for video localization and question answering

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.272772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:371ca11c0973f386670cef6980580118de24ea1c7a21a24f3c23261133b4dea8

Observation 0c0e51cf-69ea-43a2-a439-32182c12aad7 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.135251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:034b2ebbc15bdc43fa1401b06f97a4506545cd7f9f72d7753efdaa2f2da41ec7

Observation f6d84558-72db-43b3-9251-f5187fdcd617 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.275559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:2d6c3db278947ffa8373715de4b32c490fdc48344b55401a5351e824cba4286d

Observation 5fd0b994-56bd-42e3-ac4c-52e295b56df8 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.278661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:98a3ea7ca5ab84280f52b78fd7f6c14ee8bb84b542e5b4729c0281d7dfd5db98

Observation 86529275-f921-4e96-940d-fc3d046343d9 · outbound

This paper cites Zhang, Yuxiao Dong, and Jie Tang.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Zhang, Yuxiao Dong, and Jie Tang

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:22:35.281440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:051f884a7af6c1a00bb4134d03640f3f840d5754ea28870c0c979f9b7242d867

Pith citing papers

Observation 2e4f28de-4321-4ddf-9b94-5d5d7ce4b5bf · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:3235e14fe7ca3ba2097fb49ce4206972eb900bfae38679ef5980df726b24e3e7

Observation 78825fca-ed50-493d-b9f5-23a9dfe39e04 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:81237b7067870fed3425f699bcf855c5c0d143a1f091c1fa791a4a0410af0f3e

Observation 24bf758a-fda9-4c5e-a094-156e14d055f9 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:b8a68c64f561f91295945dea7fa2a039ce54d1b3339f4c235d27cbd07a703b09

Observation 54cf1153-ec5d-4c64-bf86-2e4f89809966 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.185743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:9509e591df3134ad9ef69aa760a6f84086e1fae88e459a0d6e6badc01c6bd05a

Observation 29263d61-c4f2-49bf-956f-2bab10c7f3fc · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:1bd985f71ff95885a5ec25c0f030fb91411ed81735775702d8c3d0116685f9ec

Observation 3fa8ee51-1ac0-4c2d-b14f-feace5327b12 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 223

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:20:36.427467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:80a640c4ad171ba69f083e855768f3f99fbda53245524593c5795727ffc43d17

Observation ebfefeba-e6cc-4e98-ad04-5e03f150fd7a · inbound

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos cites this paper.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.663653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.663653Z digest=sha256:0c02a10d2a908af90bbbfcc8d8354322b0f94ecd87c11a809f1b8073018eeb0c

Observation 4b36e757-7f2d-47c3-b746-91d1986490f7 · inbound

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding cites this paper.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.160467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.160467Z digest=sha256:f971701791cd78e9f1a159a8da20af113650552d53fd9d55848e312050cda7b4

Observation ae669222-29c1-4638-a93e-276574568f4a · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.333026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.333026Z digest=sha256:82f8ba7c986470dd8a359f178a7373c8e1c5b27ba931c0c84e4d572cb3a9d0a5

Observation a357ad93-e830-427d-832e-6f66742c8546 · inbound

Neptune: The Long Orbit to Benchmarking Long Video Understanding cites this paper.

Neptune: The Long Orbit to Benchmarking Long Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:29.246089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:58:29.246089Z digest=sha256:de0271d1802caec504778264f8364f4ba59d41d6f3b91a7c120b10c43ddef218

Observation e5eb1154-0f5e-4f3b-97a6-4c520c7c4bb5 · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.183821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.183821Z digest=sha256:fa1aaf353b12b79faf5873bf133b4bfc54e14b3ce9c0c3fd0d1011de674e7f7f

Observation 05de5a87-1bcb-4955-bec0-463ccbc8413c · inbound

Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding cites this paper.

Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:41:47.557078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:41:47.557078Z digest=sha256:e4faea3693034757e3e047760587dd0041bf1999b35abdf29d69afb5d0022bff

Observation f6571f36-6847-467d-b570-ca49b7c2753d · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 235

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.166701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.166701Z digest=sha256:49416548cf91512f109896cea82ef3c9054659be4c415f53a0677d122d382923

Observation 38f105f4-5c9a-4fc0-bb55-2fb2def1c153 · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:40.056441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:40.056441Z digest=sha256:4cbf6816906ecda059eaf50602665876b19801c00080c4f96089f34439bd87d3

Observation f405959f-b157-46c2-a7ad-577c02bc2f9a · inbound

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models cites this paper.

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:55.405872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:55.405872Z digest=sha256:3d3f9d5cd74c24581b66657453ab5fcf2058863ba29d32b91ca0529a1c073558

Observation 118b31ef-01dd-4a1f-88c9-4979b156b259 · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.377380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.377380Z digest=sha256:5f0a605e945850c9db6451a645a7db6120c26ef36b458cbac7ce04c360303c7f

Observation 9545123b-93dc-455a-a7da-6026fec7ccc5 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.460300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.460300Z digest=sha256:5ea79eedc5a369e1f4e11bddcbb53bf6833e7d8e6075160fc5e384eb6cdbd77b

Observation c3c770c7-f02e-41f2-97e0-60799e2bb24a · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:d188f6c42b06ecaf1c3c5896607d73a91705efdbc2f575a3c9a9411fe4e7be4a

Observation 2806aa15-9fd4-43c2-a6cc-a52db8a1468e · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.664685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.664685Z digest=sha256:f884698b3ee9b920dd8d84f2d16dc4fb933a8262e8032d894b1fabb8a964b4d2

Observation 0c98bf56-fdcf-4791-9623-3cec2e25a9f9 · inbound

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning cites this paper.

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T15:18:43.724432Z digest=sha256:561fa4046fd3dd4c2e26116d42058fb3436ab5236ca71f5915af0ee54561c378

Observation c9a7165a-78f0-4522-9a70-64d9b771a3a3 · inbound

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension cites this paper.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.498647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.498647Z digest=sha256:e02d229c3c96f1489072d7fb7aca5cfdbc243b5a85460a94800669496dd4e938

Observation ded1134d-8f6e-40d4-bcd3-355b0f706f3f · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:51.806239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:51.806239Z digest=sha256:6bc3233c67d35c8a0bcbb03372553cfa377d5459ac7932a0eb697f1fc30d36f3

Observation d490d203-fc5f-444c-adb2-d1056550f63e · inbound

PEVLM: Parallel Encoding for Vision-Language Models cites this paper.

PEVLM: Parallel Encoding for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:37:58.984224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:37:58.984224Z digest=sha256:794ab263d31ef8fd39dc2c36d6512d2007f93719cb1fd332dfd05f18b53ea806

Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.662954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.662954Z digest=sha256:9bf813f8daa38223eb6e7d8c5f218be225875dc43429b5c30005097746863b56

Observation c4787365-ae0f-474c-8cbe-961b343e0487 · inbound

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior cites this paper.

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:55.854748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:55.854748Z digest=sha256:3b051550727c0f4eeff5afabaffd0f074cfbb2c288b5c115ca5eaefc5a35fbce

Observation 0486619b-47e8-490b-ba3d-a4906dbf350b · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:01.214250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:a3a232d6f4f188fdfd7ac31da30b0b65ebc9e223102ccf670bfcd130f5944eff

Observation fe55c291-66a1-444b-b22a-99fd953b47ed · inbound

Promptception: How Sensitive Are Large Multimodal Models to Prompts? cites this paper.

Promptception: How Sensitive Are Large Multimodal Models to Prompts? MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:34:18.342680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:34:18.342680Z digest=sha256:9cc74f8a6065100e57a202277a315501f39493fee0dd15228a78116ca8fd75c9

Observation dec76cd9-6cc1-4bdc-b315-c7a48ee6782e · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:7a0bb9748fb09c97b16ef9b47018d591dd5793344c8346d2770c6b2d28811eff

Observation 30f66e08-de3e-44e7-a9d2-d0ad4ff761ee · inbound

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models cites this paper.

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T02:37:22.472754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:37:22.472754Z digest=sha256:bdfe217550e92cb544817e9a97dd836659034573b3422df473df8d9573cef995

Observation 1b4f11f8-a14f-41fe-aba9-d21691f4ad69 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:3f03ef030bd93edd7464631f8edfb111edb43157a3f76b20e7b4bff95a6ae18a

Observation f3b5ea87-3001-41a2-b6e4-a5b1c3405525 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:ee8c9595fab61c821067f4d56c617d14156002dbfdb353c90a7c52b949f5ee76

Observation 5b4fec27-de6a-47d4-9f6a-aaa57272fab5 · inbound

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation cites this paper.

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T21:14:42.021240Z digest=sha256:30f12b06681b1ab2ccc8d2da16246d25170cce2e9b7d354c4f6eb73944c19465

Observation 15d9a4bb-9f80-4457-a5fc-f25e92342bdf · inbound

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding cites this paper.

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:21:47.439019Z digest=sha256:6613d3dadb887ef640adca779f6a95e7cb778f38e512b3f673ff49b3cc8e74c0

Observation c03de85c-b9aa-4593-9672-80643601eb04 · inbound

QoS-QoE Translation with Large Language Model cites this paper.

QoS-QoE Translation with Large Language Model MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:01:57.703684Z digest=sha256:04c19072382aaa56b1ce10d28e8cdbfcad0b0928278c4c1d1c0cad2492c17379

Observation ea0e2140-881b-4e1c-aa8b-ffc97f098bfc · inbound

EasyVideoR1: Easier RL for Video Understanding cites this paper.

EasyVideoR1: Easier RL for Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T07:41:27.231098Z digest=sha256:30bb762e0809c997f5ad26df396c9cb930a2924b3c515d394d4f204f038686ba

Observation f3286376-a16a-49f5-a67c-e20e624b13ea · inbound

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ cites this paper.

The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:25:40.462773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T09:08:06.592577Z digest=sha256:853cd20625ab89419da186b92382dd641a2f0579d17696f8b4e4d033c6c42ffa

Observation 6be03a38-c61e-4b96-81c1-35ca1a0d723f · inbound

FCMBench-Video: Benchmarking Document Video Intelligence cites this paper.

FCMBench-Video: Benchmarking Document Video Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T17:14:19.186123Z digest=sha256:1ea9e0643b8db24dabd8669fcd5b9a047dd3110779be3bf4114b172a077efe75

Observation 641ea536-c5c4-46f2-b3ad-6ea54775e99c · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:76d2b5fc453153a50c32f30ddd1ab92928c567b287accb6e53e6f6587a6b8f5c

Observation bc72911b-dfb6-453a-91d8-f0bdd709f19f · inbound

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models cites this paper.

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T01:30:15.463051Z digest=sha256:efacfbe017bdc75d40ad484a27d71e3a41a1569c28986c9f662305b522c6bc8a

Observation 5809c465-a775-45a4-afe9-99d7cd7493e7 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:22:35.388231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:fb12ba4dc7f3a72c04a9c2a80a5e4a185dbe5c7869f35f94303f9cbb995fe8fd

Observation eedde43e-b378-467f-8e9d-1b4d0196a6e8 · inbound

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs cites this paper.

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:22.278869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T05:53:05.450946Z digest=sha256:5ce6aec4e78f5328cf64a53f3dac4a027a271056100ea03d96b92ccfbbadd266

Observation a260fcbd-c642-4343-975e-eeb7418c5157 · inbound

AffectVerse: Emotional World Models for Multimodal Affective Computing cites this paper.

AffectVerse: Emotional World Models for Multimodal Affective Computing MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:48:05.786338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T06:46:33.612905Z digest=sha256:a4c4064ed63e972d23eceb81d6104f93b3707ee0ac115ae4ad47c7684a04d617

Observation a521a322-6f9e-4180-8998-a1645aca73f5 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:21:20.626071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:ce698e83c0ba680e45ad3d897e0896ad6291fafe4c8a85b4007b0b34d0f4ef4d

Observation a2f87306-3810-4368-a92c-4b59cea0e442 · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:12:46.646665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:a32de9fc86c858e85ce6132520033de4968f18edb990565d73ab7c176d0b6180

Observation 1e2857d9-8d63-4ab1-b226-c19bca24d5db · inbound

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR cites this paper.

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 113

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.117632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T11:50:00.954670Z digest=sha256:344860cf562577ad93e6ab7ed579ef8c39e69341907a838b95443ec6562ed6bc

Observation 3716f920-fa42-4ad1-966c-a3105a42bbe2 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:46:56.826694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:7ab033ec07131e51fd606c21595e3bac73deffb4fdcbd698aa3a9e91baec1022

Observation fdcd9de7-c746-48ba-b93d-fb1a1da61a28 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 152

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.970891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:b0a0e3f19a2002d4f7f4d19cb28208766aa003de48b34adc915f6a4d01e93041

Observation ac9d4f6c-8f74-4f01-9182-ba005ea3cb3a · inbound

NEST: Narrative Event Structures in Time for Long Video Understanding cites this paper.

NEST: Narrative Event Structures in Time for Long Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 281

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:31.178505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T17:57:55.366051Z digest=sha256:c577c679e1c614a35b0e4e022c35d5b4bb44e888b9c7696f54a22e3f765140f2

Observation 61179361-398d-4af4-9299-cb6a3ca4c218 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 153

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:37.664149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:ad9bde9ce2986d74d0b489a98dee1467fb72903727de6de969ca55cfcac49f81

Observation 5b512191-e716-4f60-a9e1-5304ad24fb17 · inbound

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR cites this paper.

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:29:41.282601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T11:43:17.464276Z digest=sha256:98a0b91545b25c1de99a1faced48d5d5d5a86f80ff2e0beff311ae3fe688b54e

Observation a3398f72-12fc-4496-a0a0-183f4d26b4dc · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:03:32.148979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:c435fe647a5ed5df442e63da5fe8369d2a26234f8b0027ce231dd5deb7cd0964

Observation 995be05f-5f99-49d1-801c-27032a7c919d · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:54:22.384015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:c8649dcda59436d7b94175d582f32e2fd00f9bd4bcae970bf2005265b347e8fc

Observation 51f9cb7e-7a5c-4fb8-9a69-5c661e7ff162 · inbound

TimeThink: Reasoning with Time for Video LLMs cites this paper.

TimeThink: Reasoning with Time for Video LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T08:59:46.244502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:59:46.244502Z digest=sha256:c33f79fbc29a14bce44763e332f330877bff2aaa1e0da01d9067b6ec0037b467

Observation e3f76df0-2a30-4fb4-9156-65494a8acfce · inbound

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning cites this paper.

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T02:42:46.866557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:42:46.866557Z digest=sha256:d90663bf001979cc1c91d5b5e4c0a161052282557bc3bd38865d7acc6268bc46

Observation ff6fa96a-a337-43e2-b5ac-823cdd111a8b · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.590927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.590927Z digest=sha256:df79c8151a6866c54f05497c141deb35fb001c03c9b335ad63d3e7ec83697781

Observation a617dda4-f763-4c75-84bd-f36af7609cbb · inbound

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models cites this paper.

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T20:16:34.442248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:16:34.442248Z digest=sha256:8f42e16c679c4e5735024ca3de49b3c9baa5f7ea575d0109427997099c960eaf

Observation ed72f624-49b6-46a2-a63c-fd922167502a · inbound

DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models cites this paper.

DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:43:16.002980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:43:16.002980Z digest=sha256:3a7a479b3af58c6c98ffcf709925710734785df8e76e5d972aab2e51cf23ed44

Observation d91c9d61-57eb-42aa-9c96-71709166d49e · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.025621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.025621Z digest=sha256:e6f2d12c5b3c8ff59bb68fb56447a13ea2603c2d866fb47293fe73ae3d7b307a

Observation bd18a110-790c-4eee-bddb-ff6638e71253 · inbound

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence cites this paper.

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:34.635521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:34.635521Z digest=sha256:308e16759bc3874a337b2b517664ff1cc5a7596bfaf1a2eef4035223619bd6d6