Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T20:22:34.954228Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 50 inbound Pith citation observations for arXiv:2311.17005.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T20:22:34.954228Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:41:47.557078Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T08:29:41.281374Z
100 of 104 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 47b0421e-9227-4d52-b309-3a812e2031e0 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Flamingo: a Visual Language Model for Few-Shot Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90f2ffa5-3409-475a-a4eb-0056df666e3f · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a781def4-bb44-416f-9b1c-4d89c54d6663 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 995854ba-d09c-4e4e-9e25-c06869744d6a · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 52cfa3b3-a378-41d7-8322-8ffb0a02d404 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Language models are few-shot learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c3e1ca57-0421-4234-88e3-88fda803c644 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 415a9afd-4921-4a85-96bb-8488c299c38e · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Chen and William B
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bf002db6-6836-4ce4-97d2-0bdbc78a7c41 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5b8838c0-e6d1-4221-b067-239d76fb8ecd · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0613568e-bef9-4bb2-9dde-031db678094c · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Shazeer, Vinodkumar Prab- hakaran, Emily Reif, Nan Du, Benton C
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b185b69e-4a16-4aa6-a4e2-e8a43e68198a · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4359a529-2860-4cb3-9520-561c8dbb8953 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c48d2239-c31a-40c2-971a-a003b75eba68 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Doell, and Jason J
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e34e3db-000d-468f-8351-88cef33497d4 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Imagenet: A large-scale hierarchical im- age database
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5ad86f32-7955-4db5-89af-1ec1fe8cfd93 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7d4ecb7f-b028-4a26-a347-e016f26f4a55 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Xia, Mehdi S
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 490064e9-e5ef-4873-b3e8-33713a887665 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 067cc1fb-e7f2-489f-a50a-ad1df38a4057 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a851186a-6e04-4636-8230-9472ca2a0547 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Mist : Multi-modal iterative spatial-temporal transformer for long-form video question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d21238a-6de4-4637-87ba-c1b3689434c6 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Gao, Chen Sun, Zhenheng Yang, and Ramakant Nevatia
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3f44ad43-9dd2-4b8c-8bd0-78657d63d7b0 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7a0662b-e956-418c-a4ac-b0879b911ced · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark something something
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c708d15-c26e-48a9-8c76-ea2581ac6a17 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Making the v in vqa matter: El- evating the role of image understanding in visual question answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3b2ce860-54ea-43a8-ab4b-40115579d15f · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Z
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 066cf6d1-4f64-4f50-bc53-d3990f8098b4 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 70db8531-0909-409d-bbf5-548bb0ad9086 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Language Is Not All You Need: Aligning Perception with Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef3c492b-4165-4354-892d-41ee5ca3db79 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Hudson and Christopher D
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d0946226-b6b3-4a07-91db-745b0bc46d52 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gun- hee Kim
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd232c85-9449-46ce-a234-e511dc54e46c · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Mistral 7B
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b7ee2ce7-d33a-4b24-83f9-95e86febe9e3 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Lawrence Zitnick, and Ross B
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6867e38-db18-4f71-8bd6-8a4ee3a4dd8f · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark The Kinetics Human Action Video Dataset
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae08a2f6-519e-420e-8d3b-b0d991e68af2 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Beyond the nav-graph: Vision-and- language navigation in continuous environments
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bcbeb098-4575-4909-bd39-833c1a54f38a · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A hierarchical approach for generating descriptive image paragraphs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a695e460-0da8-4aae-a0f8-ac4b3f3eb42d · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca74c5b5-1186-4cce-a794-1f3d61e0311c · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 690b9d0a-9d42-4bac-af25-04d8b6a2d0af · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Moreno, and Jes ´us Lov´on-Melgarejo
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30eeb6ab-3cc1-4e62-a226-dd766552f830 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation baca8caa-4445-42e7-8a03-44098fc3dba0 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 14800046-ca1c-40ce-8666-66ef67006f86 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd1694ba-3425-4fe8-b8d6-7a3d352f70be · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Inten- tqa: Context-aware video intent reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb5d0769-acbf-4ce3-9e46-8105f8c30d1f · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 884e9606-2235-40c7-b760-7076483c38d2 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VideoChat: Chat-Centric Video Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5c253939-cc34-4738-ba97-ff8401600cc5 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unmasked teacher: Towards training-efficient video foundation models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ffad8371-cabe-4fd4-b398-09404bcf7c7a · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15f2eb10-6ccb-41ef-bc89-8b39ce753c70 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Evaluating Object Hallucination in Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac434c46-3150-4eed-9857-dcdd0b704875 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Microsoft coco: Common objects in context
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 60236d7c-dc73-4f1e-8d74-fd71425a3de0 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Visual instruction tuning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ff17c2e8-d1ca-4029-946a-8a09807addf5 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ntu rgb+d 120: A large-scale benchmark for 3d human activity understand- ing
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d779c7c-2c89-45b0-aa56-dc03a5e0ff35 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MMBench: Is Your Multi-modal Model an All-around Player?
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 706d5a7a-6d15-4eed-a84f-052294720fe6 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 64afd474-cd84-470f-a8c7-086f8848c8d4 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 608c3e73-431a-4faf-a295-280d692164e7 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5f0ace5-ded8-4df3-9488-63f134d96c0e · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Manmatha, and C
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8d385301-dd3a-4c13-a487-c90ce82e7b19 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Ocr-vqa: Visual question answering by reading text in images
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90ea304d-a46d-4428-ad42-b05a0c1866ab · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Spoken moments: Learning joint audio-visual representations from video de- scriptions
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0faf0da3-d353-42d0-a1a7-57b4ccf3ecfe · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Brown, Quanfu Fan, Dan Gutfreund, Carl V ondrick, and Aude Oliva
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82c7baf6-c260-47c2-9721-59e293ce147e · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f9097022-6d10-46cc-94b3-4f05fff240a0 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Gpt-4v(ision) system card
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b73ffd91-2fd3-46d4-99b8-b47f03135d92 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Im2text: Describing images using 1 million captioned pho- tographs
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6499d076-800b-459b-8d4d-479ba6c6d609 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Koster, Junlin Zhang, Stephanie, Winkler, Yusuf Aytar, Si- mon Osindero, Dima Damen, Andrew Zisserman, and Jo˜ao Carreira
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a39c1eec-c583-4a72-9465-f206b833dd48 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d000b56b-1921-41be-9edb-3d2764562e6a · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation de097274-2a17-4332-8f47-9759e0228f82 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A-okvqa: A benchmark for visual question answering using world knowledge
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e6902d31-2215-4dc4-ac90-278be090f162 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 837d356b-1914-4a8e-aef3-bf846d1da07f · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Textcaps: a dataset for image caption- ing with reading comprehension
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a7e08426-b268-4fe6-ac78-bbcd4b051f4c · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Towards vqa models that can read
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c9fbff2-1c1d-46ed-8d07-0a09f3a29daf · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c66ae51-ea23-4424-a939-b9fd94be8923 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Vi- sualmrc: Machine reading comprehension on document im- ages
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab0672ca-c800-4850-9011-4c96415d878e · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Internlm: A multilingual language model with progressively enhanced capabilities
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8b001f02-1935-40b6-b1e9-07b8d5b49dc0 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Vicuna: An open-source chatbot impress- ing gpt-4 with 90% chatgpt quality
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5660cc0-b6fb-4c52-a400-f2744745d926 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark LLaMA: Open and Efficient Foundation Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d68bc025-cdea-4ebb-a0f5-38e38cf1201c · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e895c53-b0ff-4461-8984-0ecf93c6461a · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark All in one: Exploring unified video-language pre-training
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e73b2057-7119-4098-b954-52adf08ff6bc · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Temporal segment networks: Towards good practices for deep action recogni- tion
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9d3a4a30-1b12-4082-99cf-9ce870ff4ee8 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Videomae v2: Scaling video masked autoencoders with dual masking
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 107448f9-83d7-444d-b0ef-af0b68ae0758 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4ebd95a4-819d-48a7-8632-f753c3b93a26 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d921c129-0c8c-43f4-81b4-a572d9209444 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Pax- ion: Patching action knowledge in video-language founda- tion models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 02f8d3ea-133d-4dc4-a69a-d21891a158ad · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Dai, and Quoc V
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2dabd89e-dac9-49a6-87b8-96e0ed1551bf · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Chi, Quoc V Le, and Denny Zhou
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2b684e64-4a4d-45f2-b4df-d9b2da82d432 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Tenen- baum, and Chuang Gan
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b30f322-f650-49d0-a240-162c68223b5d · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7cd696e1-d891-48f1-bdbd-1e220a9dd4d9 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Next-qa: Next phase of question-answering to explaining temporal actions
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7e7b6dab-6a20-4f87-b281-d19dcd8f2150 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video as conditional graph hierarchy for multi-granular question answering
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dccfa900-c152-417b-a495-35ffd433cc0a · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video graph transformer for video question answering
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b724392e-2dc2-40f1-99df-59cc4c16f0bd · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark FunQA: Towards Surprising Video Comprehension
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b2e3657-ec4f-4b5a-a812-8c5fb40b800d · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aa04807c-dc3d-48f3-9d7e-940b3781dbb8 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Msr-vtt: A large video description dataset for bridging video and language
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8315d76d-f51f-4b95-b08e-8b1953fce6aa · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b7637b85-bb7b-4170-97b1-1bfcc189b629 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Just ask: Learning to answer questions from millions of narrated videos
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a7ea0bb-f1e3-46f7-a52d-746ac6284090 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Zero-shot video question answering via frozen bidirectional language models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2b566d4-c165-49bc-b3dd-c8336ae6afa0 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Hitea: Hierarchical temporal- aware video-language pre-training
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54661be6-1b54-4002-b669-da2f5e91851d · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b79f101-d438-4515-b2c7-c0540d3c4f1c · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Tenenbaum
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 88e9485e-5ce6-4160-b357-820ea91ae442 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Self-chained image-language model for video localization and question answering
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c0e51cf-69ea-43a2-a439-32182c12aad7 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f6d84558-72db-43b3-9251-f5187fdcd617 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5fd0b994-56bd-42e3-ac4c-52e295b56df8 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 86529275-f921-4e96-940d-fc3d046343d9 · outbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Zhang, Yuxiao Dong, and Jie Tang
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2e4f28de-4321-4ddf-9b94-5d5d7ce4b5bf · inbound
TempCompass: Do Video LLMs Really Understand Videos? MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78825fca-ed50-493d-b9f5-23a9dfe39e04 · inbound
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24bf758a-fda9-4c5e-a094-156e14d055f9 · inbound
MLVU: Benchmarking Multi-task Long Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54cf1153-ec5d-4c64-bf86-2e4f89809966 · inbound
LVBench: An Extreme Long Video Understanding Benchmark MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 29263d61-c4f2-49bf-956f-2bab10c7f3fc · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fa8ee51-1ac0-4c2d-b14f-feace5327b12 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 223
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05de5a87-1bcb-4955-bec0-463ccbc8413c · inbound
Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f105f4-5c9a-4fc0-bb55-2fb2def1c153 · inbound
Online Video Understanding: OVBench and VideoChat-Online MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f405959f-b157-46c2-a7ad-577c02bc2f9a · inbound
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118b31ef-01dd-4a1f-88c9-4979b156b259 · inbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9545123b-93dc-455a-a7da-6026fec7ccc5 · inbound
Visual Large Language Models for Generalized and Specialized Applications MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 147
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3c770c7-f02e-41f2-97e0-60799e2bb24a · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2806aa15-9fd4-43c2-a6cc-a52db8a1468e · inbound
$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c98bf56-fdcf-4791-9623-3cec2e25a9f9 · inbound
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ded1134d-8f6e-40d4-bcd3-355b0f706f3f · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · inbound
Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4787365-ae0f-474c-8cbe-961b343e0487 · inbound
PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0486619b-47e8-490b-ba3d-a4906dbf350b · inbound
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fe55c291-66a1-444b-b22a-99fd953b47ed · inbound
Promptception: How Sensitive Are Large Multimodal Models to Prompts? MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec76cd9-6cc1-4bdc-b315-c7a48ee6782e · inbound
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 30f66e08-de3e-44e7-a9d2-d0ad4ff761ee · inbound
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4f11f8-a14f-41fe-aba9-d21691f4ad69 · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f3b5ea87-3001-41a2-b6e4-a5b1c3405525 · inbound
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4fec27-de6a-47d4-9f6a-aaa57272fab5 · inbound
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15d9a4bb-9f80-4457-a5fc-f25e92342bdf · inbound
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c03de85c-b9aa-4593-9672-80643601eb04 · inbound
QoS-QoE Translation with Large Language Model MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea0e2140-881b-4e1c-aa8b-ffc97f098bfc · inbound
EasyVideoR1: Easier RL for Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f3286376-a16a-49f5-a67c-e20e624b13ea · inbound
The category of Whittaker modules over the Cartan Type Lie algebra $\bar{S}_2$ MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6be03a38-c61e-4b96-81c1-35ca1a0d723f · inbound
FCMBench-Video: Benchmarking Document Video Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 641ea536-c5c4-46f2-b3ad-6ea54775e99c · inbound
Co-Evolving Policy Distillation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc72911b-dfb6-453a-91d8-f0bdd709f19f · inbound
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5809c465-a775-45a4-afe9-99d7cd7493e7 · inbound
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eedde43e-b378-467f-8e9d-1b4d0196a6e8 · inbound
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a260fcbd-c642-4343-975e-eeb7418c5157 · inbound
AffectVerse: Emotional World Models for Multimodal Affective Computing MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a521a322-6f9e-4180-8998-a1645aca73f5 · inbound
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a2f87306-3810-4368-a92c-4b59cea0e442 · inbound
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e2857d9-8d63-4ab1-b226-c19bca24d5db · inbound
Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3716f920-fa42-4ad1-966c-a3105a42bbe2 · inbound
VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fdcd9de7-c746-48ba-b93d-fb1a1da61a28 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac9d4f6c-8f74-4f01-9182-ba005ea3cb3a · inbound
NEST: Narrative Event Structures in Time for Long Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 281
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 61179361-398d-4af4-9299-cb6a3ca4c218 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5b512191-e716-4f60-a9e1-5304ad24fb17 · inbound
When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3398f72-12fc-4496-a0a0-183f4d26b4dc · inbound
TuringViT: Making SOTA Vision Transformers Accessible to All MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 995be05f-5f99-49d1-801c-27032a7c919d · inbound
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51f9cb7e-7a5c-4fb8-9a69-5c661e7ff162 · inbound
TimeThink: Reasoning with Time for Video LLMs MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3f76df0-2a30-4fb4-9156-65494a8acfce · inbound
MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6fa96a-a337-43e2-b5ac-823cdd111a8b · inbound
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a617dda4-f763-4c75-84bd-f36af7609cbb · inbound
CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91c9d61-57eb-42aa-9c96-71709166d49e · inbound
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd18a110-790c-4eee-bddb-ff6638e71253 · inbound
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.