Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.917814Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 235 outbound references and 32 inbound Pith citation observations for arXiv:2504.13180.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.917814Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:35:44.870369Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 235 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 16ca75ab-6316-46f7-87ef-10347eaf7f55 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Visual instruction tuning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a048876-d40d-4740-97ab-02b68cc5099d · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b611945f-069b-4791-9049-12a434cfeff1 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Sharegpt4v: Improving large multi-modal models with better captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a525e0e-bb21-4904-b92e-29df26792509 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Finevideo: behind the scenes, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c492a6bb-b5d3-4bc5-87e5-ac11f57d5a76 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video instruction tuning with synthetic data, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c44d3654-2579-4e20-b1e4-fb6f56135dc3 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 325ab784-ef84-425f-92c9-a3c7369a4249 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43398dab-a72d-4441-b1e2-55da1f3df5d2 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f6e097-9e1a-4291-9a8a-0ac7eeac1f9e · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Where does it exist: Spatio-temporal video grounding for multi-form sentences
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1af8518-3d95-479e-8906-9f88ae1a770f · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a6f1ea-4eb4-4391-b086-a6c76dc81130 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cfed3d-5642-4c01-a041-1649df3fb22f · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142ca518-f2e4-4386-9f7b-9f524b8f3bb6 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92330894-b3d3-4abe-b053-4f782f2550dc · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 108f75cf-f89d-4de6-8385-6eba40d76e9c · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b337e501-00b9-4dd9-b1d4-deac0b119bbe · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Flamingo: a Visual Language Model for Few-Shot Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27fea407-90ae-4867-a924-29b541458a28 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vila: On pre- training for visual language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 830a7ec6-e930-45ed-9711-0f60cdbed056 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3988dee-b4ee-4bf1-ad25-ca63bd140d40 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d3b4f6-dbe0-4ad2-a728-7944bf161e53 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a44779-8703-4016-99c9-c504bc4ec177 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoChat: Chat-Centric Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aa03cf5-150d-41a0-bfb9-9eb7bdf42f01 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8098cc-cf6b-4254-ad36-9324668d7b77 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9767c1c-77f6-449e-8e2b-03b2b93abe5d · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59e0e807-d82b-402a-842e-9e2b9a188f77 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7b1ea1-0e59-4378-bfe5-814968b28754 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Longvlm: Efficient long video understanding via large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e779e7f-4f47-466b-9110-9869485177f9 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5dbe1b-4619-4bd7-a977-98491a6aab08 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fab2509-b331-4909-a6cc-0558d4851560 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Anymal: An efficient and scalable any-modality augmented language model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 000576c0-9e8e-48ab-b6e4-9586afe5f8b4 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0bb014-6a46-406a-a79a-3156e534898d · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1862ffb-7313-4922-bce3-aaba6a4eda55 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gpt-4v(ision) system card, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338771fa-3acd-446f-bf72-03e0d55cbdb9 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gpt-4o system card, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f2af48-f529-42b6-9720-0c198ed5e422 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 940f928e-1b35-4234-b7db-4386b4679e9c · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11900aaa-fc6c-4aa7-a5d1-174a2b804596 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding The claude 3 model family: Opus, sonnet, haiku
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c513db-7dfc-4a03-8fb9-0c44daddaa26 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1480b762-335e-4750-93ba-9af59537700a · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a66b0859-6aa2-4deb-a6c8-7dbe8c9448bb · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding A-okvqa: A benchmark for visual question answering using world knowledge, 2022
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a069c1-855d-4703-92cf-de4b0b2154bd · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vizwiz: nearly real-time answers to visual questions
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec5040a-9dd1-4619-b2ac-7cee39082a1e · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68a8351-1612-4a5a-87ae-24d735c84c96 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d40bcc-f888-44b8-b6f0-c21f3066f810 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b666324-d314-4bcb-9635-c0a985489d67 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Blink: Multimodal large language models can see but not perceive
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524002db-2b40-4b4e-9d14-50f16a6f5e8a · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Realworldqa benchmark
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11cfde4b-e409-404d-a7fc-9fa47b040793 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c25d4d-1d58-473a-98bd-914753cb556f · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 369c0858-dd28-4ac2-9d44-11c325b56693 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f6672c5-1ff3-480b-a74f-fa1c8e916947 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Microsoft coco: Common objects in context
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1506e30-27c7-4c7b-bfb5-cc69f77f826d · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Nocaps: Novel object captioning at scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73ebd77-8a9d-4c79-99f3-48a282097279 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c714fac2-7c3c-4f18-bb3e-05cdaa8dee75 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Towards vqa models that can read
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7a51e6-8026-4ef4-8dc2-1d0433f71800 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c43630a-20ec-4d84-a900-3e0e80d9c64f · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Advancing Chart Question Answering with Robust Chart Component Recognition
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation de527b50-6dbc-42b0-b065-316309f9b995 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding A diagram is worth a dozen images, 2016
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f64dd8-f3d2-46e3-923f-499ee0a04cf9 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be40560-095c-491e-9eea-0e5f9b3b3599 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150637a6-f450-4ba0-88be-2bbf45db1a25 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd37f51-f0fd-4f8b-8147-8f7e4cecb931 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aceb6ef3-34d4-49b2-8196-2d1aea9d99f3 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding From recognition to cognition: Visual commonsense reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70de624-4752-41e8-b45b-e32d1a8f81e0 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4246225f-b79b-4929-995e-e7f9123c6bc3 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186, 2025
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3ceeac-ffca-461d-8a67-d4095ac5b3b1 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed44cf2-e708-48a7-b235-535dcb997b59 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 750f8264-b226-4935-b8ae-ca15a37c60bd · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f746fb6-72f7-48a5-9544-af59859f16b3 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f32b7c-508a-4a05-88fe-ca0109527ca1 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5822dbea-4472-43ad-af67-ebd6fc50fbd3 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Evaluating Object Hallucination in Large Vision-Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92d91a9-eed8-4b6e-bb55-8b476e87e987 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Next-qa: Next phase of question-answering to explaining temporal actions
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d48eb5-8296-4901-9411-78349347a32b · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d399e52-fd78-4879-af69-ec2ca6b00792 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Perception test: A diagnostic benchmark for multimodal video models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6c9ba23-dfd0-416a-86bc-897eb47a4de9 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Star: A benchmark for situated reasoning in real-world videos
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c749c94c-1f2b-4c7f-9d96-db90b140502c · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Tgif-qa: Toward spatio- temporal reasoning in visual question answering
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950d872f-70aa-491e-ab8b-9a6241880bdd · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding TVQA: Localized, Compositional Video Question Answering
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a16e8d0c-ca51-4034-9d34-8f68480ce1d5 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8843b2a-2896-4094-b6ea-0a3e5c1062b1 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99046912-b16a-4197-9d4e-81770bcf636e · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f04011-2e93-4114-a36c-edf06bd2dc7b · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab3da26-6402-45d4-aecf-bf17ed42d31b · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4cdb90-95e4-4514-9f5a-af4e447517b5 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3cb6c5c-ddec-4d68-8ebf-7798c51c6767 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Msr-vtt: A large video description dataset for bridging video and language
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14823033-ad53-4c92-9094-18d84c93a1ae · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Collecting highly parallel data for paraphrase evaluation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab77670a-6332-4abf-91d1-a345d88ece21 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Towards automatic learning of procedures from web instructional videos
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d279b101-bf4c-4bd8-94ac-18a0610e604d · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c503fa1c-0df3-4971-b880-a5973c11678d · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Dense-captioning events in videos
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc87307-2b8a-4df9-b86e-d9a42f27b391 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24faaf20-db2e-40d4-8566-e87b52a3e902 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39150a31-3698-4398-b3eb-dc88c042dacf · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec2fe659-1e9f-4afc-8c91-c673d2b17594 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Eventhallusion: Diagnosing event hallucinations in video llms
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f2f325-5907-4713-8906-369d2cdd9d25 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d9bd916-b723-42d8-add9-51debe881cef · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f7aaed4-1cf8-4b8d-8b51-692555e12207 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding LVBench: An Extreme Long Video Understanding Benchmark
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3c861c-1a44-475d-9f38-e90afe065029 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Movieqa: Understanding stories in movies through question-answering
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6d55e2-6d24-42ea-b596-0d77650b80a5 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828– 28857, 2025
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0cc2070-28b6-4b01-a9d2-7f5f2fd686ac · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Moviechat: From dense token to sparse memory for long video understanding
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80e9632d-b4d5-4aaa-9e94-687a4322bdcf · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding MLVU: Benchmarking Multi-task Long Video Understanding
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4cd5b5c-dcb8-464a-b26e-baa7c1ec1267 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d565190c-c7ad-4575-ad02-0d6476b5a4a0 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9658c1c3-82cd-45bc-8251-b58c942b3496 · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b008b2-5910-4ab5-a082-1123a2b09afa · outbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f51e9682-09fb-47b3-8924-32619f16dc8f · inbound
Perception Encoder: The best visual embeddings are not at the output of the network PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 36bbce6e-2e67-485a-90eb-aae6f5bd0c0e · inbound
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb10324-1938-47ff-a4ff-fc8f6981218f · inbound
BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab285409-aff3-450c-b2d9-708aa346803f · inbound
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5481e072-022d-4fb6-974c-a1f73ccbf903 · inbound
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 254383df-acec-4f27-a082-f36dbfe6d5b6 · inbound
Group Relative Augmentation for Data Efficient Action Detection PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 048c6c3c-bdd1-4fc5-86b3-6ebce32f8450 · inbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b872af4-c4ba-4404-9fd9-9132e9fe6cae · inbound
SAM 3: Segment Anything with Concepts PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 28b1ff73-f9ff-4cb7-8a2a-d35a64972510 · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c87468a0-646d-48e9-baee-b10aa5b2c2b6 · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5638f5c9-e25e-49de-bf42-b0e8b754930e · inbound
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3ab36bd7-1861-4b62-bfd4-6b1b14d6faa5 · inbound
InstrAct: Towards Action-Centric Understanding in Instructional Videos PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 303a81e6-0a18-414c-a1da-9dcd0f3f3662 · inbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 90b75ce0-2758-4e5f-b709-91725c335095 · inbound
Building a Precise Video Language with Human-AI Oversight PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f0a3f67e-6673-468f-b957-ee4e6c81beb8 · inbound
Don't Pause! Every prediction matters in a streaming video PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d1335f95-8139-4879-9e57-123cf6c50c59 · inbound
ZAYA1-VL-8B Technical Report PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8542f057-34bb-4891-9e1f-a8d89a7d3f16 · inbound
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a1e13608-5d35-46fa-88cf-ec4009c487b9 · inbound
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 883c0a7a-fb6d-484a-9887-2490acff76e2 · inbound
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fcf13a7e-7ada-4806-b127-201796c1e010 · inbound
Zamba2-VL Technical Report PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation df891798-c9bd-4693-95c3-3f58cc7d1c13 · inbound
AdaCodec: A Predictive Visual Code for Video MLLMs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1765e026-7c55-4e84-8cd2-5ce4315e9981 · inbound
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 84eb08e2-7c7b-4151-adec-6a76caa51cbb · inbound
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 256df914-618f-4b03-9e49-28a8ceef1490 · inbound
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 73f2b3c5-bdef-455f-ac15-b2fe8c5fe847 · inbound
DiffusionBench: On Holistic Evaluation of Diffusion Transformers PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 479012f2-b78e-486c-b73e-fc149cfee7d5 · inbound
SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9bc7a457-a71c-40a0-b29b-1d5a706ad9e7 · inbound
Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 070e822c-dcc7-4a43-8b9f-8a33109367ae · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 115430e1-3ed7-4c6d-be2e-554e0056aa28 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 22c24511-54df-4f63-ba2a-427fe098fd85 · inbound
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0d9a6198-a0f5-4191-9a08-107af9fc74e1 · inbound
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e667046c-a725-4642-b587-c2328b08047c · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.