Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:44.884301Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 10 inbound Pith citation observations for arXiv:2505.14640.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:44.884301Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:02.846860Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T19:07:17.426534Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6a88210c-b53b-4908-b4e2-4d404068d158 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LVBench: An Extreme Long Video Understanding Benchmark
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10ed969e-dfad-49c1-9133-7e0750fa635e · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7ad542-9a07-4b88-bba4-3cd40cdc2b28 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video Anomaly Detection and Explanation via Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5f0c6b-03c9-4506-9af8-f0464c479552 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77b6f11-b337-4d0a-9a24-bfdfca92ecb7 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Towards automatic learning of procedures from web instructional videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb17e75-ffea-475a-947c-10ab364bd87b · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Long Context Transfer from Language to Vision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f52197b2-3d0b-4a6e-8276-e907aa5a7935 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e55e230-5270-4ea2-ba9d-10a36c347294 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a97207d3-7a08-4b04-b7be-0cd06df75772 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08eef1c9-d3c6-4b8e-a355-ced9c4eaa905 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 809ec41f-47d5-4412-a93c-ca8d346ecae4 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Token-efficient long video understanding for multimodal llms.arXiv preprint arXiv:2503.04130, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac6fee9-ec53-406f-b2b4-f7c2e83cff95 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aece33d9-7349-46da-b737-66bdd2dbf49b · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c716fa61-58fb-4624-89ce-fe750b5d2a98 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34216a65-fbda-4284-9c4c-1b6854591650 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 157d2750-e5aa-4ea4-9f87-c7e1e62a4782 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4b19fd-09e3-4e4d-95d7-a4aaf5e427f6 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e33eeb5-d633-4f8a-b2ed-99519e4ae733 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d257b92-d56b-4914-867c-72506d73fc7e · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation MLVU: Benchmarking Multi-task Long Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f84637-01cb-4874-a219-6fbf78ea032e · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f3d0a31-25bb-492d-a763-2c477de4a226 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4cbf0ae1-1f36-4988-99b6-e62437f5ba40 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f383376-799c-40e8-b0ff-10aae3869e41 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Qwen2.5-VL Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f9b5974-c0ef-464f-93fd-7d8253ab8b18 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat: Chat-Centric Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe5c5d6-7f0e-4d57-a531-0314a6b3bbd8 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9db9ddf-06ec-471d-8a9e-933d2d1fbcb8 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc29b80a-afa0-4295-88e0-6ee1ce4abb53 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911ec23a-c43c-4dbf-989f-a33908344c6f · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce17acaf-da91-44b8-a33a-65f9f5f53e0c · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0cc6776-42c0-4274-841d-926965e00b77 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cac2fae-867a-4ca3-81ec-80b4cd26c169 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5523b6de-9a63-4e83-8221-c14dd59cc6bc · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56cf553-118f-4410-91a7-10e529619f69 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d083842-2e43-4ed1-97ad-f4013fb07b65 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Vript: A video is worth thousands of words.Advances in Neural Information Processing Systems, 37:57240–57261, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 782cb605-d43c-49b8-be75-4e6046766571 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 085e1bbd-1e7e-441c-bca8-6507ba764d3f · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video question answering via gradually refined attention over appearance and motion
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2e9a99-d3cd-4e3a-bbe4-0395d87eacb7 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647e48b4-0e88-4c5d-82ce-b6601f34a514 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation TempCompass: Do Video LLMs Really Understand Videos?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29aaa453-09ca-4625-84b3-d9457014c874 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee2d8f7-c7d7-4872-8ffd-955bf72fc76f · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Measuring short-form factuality in large language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab987bd-e7c2-4885-ba42-fd30ab5413d9 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73cd584-1302-45ee-9b4f-ad87622b26ea · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation A Survey on LLM-as-a-Judge
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490bec7a-661d-4ac6-85fa-759897b69dbd · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76730be-8774-46a2-a7c2-37b02a43f9a7 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d87df07-319c-4926-b206-0ca0809dfd24 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gpt-4o mini: Advancing cost-efficient intelligence, July 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0f93d1e-fae4-425a-bff1-0880eecedd1f · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Introducing gpt-4.1 in the api, April 2025
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f669b362-f506-4ab3-97ad-60abaef86525 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gemini 2.5: Our most intelligent ai model, March
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1bebcfb6-ddfd-44f2-a173-13b10dc86f4f · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0872ad-c2e1-40e5-a047-d42e50b38f36 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2f50624-c532-4090-9107-576c2a735bad · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb22c9ce-78de-42ef-8966-21a2949b6f12 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Phi-4 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25113abc-91cd-44f7-93f6-2d069f44df5a · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture.arXiv preprint arXiv:2409.02889, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16b9d63-0c2a-4bd7-8b01-f283f01c91d8 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff54a88f-14c1-4d97-958e-95fd6b6afaba · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Keep" if the question can be answered by someone who has watched the video, even if the answer requires reasoning or summarizing visual or auditory evidence. -
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a5fdf07-81df-48d0-b405-ba843d7a7362 · outbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Accessed: 2025-05-08
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a0e8be73-74c8-40ed-ac7e-74a5cf6f213b · inbound
Vid-SME: Membership Inference Attacks against Large Video Understanding Models VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7fc560d-b42f-45f3-a437-4be1ade2a1d1 · inbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39eb555c-eb3a-4d13-b80d-f9c4fac7983e · inbound
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 47adbc6b-9538-4066-9e0a-a2ccd64bb11c · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 920533a0-f3fe-4684-a31e-1a4b0b8223d9 · inbound
TrajTok: Learning Trajectory Tokens enables better Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13235850-057d-420d-b4cf-99c338778fd2 · inbound
Video-Oasis: Rethinking Evaluation of Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5eeacb-ba77-4cfc-97f9-0dba3b9c2840 · inbound
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 280e7169-00b3-4883-bce5-1fc4884e04b7 · inbound
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84ee4e48-c29d-4129-aa2e-0182b72b4e4e · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40587f60-b205-4071-bfe4-bd9aa258a220 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
Reference 155
Source-reported events for the cited work
Unavailable: canonical work link unavailable.