Pith. sign in

Paper Citation Record · LEDGER

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 10 inbound Pith citation observations for arXiv:2505.14640.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14640 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:44.884301Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:02.846860Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:07:17.426534Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a88210c-b53b-4908-b4e2-4d404068d158 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LVBench: An Extreme Long Video Understanding Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.208786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.208786Z digest=sha256:a967750bfe88f706e7d7af311751ef7c0821cb8982ce91440a30ae994c1302b9

Observation 10ed969e-dfad-49c1-9133-7e0750fa635e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.317680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.317680Z digest=sha256:1a5f7b7aca98d8c7c25ec367217e27acd17cdbfa31c958d27d91c44019122eec

Observation 1e7ad542-9a07-4b88-bba4-3cd40cdc2b28 · outbound

This paper cites Video Anomaly Detection and Explanation via Large Language Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video Anomaly Detection and Explanation via Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.477549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.477549Z digest=sha256:5cad907fcd3bdf16ff3e2b1cd6cf87432b3d9277f5b3a5e9ef32567617ecf5b2

Observation 9d5f0c6b-03c9-4506-9af8-f0464c479552 · outbound

This paper cites Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.586160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.586160Z digest=sha256:f166461650134a20cf66aeb3df6df1396f0e77f0b8fecf7eba43ac16912b0cd6

Observation e77b6f11-b337-4d0a-9a24-bfdfca92ecb7 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Towards automatic learning of procedures from web instructional videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.675590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.675590Z digest=sha256:79aed327b21099a15b16c48a78fc4cf96ceefaf2e396b1387f64ce47839b2f4f

Observation fbb17e75-ffea-475a-947c-10ab364bd87b · outbound

This paper cites Long Context Transfer from Language to Vision.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Long Context Transfer from Language to Vision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.769437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.769437Z digest=sha256:fcaa8fa2b456358cea4d74d9222f091e1adfb2d3090d21f97a1617b181290c89

Observation f52197b2-3d0b-4a6e-8276-e907aa5a7935 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.848837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.848837Z digest=sha256:11430d0dadc09e174df625c926941468c19559090d91dc081704a72387425340

Observation 0e55e230-5270-4ea2-ba9d-10a36c347294 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:40.968075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:40.968075Z digest=sha256:68e33e57ab1711bac40361116615eea2efb63dfe68f57e81609997e2fba47b72

Observation a97207d3-7a08-4b04-b7be-0cd06df75772 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.079838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.079838Z digest=sha256:4d46265f395425fbcdbe791d80f8f1d01b9f19e66cb6c514ff69876ebdf90064

Observation 08eef1c9-d3c6-4b8e-a355-ced9c4eaa905 · outbound

This paper cites Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.182215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.182215Z digest=sha256:aabc7f8acc34554ba246fda2f559e2ff836af5368e7a3e7d53aed8bae166f1af

Observation 809ec41f-47d5-4412-a93c-ca8d346ecae4 · outbound

This paper cites Token-efficient long video understanding for multimodal llms.arXiv preprint arXiv:2503.04130, 2025.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Token-efficient long video understanding for multimodal llms.arXiv preprint arXiv:2503.04130, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.295575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.295575Z digest=sha256:13a45adbf4036f07aa3067e285147a5df038c4f98c86429ce7f910ab10ea7b1d

Observation 3ac6fee9-ec53-406f-b2b4-f7c2e83cff95 · outbound

This paper cites BIMBA: Selective-Scan Compression for Long-Range Video Question Answering.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.423639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.423639Z digest=sha256:2a370a65b1291268346c7cae1a3202848c48dcd5e5f041a2aadb4c999acce0dd

Observation aece33d9-7349-46da-b737-66bdd2dbf49b · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.495843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.495843Z digest=sha256:00e1313d4cf39399506f3cce636bb84c2761298ce443b890aa1f00a96dddde03

Observation c716fa61-58fb-4624-89ce-fe750b5d2a98 · outbound

This paper cites VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.595618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.595618Z digest=sha256:fe23fa3ff9e5b191ab66798918b71e7d45203f26a733facbe1f2063d8f62831c

Observation 34216a65-fbda-4284-9c4c-1b6854591650 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.669333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.669333Z digest=sha256:571d9bbf1aa0f6aff28b34a6785fac4389d6c08b24b5a7ab73e530d7863035fc

Observation 157d2750-e5aa-4ea4-9f87-c7e1e62a4782 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.743812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.743812Z digest=sha256:ce2689b8df3b0bf4efcd9561b360e4ebf19ea18b9afab7fb59037766ff2d6b21

Observation 4f4b19fd-09e3-4e4d-95d7-a4aaf5e427f6 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.886033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.886033Z digest=sha256:2acdd5bac86545c43bbb63b7b04cffdb77e585a57f625ea59a07f54a1a31c5e0

Observation 3e33eeb5-d633-4f8a-b2ed-99519e4ae733 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:41.992087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:41.992087Z digest=sha256:9d53609f9bc99606aea78e71de9287d21daf4ffd89ce138fd78722b8b136f98e

Observation 5d257b92-d56b-4914-867c-72506d73fc7e · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation MLVU: Benchmarking Multi-task Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.108152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.108152Z digest=sha256:c77ad72843db1b35333d6b8d3f2d72d7ac73bdd3ba47cef76bc3e0258b1650a9

Observation c0f84637-01cb-4874-a219-6fbf78ea032e · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.202827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.202827Z digest=sha256:2c47f548f6e946f26b08395824cda0fc4aa7b94dd559f4d8e0b88b62834199d7

Observation 4f3d0a31-25bb-492d-a763-2c477de4a226 · outbound

This paper cites Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Hourvideo: 1-hour video- language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:46.268970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:42.274637Z digest=sha256:161ad728b249c8ec9a7f5143e0fcc83a6ca880eaa0f377f0fdc2d420edbba44e

Observation 4cbf0ae1-1f36-4988-99b6-e62437f5ba40 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.341438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.341438Z digest=sha256:52d315b0bb3a5a232290e8bd9e9ea3f2374543e682cde4b8d3c8dd1be5f3924b

Observation 5f383376-799c-40e8-b0ff-10aae3869e41 · outbound

This paper cites Qwen2.5-VL Technical Report.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Qwen2.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.438272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.438272Z digest=sha256:f64995f222db2a4ae0c492169dd80ba55f9a04bd3e413bb71d4e3655c40e131f

Observation 3f9b5974-c0ef-464f-93fd-7d8253ab8b18 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.546538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.546538Z digest=sha256:d8cff89c3defd5bf3cee4aa3794e8ce97f4ec5f9c355a0de656acccbf2846e4d

Observation 3fe5c5d6-7f0e-4d57-a531-0314a6b3bbd8 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.650236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.650236Z digest=sha256:5b0c6d7fcc23d22f620689960133a96635c747c97ae76e37b0ddf9e35f99584b

Observation f9db9ddf-06ec-471d-8a9e-933d2d1fbcb8 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.725167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.725167Z digest=sha256:c247c972ba4eed78faa2bc5b1f9e8bcd39b76c13d8a20a583efde52cf285f49c

Observation cc29b80a-afa0-4295-88e0-6ee1ce4abb53 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.813436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.813436Z digest=sha256:cb6290ded2a322686993acf0cb91f49b71221c18e87e809fa02410336ec0c8af

Observation 911ec23a-c43c-4dbf-989f-a33908344c6f · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.886209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.886209Z digest=sha256:0d74369cacb8a754e09ea3daf192ead4c7d9616312662e279c12aac19d48a92f

Observation ce17acaf-da91-44b8-a33a-65f9f5f53e0c · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:42.981605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:42.981605Z digest=sha256:39a6d086de03298ee315b5d5906b6cbe083ff4c17e575bc157bc4ed7cdf478e1

Observation b0cc6776-42c0-4274-841d-926965e00b77 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.096309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.096309Z digest=sha256:f5ee59a0b6559ddf434868b172d64d108fd1fcb97cbde4e6530ee97aa496e1b7

Observation 4cac2fae-867a-4ca3-81ec-80b4cd26c169 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.173823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.173823Z digest=sha256:1ac6ef651008241601223254f7e1b88cf58b0093f59a3cbc001cdf4bb3562edf

Observation 5523b6de-9a63-4e83-8221-c14dd59cc6bc · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.251039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.251039Z digest=sha256:3de4e5aa223864533b8d68d2649d4f0b85ea58c079c61f6b54a42e06c6d06c26

Observation a56cf553-118f-4410-91a7-10e529619f69 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.353762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.353762Z digest=sha256:513b7076d05d6b944653b7f801ac2c55f2b070b48e136557da55db3ebe75c91a

Observation 7d083842-2e43-4ed1-97ad-f4013fb07b65 · outbound

This paper cites Vript: A video is worth thousands of words.Advances in Neural Information Processing Systems, 37:57240–57261, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Vript: A video is worth thousands of words.Advances in Neural Information Processing Systems, 37:57240–57261, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.440428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.440428Z digest=sha256:cce8e37a957116e5570d56f9f1cc82c09f5e039571bf3ab8f75b610706acd76a

Observation 782cb605-d43c-49b8-be75-4e6046766571 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.506440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.506440Z digest=sha256:966671948969f67b18ec7c434404a918037f139fa3ef518f1d2d8f414b91c765

Observation 085e1bbd-1e7e-441c-bca8-6507ba764d3f · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video question answering via gradually refined attention over appearance and motion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.574670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.574670Z digest=sha256:8df978574716e14ff12155a0476b3278e9404ebdd3918f02ff5dd1fd538a6634

Observation 0d2e9a99-d3cd-4e3a-bbe4-0395d87eacb7 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.650831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.650831Z digest=sha256:3940ac957a92a9a6dfd600ede0536c5ca2bf9c1d5925ec867929eab94cb9e5ab

Observation 647e48b4-0e88-4c5d-82ce-b6601f34a514 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation TempCompass: Do Video LLMs Really Understand Videos?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.723261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.723261Z digest=sha256:a4d1705aff41b5839db2dcde4fc32e145a9edc862474c8f84a325b42820c0885

Observation 29aaa453-09ca-4625-84b3-d9457014c874 · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.788047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.788047Z digest=sha256:3cee0a0d7ab97f77aa63d6eb98ae8e03e159fe491573636e28df3d7bd3d0a5ae

Observation cee2d8f7-c7d7-4872-8ffd-955bf72fc76f · outbound

This paper cites Measuring short-form factuality in large language models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Measuring short-form factuality in large language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.846476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.846476Z digest=sha256:01e7dfa3512dc2e5bb6e429c605a219c3076921b2f4f092220ebed1227b3de86

Observation 6ab987bd-e7c2-4885-ba42-fd30ab5413d9 · outbound

This paper cites Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.911417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.911417Z digest=sha256:c54a6b639e53eea603d729916d574e78bf20b11e098cfd7efd28343639489cf5

Observation e73cd584-1302-45ee-9b4f-ad87622b26ea · outbound

This paper cites A Survey on LLM-as-a-Judge.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation A Survey on LLM-as-a-Judge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.002274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.002274Z digest=sha256:d58c029e3e7b1a05280cb94e6451544f39feaba4f68b496bf669eeccefda16e9

Observation 490bec7a-661d-4ac6-85fa-759897b69dbd · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.072681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.072681Z digest=sha256:c2b445eb1a73906711c7945f492594d954cb521570575242589a5ae5e7c48947

Observation c76730be-8774-46a2-a7c2-37b02a43f9a7 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:46.047734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:44.129970Z digest=sha256:7ce3f7770dde9cfcd0ec9ca206e73523cea5bb85f1ce38d9eed345c0d845a2b5

Observation 4d87df07-319c-4926-b206-0ca0809dfd24 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence, July 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gpt-4o mini: Advancing cost-efficient intelligence, July 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.877992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:44.192407Z digest=sha256:d2d033081ec02454ad1526d8ed6dbc352828439342f84f55b0e5f27d23dc2942

Observation d0f93d1e-fae4-425a-bff1-0880eecedd1f · outbound

This paper cites Introducing gpt-4.1 in the api, April 2025.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Introducing gpt-4.1 in the api, April 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.790494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:44.250379Z digest=sha256:867fedd0783366e60c2be8b587c656a28dce8205982a632689acaa3e11a553dd

Observation f669b362-f506-4ab3-97ad-60abaef86525 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, March.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Gemini 2.5: Our most intelligent ai model, March

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.725060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:44.319389Z digest=sha256:486b2f674dbfab3d46a46062e0c30a4af837f32b57b7462cb485d561b76c22e6

Observation 1bebcfb6-ddfd-44f2-a173-13b10dc86f4f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.468889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.468889Z digest=sha256:5ee9fa91ac3bc5b126f775f2594fedcaaace3170c3cce214e801ed01cdb65367

Observation 5d0872ad-c2e1-40e5-a047-d42e50b38f36 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.518133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.518133Z digest=sha256:5b8a79c57737ebdb9b8de982bcdcfb8d54083502e4c198c0c2f63d5e7692c6a9

Observation e2f50624-c532-4090-9107-576c2a735bad · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.591263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.591263Z digest=sha256:79b7b75243fd4b1a57f76763b93fec633786e6f1921401e4ec3d9f5c2f08fca0

Observation fb22c9ce-78de-42ef-8966-21a2949b6f12 · outbound

This paper cites Phi-4 Technical Report.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Phi-4 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.670107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.670107Z digest=sha256:bb34b6bbac85b4ff5106b065a64bd94e793f5e20f7d7a34c82126099ac517872

Observation 25113abc-91cd-44f7-93f6-2d069f44df5a · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.746446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.746446Z digest=sha256:7f7576641843b9cdfbb34362090ece1129582d090fe5350a10a03ef1d186714c

Observation f16b9d63-0c2a-4bd7-8b01-f283f01c91d8 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:44.818600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:44.818600Z digest=sha256:4b59f41f46c5217623d99dc4a2e0780cdba533db75e85ffc42bd4f6404f1d991

Observation ff54a88f-14c1-4d97-958e-95fd6b6afaba · outbound

This paper cites Keep" if the question can be answered by someone who has watched the video, even if the answer requires reasoning or summarizing visual or auditory evidence. -.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Keep" if the question can be answered by someone who has watched the video, even if the answer requires reasoning or summarizing visual or auditory evidence. -

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.565691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:44.884301Z digest=sha256:581c4fcc98ab4002b44478dabe09fc7a9563cdddc51b6df2553859470a24de19

Observation 0a5fdf07-81df-48d0-b405-ba843d7a7362 · outbound

This paper cites Accessed: 2025-05-08.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation Accessed: 2025-05-08

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:45.668732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:34:44.396406Z digest=sha256:26ec5fcc99cb1035e3f874d72591a4d818d1ef36e1fa139f0389d53801ed86c5

Pith citing papers

Observation a0e8be73-74c8-40ed-ac7e-74a5cf6f213b · inbound

Vid-SME: Membership Inference Attacks against Large Video Understanding Models cites this paper.

Vid-SME: Membership Inference Attacks against Large Video Understanding Models VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:02.846860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:02.846860Z digest=sha256:d1ecaeeb3ad772ce07f827ec448163c69318871c2a73fddd7ea87ff68a8b26ab

Observation c7fc560d-b42f-45f3-a437-4be1ade2a1d1 · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.304314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.304314Z digest=sha256:d1d02b72ad653bf12590c4a3a0da4455fc825efd664aac78ce6e03ef39031ea4

Observation 39eb555c-eb3a-4d13-b80d-f9c4fac7983e · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.796098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:38baeb34ceae4b6564f5b12ca50c4eebc99266bf1aa8d24589f08d6b93c0baa9

Observation 47adbc6b-9538-4066-9e0a-a2ccd64bb11c · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:16:31.811999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:11:52.694778Z digest=sha256:20a98871a6281b64e1e7fc5463f4deb483254726fb70723d1d159450d1187a41

Observation 920533a0-f3fe-4684-a31e-1a4b0b8223d9 · inbound

TrajTok: Learning Trajectory Tokens enables better Video Understanding cites this paper.

TrajTok: Learning Trajectory Tokens enables better Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:40.952686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:40.952686Z digest=sha256:d22562a43de40a0b49f8ada8ce35f4ecffd166e43a22aacf555e444e43d4320f

Observation 13235850-057d-420d-b4cf-99c338778fd2 · inbound

Video-Oasis: Rethinking Evaluation of Video Understanding cites this paper.

Video-Oasis: Rethinking Evaluation of Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T15:38:41.390945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:38:41.390945Z digest=sha256:15b710867d0e37ff0ac420b06c4486c949c03688d16f4a6fd586edf97d51dbd8

Observation 4b5eeacb-ba77-4cfc-97f9-0dba3b9c2840 · inbound

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors cites this paper.

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:25.138117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:39:01.254916Z digest=sha256:a92bb92120c3800c4d630c5d91a6783112ee905dc089af0056766ddece0cda74

Observation 280e7169-00b3-4883-bce5-1fc4884e04b7 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.427951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:b5fbc8bb0799d7b71e988ba4cea573601751f82bcea6a722389599048d2d09c9

Observation 84ee4e48-c29d-4129-aa2e-0182b72b4e4e · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:46.835407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:46.835407Z digest=sha256:d4af34e0bb239f8267ef076d18925b1f4cc13b809336446bc8c5b6c2440343f5

Observation 40587f60-b205-4071-bfe4-bd9aa258a220 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 155

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.256778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.256778Z digest=sha256:fc145643c343ec9e097b8e47c23f16d05f917b6cd333e02f3f85c3fd3a49aa43