Pith. sign in

Paper Citation Record · LEDGER

Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2311.16103.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.16103 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:07:09.556330Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.654958Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a496b85-6cc0-457c-81f2-327b2ba5c964 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:41.875900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:f44c05a8604aac982d7009597c8a128ae9409c69e9a33166fa5d019b04801f7a

Observation eeefd3dd-f688-43d6-80e2-93659ab49552 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.779312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:13495d2cb070df56558164819e23eb7bd7231cadfb6dba1fe7f3125621b4fecc

Observation 0fe0559f-075c-4d14-a087-1de07c63ef1d · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.421249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:ee7e0b624639b6cbc3f94107d2548d8050581908764bc67e559d376450f6715a

Observation 4380602a-a411-4aec-a9ef-e9139b5d7aa6 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.564426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:fc8ff1b7a5a8c6f73b757901659e8c87ff6e272b14236d09b1629659221c3211

Observation e5477cd8-9c92-4497-8925-47646329c903 · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.556330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.556330Z digest=sha256:dfc9ec769a7d262e5b06c837ee5fc4f7b7cd2cd31f60608f30fed4679b5c6fc9

Observation 081f7dd8-e3a6-4ac6-a1ad-3459c8ad70bb · inbound

On the Consistency of Video Large Language Models in Temporal Comprehension cites this paper.

On the Consistency of Video Large Language Models in Temporal Comprehension Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:00.809662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:07:00.809662Z digest=sha256:26fa95c74c5aec24ff1006f91a6098b7db5f24380d5362430191b4e2ce2fceaa

Observation 75f5f241-f1d3-4427-b93b-433224b632db · inbound

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs cites this paper.

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:12.002715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T16:57:12.821916Z digest=sha256:c01cf2ba0bf6ddbd22d6f6598d39f71b851f2f6ebd4d8509249df1c3488be323

Observation e2b48806-8901-43d2-9c9c-310312ee36ff · inbound

Neptune: The Long Orbit to Benchmarking Long Video Understanding cites this paper.

Neptune: The Long Orbit to Benchmarking Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:29.460476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:58:29.460476Z digest=sha256:08489899724211c6c60ab63772d5c8aae4a621167dab5beb8ac15cea483e3d20

Observation facc0219-2a24-4ea2-a415-bd5a42daf1f1 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.311193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.311193Z digest=sha256:112e43f59d160d7bdba3bb9294565894dd505e8adf96d642b3c9b22eff1bfea9

Observation 0d49752f-6432-41a9-b1a4-dcbdc4b942ad · inbound

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding cites this paper.

CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:22:51.012539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:22:51.012539Z digest=sha256:aec146b0e08b44a36979764008f181e7378be54f3cce428cb34e8f58282ba6b4

Observation 59d1c639-d981-4237-b57d-e4f8c904d5b8 · inbound

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces cites this paper.

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:27:44.178010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T09:27:43.919941Z digest=sha256:5f8ae898d93b8ed7be785ca968ef8b1fbd3fb06c9a7e80ae2dea111c4fce9abb

Observation 65f5c485-f92e-4fb4-a78c-7f68d539b65a · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.328262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.328262Z digest=sha256:25052c44510af897b72d132382f2e20cea4455c20cfe1aa7b065e02db246e75b

Observation 790bf743-95bc-4de3-be39-e8ff36a81b80 · inbound

SCBench: A Sports Commentary Benchmark for Video LLMs cites this paper.

SCBench: A Sports Commentary Benchmark for Video LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.780895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.780895Z digest=sha256:978052ba39d0701a887206d08536c50d0a4c0396360e0a9b5de67f30dc9e4bf9

Observation bbd32194-8fee-41c4-9768-d7b7d9e0d75e · inbound

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs cites this paper.

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:53.440780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:53.440780Z digest=sha256:fde0d2febc92a276092e13c3d0fed4cee9adf6ab841c118a4fc7b0efb1ad3358

Observation d26d89f2-222c-48f3-b906-17ba2aab3661 · inbound

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding cites this paper.

MMVU: Measuring Expert-Level Multi-Discipline Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T17:15:23.823071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:15:23.823071Z digest=sha256:b3eb793193c667e77d86f29afdd52e7400d9fa9a0447cad2a3c5eca2d27bcca6

Observation c06a4d2c-520e-4fbb-aff2-933b4f80fb63 · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.171272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:56e0adb6756d688306e1127201244ece7dab825c73d93cd91c43d086900d75fd

Observation 832a929d-004e-4f4f-9261-e952cac7cfb0 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.388526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:1a11b0c6cfd39f85746766e1cfc37e5b7f24b2a22384955f67b5573606933c92

Observation c9c0aa41-0ace-4d52-82ef-682ec8b0919c · inbound

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models cites this paper.

TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:21:05.497735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:21:05.497735Z digest=sha256:1e9a57eaffbc96c9b02ee5251f1ff7884299fb2c6b58ef50ecc9c63e33088dc8

Observation e79e1b7e-d904-4d9e-bac3-d687bd900045 · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:47.827826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:47.827826Z digest=sha256:20d83b219089962d44ade919c91aedf7e8a22219d8794ef66c6de9b1dc220547

Observation 64cc6b8b-9e79-4452-8c8f-43cbbc82b203 · inbound

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames cites this paper.

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:18.562580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:18.562580Z digest=sha256:a88e71c86b90fc15c448578ab5e662d96aec9bbe0c335d9c0c7ba89f26cc95a7

Observation dec74c2f-daa6-4af3-993d-154d4e839773 · inbound

VUDG: A Dataset for Video Understanding Domain Generalization cites this paper.

VUDG: A Dataset for Video Understanding Domain Generalization Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:12.770007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:12.770007Z digest=sha256:18b99e1509e005217f176acd2a48687133f17b158dfda5dd4e7e3f8b27d3d71f

Observation 7015e766-957a-4caa-ae1a-f5ee80c334fc · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.887571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.887571Z digest=sha256:606452e071e9747ace4e6526ee345218d4dcf30d323ab11ab7cfc3b0291fe180

Observation 6521a2b5-eede-41bd-af83-a54a6c289419 · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.771494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:97fb04e04d6029fbf2f2342cbd26ce94aada9fcf2a56c3afecdafc1184a84d90

Observation 8a1f05c0-d91a-4c7a-830c-077a1a9b263c · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.591063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.591063Z digest=sha256:cf060ac250c428ad031154dd7e7f050614de3b1359dadb95a788f96f1d559412

Observation ad1ab0a5-8d56-42b7-abc7-9ab6202c0e16 · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.944763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.944763Z digest=sha256:877b3eb32497079e766ef98fd6d8a78ecc7ab092f2a6551e3aa159e630d3a40e

Observation 7d870a68-30b1-48c1-adc5-d6dd3459f69a · inbound

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding cites this paper.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.696016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.696016Z digest=sha256:8b4ec9a6a636e158fb6117323d9520c59c8b100408f0e9573499c2fb152f6a9e

Observation 5ed3b0e7-ba83-42af-a584-0aec14638407 · inbound

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step cites this paper.

CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.419558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.419558Z digest=sha256:b1d86de3d1ed68a1f55ac5f20ae418f6973c2cfafa6afeb378ffc5b28dca01cc

Observation 6fe0b637-3e4b-400a-9be1-0b50d4404bfb · inbound

VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations cites this paper.

VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:27.002204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:27.002204Z digest=sha256:a96083caa173f603784d5777cb47e0a31e203af722e2283c47b97685681bd2db

Observation e3afeb3f-c265-4265-b6d6-c827bc88199f · inbound

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? cites this paper.

GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:58:27.742728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:58:27.742728Z digest=sha256:9cf458c51e6e38c70b8c96d0b4173adaa448757adc4d345ea30eb25cee82e05c

Observation ee9acb31-5a21-4045-b2e0-84dd3f2f04f2 · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:42.629926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:42.629926Z digest=sha256:b6c82eff3cf9eec37f380f87b7dec05f2340b38cefc3471d321e83e364083579

Observation 67684500-24fc-4457-a575-c13145b72f96 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:37.260203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:37.260203Z digest=sha256:e3bd81413ce3ff891ddaa2ed092cf256fb9486e7f72fb769a9f1dceabefc1fa2

Observation bfe8e4e3-9465-4421-9f4b-1e4d1c9e580b · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.297857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.297857Z digest=sha256:e39a166580a0f67111055652255510a10bd6e5f38085490688ed6a89e0b6bab2

Observation 21a2c4cd-3046-4363-94bf-87e42334fa85 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.736581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.736581Z digest=sha256:aeedb1d224024ad0132a49722bca696186a345558bc4268021585fa6b57782ae

Observation f481b5b8-2317-4325-bedb-abeca7dfaecd · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.840834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.840834Z digest=sha256:313aa1f6ea232aeaa87f5964c2065ba0789325d6124044feaea52679fd8b3bb3

Observation ab2acae2-096e-49db-a432-7471022ac7b0 · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:19.068356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:19.068356Z digest=sha256:e36f182276df76e1f332d011fcb88fecf7b97b2f46e0db6171875d49177871f7

Observation f8fb1a83-acef-479f-a08f-2d2b8d26c3f7 · inbound

Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models cites this paper.

Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:00:39.537569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T01:59:23.928725Z digest=sha256:ed8c5b4e6634004fbedaabdc8634fbc063d2012adb04b9b724540d373d9c76a1

Observation fd65e0bd-1307-4f2b-8dca-5361e2f58090 · inbound

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? cites this paper.

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.548699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T21:46:43.353305Z digest=sha256:ca29b7a5956db5b75b9f87922a7063d3309a18d00ee11e05ee0fd272015e37b9

Observation 472bbbbf-c89c-4212-a656-713ea518a437 · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:51.900773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:877142080353e8362c9dc765809388374f567226ae7aae3e742217ec0dd4c8d8

Observation b65ff505-7262-4668-8e46-ac9ac36a94b7 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.688660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:ae07a4db41f91ded64433fd57d03dc2d916437e09da0dd94898fbd0b799f1376

Observation a5d7c21a-0f95-471f-8249-cbd7ba60f8a8 · inbound

Can Multimodal Large Language Models Truly Understand Small Objects? cites this paper.

Can Multimodal Large Language Models Truly Understand Small Objects? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:01:19.368825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T12:49:53.645987Z digest=sha256:5a9cc76220e34ad3bae1e0e568c1a064a5e25949ca723b8b82b0228bc5957b55

Observation e7457510-7bcc-4a34-abf6-c7bef05eca4e · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.414269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:a5559245fa9ab41e4c83cb9368ebc4cd4206770e6fe2887680bf5774fd92ad63

Observation 54501be4-b06e-4906-b391-dc0aef2ece7b · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 183

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.656222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:8698362a66747989295638bf62aeba602af5c20a91b5b5493f24ec7cbcd546f0