Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2311.16103.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:07:09.556330Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:39:37.654958Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6a496b85-6cc0-457c-81f2-327b2ba5c964 · inbound
A Survey on Multimodal Large Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eeefd3dd-f688-43d6-80e2-93659ab49552 · inbound
TempCompass: Do Video LLMs Really Understand Videos? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0fe0559f-075c-4d14-a087-1de07c63ef1d · inbound
MLVU: Benchmarking Multi-task Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4380602a-a411-4aec-a9ef-e9139b5d7aa6 · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e5477cd8-9c92-4497-8925-47646329c903 · inbound
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 081f7dd8-e3a6-4ac6-a1ad-3459c8ad70bb · inbound
On the Consistency of Video Large Language Models in Temporal Comprehension Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f5f241-f1d3-4427-b93b-433224b632db · inbound
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e2b48806-8901-43d2-9c9c-310312ee36ff · inbound
Neptune: The Long Orbit to Benchmarking Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation facc0219-2a24-4ea2-a415-bd5a42daf1f1 · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d49752f-6432-41a9-b1a4-dcbdc4b942ad · inbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d1c639-d981-4237-b57d-e4f8c904d5b8 · inbound
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 65f5c485-f92e-4fb4-a78c-7f68d539b65a · inbound
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 790bf743-95bc-4de3-be39-e8ff36a81b80 · inbound
SCBench: A Sports Commentary Benchmark for Video LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbd32194-8fee-41c4-9768-d7b7d9e0d75e · inbound
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d26d89f2-222c-48f3-b906-17ba2aab3661 · inbound
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06a4d2c-520e-4fbb-aff2-933b4f80fb63 · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 832a929d-004e-4f4f-9261-e952cac7cfb0 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c9c0aa41-0ace-4d52-82ef-682ec8b0919c · inbound
TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e79e1b7e-d904-4d9e-bac3-d687bd900045 · inbound
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64cc6b8b-9e79-4452-8c8f-43cbbc82b203 · inbound
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec74c2f-daa6-4af3-993d-154d4e839773 · inbound
VUDG: A Dataset for Video Understanding Domain Generalization Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7015e766-957a-4caa-ae1a-f5ee80c334fc · inbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6521a2b5-eede-41bd-af83-a54a6c289419 · inbound
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a1f05c0-d91a-4c7a-830c-077a1a9b263c · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1ab0a5-8d56-42b7-abc7-9ab6202c0e16 · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d870a68-30b1-48c1-adc5-d6dd3459f69a · inbound
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed3b0e7-ba83-42af-a584-0aec14638407 · inbound
CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe0b637-3e4b-400a-9be1-0b50d4404bfb · inbound
VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3afeb3f-c265-4265-b6d6-c827bc88199f · inbound
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9acb31-5a21-4045-b2e0-84dd3f2f04f2 · inbound
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67684500-24fc-4457-a575-c13145b72f96 · inbound
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfe8e4e3-9465-4421-9f4b-1e4d1c9e580b · inbound
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a2c4cd-3046-4363-94bf-87e42334fa85 · inbound
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f481b5b8-2317-4325-bedb-abeca7dfaecd · inbound
AdsQA: Towards Advertisement Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab2acae2-096e-49db-a432-7471022ac7b0 · inbound
NeMo: Needle in a Montage for Video-Language Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8fb1a83-acef-479f-a08f-2d2b8d26c3f7 · inbound
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd65e0bd-1307-4f2b-8dca-5361e2f58090 · inbound
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 472bbbbf-c89c-4212-a656-713ea518a437 · inbound
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b65ff505-7262-4668-8e46-ac9ac36a94b7 · inbound
Evolution of Video Generative Foundations Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5d7c21a-0f95-471f-8249-cbd7ba60f8a8 · inbound
Can Multimodal Large Language Models Truly Understand Small Objects? Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e7457510-7bcc-4a34-abf6-c7bef05eca4e · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 54501be4-b06e-4906-b391-dc0aef2ece7b · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.