Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:07.215828Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 3 inbound Pith citation observations for arXiv:2505.23484.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:07.215828Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:20.756229Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T13:46:04.546310Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 03050e34-d999-4475-85d2-d04b690f7b8f · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b8f65d-e99f-4a01-bf41-326df0955172 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation St-llm: Large language models are effective temporal learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c67106-19dc-4238-9ca3-0e6c841ed35d · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation CogVLM2: Visual Language Models for Image and Video Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7356ba04-1818-41c7-a294-66f14449f1a2 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10b817f3-d4e9-4cb8-8ab9-e3a578ea3b04 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video-language alignment pre-training via spatio-temporal graph transformer
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7e7399e0-866e-4842-a160-4c0336e41f05 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c8e754-b3e5-4bfd-a3a7-ed787af19b41 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video generation models as world simulators, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 108ea320-d8ba-4fa5-945f-d231ff320481 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d87c70b1-fb7d-4622-82db-1c8158a86890 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoTetris: Towards Compositional Text-to-Video Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be41c9c-1b62-4753-a3fc-9a370c66d0f2 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604cc1a4-c2bd-4554-9788-ffd5f7ee5bf6 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c089519-5989-46a6-8049-06dae29c2cec · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87dc9914-32eb-456c-8fe1-f132e003f63a · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca8f7df3-eee6-4c56-a935-4718d59fa739 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LLaVA-OneVision: Easy Visual Task Transfer
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18be2f0e-2dec-4930-8261-00781e08b1b3 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7514bf63-6fd6-4d1f-91bb-5fac0532564c · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca62ed7-be29-483d-90e8-69510c4c986d · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LVBench: An Extreme Long Video Understanding Benchmark
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e391e85d-caf0-4be4-ba9d-4bfa38b63f02 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92774535-73c5-46e7-b9dd-0888bd1107d0 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation MLVU: Benchmarking Multi-task Long Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e2bce3-dc08-4307-8092-8bb49921bed9 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d0d03a3-e88d-47fd-89c1-0b0c68b212a7 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 27c90ae5-9b04-4307-97ac-d0e275bd7e07 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Bleu: a method for automatic evaluation of machine translation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf5d8e2-a9ab-4460-b1b3-ecdebf2b2111 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Spice: Semantic propositional image caption evaluation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8eae858b-f127-4ec1-8692-fecdf326cbc3 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Cider: Consensus-based image description evaluation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e41919a3-e421-4640-9ab7-bfc5195ba06b · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f9a44b7-a0a1-4be7-b1ce-28d7858af71c · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a60773-3653-4a53-b54e-5e1419e2f3bc · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation TIGEr: Text-to-Image Grounding for Image Caption Evaluation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e99203-a36b-4241-83d1-e51d4b59a629 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Faier: Fidelity and adequacy ensured image caption evaluation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 80f9082b-c0bc-438b-87d0-f70cf06cbf52 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation QACE: Asking Questions to Evaluate an Image Caption
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94835fee-ad4d-4552-bfd3-c7a69e491636 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca0c2cc0-a29f-4877-8745-c2da3f8cebe0 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715e4c1a-801f-4b0a-ba64-1693121f496e · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6aab97-6f55-4ebe-b869-2ad17a830399 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Learning transferable visual models from natural language supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0e6ff2-0a63-45fa-9211-af5ce307f4b5 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e981fc05-c3e7-4b6e-9485-436e038bc44c · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Ego4d: Around the world in 3,000 hours of egocentric video
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cce918-10a6-4013-896b-4d5db85bf45d · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Bdd100k: A diverse driving dataset for heterogeneous mul- titask learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 519aae23-4062-4797-a4bb-98b37ebac3b9 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Sharegpt4video: Improving video understanding and gener- ation with better captions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 01bdbe61-9ce9-47c0-a549-2bbcf545fcdb · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c83f4534-ca87-450a-b9dd-c589c8e2b750 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95346c4f-8a5d-4ab4-8130-d501d820c0f4 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Finevideo
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 51bb15f3-bdcb-4d6f-8df0-9038d79f7a30 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a451c7-54ff-418e-b871-b1b8cb46351d · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Qwen2.5-VL Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fef6ac-c440-4e1f-a169-4b430eaf73a4 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video instruction tuning with synthetic data, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f94370-a633-46b5-aad4-a98a9265096e · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation NVILA: Efficient Frontier Visual Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90acecf-d0fb-429c-9fcf-a05ef4a08058 · outbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44abbe4-452e-4abe-804d-ad4bb0f2d803 · inbound
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f9c5f80-2681-4dc3-aed7-67647c00322f · inbound
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567f4ded-f209-43b9-981a-2df0fcec0af6 · inbound
Building a Precise Video Language with Human-AI Oversight VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.