Pith. sign in

Paper Citation Record · LEDGER

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

As of 22 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 3 inbound Pith citation observations for arXiv:2505.23484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23484 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:07.215828Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:20.756229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T13:46:04.546310Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03050e34-d999-4475-85d2-d04b690f7b8f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.520983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.520983Z digest=sha256:ae7e22a6695ebcf3f8d5100aba3f2d6bcfaca97998f31bca7a0caa4d6c9b2217

Observation b0b8f65d-e99f-4a01-bf41-326df0955172 · outbound

This paper cites St-llm: Large language models are effective temporal learners.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation St-llm: Large language models are effective temporal learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.586430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.586430Z digest=sha256:fd7950a97bce34ca3bd90edd68e68e16bebdf648504a3ec416df019f241f03e2

Observation 91c67106-19dc-4238-9ca3-0e6c841ed35d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation CogVLM2: Visual Language Models for Image and Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.656403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.656403Z digest=sha256:f8dc53058b58558112719fdfd81e95a619532e91d008f6eaee7e835fc870920f

Observation 7356ba04-1818-41c7-a294-66f14449f1a2 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.748763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.748763Z digest=sha256:7aa8d9ab07f6877edfa2341712f5eb2dd09238d03f1a9cd032704f65acdead3e

Observation 10b817f3-d4e9-4cb8-8ab9-e3a578ea3b04 · outbound

This paper cites Video-language alignment pre-training via spatio-temporal graph transformer.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video-language alignment pre-training via spatio-temporal graph transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:09.474243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:02.819957Z digest=sha256:6b9729131dabe4fa20c9cb3d4c1842ac8dd4dfcc04a8cf091ef142c17cf6bc3a

Observation 7e7399e0-866e-4842-a160-4c0336e41f05 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.870434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.870434Z digest=sha256:0edebbaedeae11b879932821c57ce63a9d5d0bd18c68166969030e6368b227b4

Observation 48c8e754-b3e5-4bfd-a3a7-ed787af19b41 · outbound

This paper cites Video generation models as world simulators, 2024.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video generation models as world simulators, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:02.947034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:02.947034Z digest=sha256:dfb6d6d1f2ef1a97bfd190c489d5c23cec4b33bac6f9227a39556aad1c4ac466

Observation 108ea320-d8ba-4fa5-945f-d231ff320481 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.021385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.021385Z digest=sha256:9490be87fce3d918398b047d009049c8f6ed4989f66907d6e9fbe514aca590ac

Observation d87c70b1-fb7d-4622-82db-1c8158a86890 · outbound

This paper cites VideoTetris: Towards Compositional Text-to-Video Generation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoTetris: Towards Compositional Text-to-Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.123192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.123192Z digest=sha256:0d74ef88807091e1bb93956190bae37e31c3062dcddc343ae23edfb872dde7e0

Observation 4be41c9c-1b62-4753-a3fc-9a370c66d0f2 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.216736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.216736Z digest=sha256:bbfd7e829335c7454ebb609a581ca4ab1f55a533ec549e81b1904740f4e180cd

Observation 604cc1a4-c2bd-4554-9788-ffd5f7ee5bf6 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.349403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.349403Z digest=sha256:a616c6f08439daa9ca79bc60901a9a36cca21b6b0561dd217f477dd26c20eb3b

Observation 1c089519-5989-46a6-8049-06dae29c2cec · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.443005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.443005Z digest=sha256:61c3fe5aea6d1586d0c0f49dc21bed23bba5e5cfdffa61ddb19f56064d4416bd

Observation 87dc9914-32eb-456c-8fe1-f132e003f63a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.522560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.522560Z digest=sha256:45ef45e18230859b30367c18229c6777a50e387980c95d782786a788447d9607

Observation ca8f7df3-eee6-4c56-a935-4718d59fa739 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.630365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.630365Z digest=sha256:a555ba7d557269819296e499366555f45477faa32f0e0e4819a3e518bf572062

Observation 18be2f0e-2dec-4930-8261-00781e08b1b3 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.678714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.678714Z digest=sha256:d1fff245aa0f268327a13edd4b0b2d6ce5760afc3737fb70af30ccb0b4130fa5

Observation 7514bf63-6fd6-4d1f-91bb-5fac0532564c · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.738031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.738031Z digest=sha256:7b4d3e8e7cfaeca5ada9618df563ecbdea936522adf6194fccc7faaacfafc030

Observation 1ca62ed7-be29-483d-90e8-69510c4c986d · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LVBench: An Extreme Long Video Understanding Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.883798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.883798Z digest=sha256:3f410a461c697002d62fdab351fc0ae4ed835c6688ee117a37ea7fe3b1cb9892

Observation e391e85d-caf0-4be4-ba9d-4bfa38b63f02 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:04.018926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:04.018926Z digest=sha256:14b14d44341b77aadba5128ed407b8dc197b2dc66478352c0409e8c5a22f8e18

Observation 92774535-73c5-46e7-b9dd-0888bd1107d0 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation MLVU: Benchmarking Multi-task Long Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:04.125937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:04.125937Z digest=sha256:5db2f0cf601ef148fe820c508c2f57194aed4c286d4bad4f7376f473bae86237

Observation 91e2bce3-dc08-4307-8092-8bb49921bed9 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:04.238595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:04.238595Z digest=sha256:b1888fea317dca38da65c03cbd226e0be42e9fecffb77730061ec03c94a68962

Observation 1d0d03a3-e88d-47fd-89c1-0b0c68b212a7 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:09.283671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:04.401517Z digest=sha256:ab029ac8c76fcb811b6846af53a94368448db0d61bb2fabcd2889939be0d3c68

Observation 27c90ae5-9b04-4307-97ac-d0e275bd7e07 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Bleu: a method for automatic evaluation of machine translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:04.531335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:04.531335Z digest=sha256:4c843d58cfecfe05e3d7e1f2cd925ff721bdd269de09ec46b5896d8b72173889

Observation aaf5d8e2-a9ab-4460-b1b3-ecdebf2b2111 · outbound

This paper cites Spice: Semantic propositional image caption evaluation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Spice: Semantic propositional image caption evaluation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:09.162468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:04.657753Z digest=sha256:0750f0731bc29323abcf076d3b02f1797534c3e58376ede48e0bb52eac1e3c7c

Observation 8eae858b-f127-4ec1-8692-fecdf326cbc3 · outbound

This paper cites Cider: Consensus-based image description evaluation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Cider: Consensus-based image description evaluation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:08.993453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:04.812105Z digest=sha256:a1cf64064a43fd1379024791f2e420b4be731c3b662e74dab5f127e566e16e05

Observation e41919a3-e421-4640-9ab7-bfc5195ba06b · outbound

This paper cites InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:07.629130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:04.880205Z digest=sha256:aab63909feef57545f0ca162958c4b43dad3368cce2a50e206e3fb8968ea62b6

Observation 0f9a44b7-a0a1-4be7-b1ce-28d7858af71c · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.036979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.036979Z digest=sha256:75e3eec6662c229799f2ca46c9d4215c108b08f1877e3021cf0ec7e657c46255

Observation 51a60773-3653-4a53-b54e-5e1419e2f3bc · outbound

This paper cites TIGEr: Text-to-Image Grounding for Image Caption Evaluation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation TIGEr: Text-to-Image Grounding for Image Caption Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.100035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.100035Z digest=sha256:2d3592ca509dfa10ec4c4abc4911d747aaa314a3431c963f6eac910e376fabaa

Observation e4e99203-a36b-4241-83d1-e51d4b59a629 · outbound

This paper cites Faier: Fidelity and adequacy ensured image caption evaluation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Faier: Fidelity and adequacy ensured image caption evaluation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:08.852277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:05.225598Z digest=sha256:a86cc207a5dbd913a7c252abc02d91c6aeb41db09cef45e7629ec62039eb1ce1

Observation 80f9082b-c0bc-438b-87d0-f70cf06cbf52 · outbound

This paper cites QACE: Asking Questions to Evaluate an Image Caption.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation QACE: Asking Questions to Evaluate an Image Caption

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.340321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.340321Z digest=sha256:9372fa154c46ebc971944be684ffa1648f0a1e874561ca7a4fa0edf4da2f4351

Observation 94835fee-ad4d-4552-bfd3-c7a69e491636 · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.447336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.447336Z digest=sha256:04e9c20e94bff63ed0f7d4256b8e9dd581025c8db8dd8710e309a375bcaa79ef

Observation ca0c2cc0-a29f-4877-8745-c2da3f8cebe0 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.569130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.569130Z digest=sha256:fa960c484b30348e28bfc1d47e8b57788f9780f011dc3b0bcad6fcf78840e820

Observation 715e4c1a-801f-4b0a-ba64-1693121f496e · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.725545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.725545Z digest=sha256:bd85cc67a9b2bcc40be558aa4f93f9b69337e88700397830653e5bad2a680760

Observation 1e6aab97-6f55-4ebe-b869-2ad17a830399 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:05.832538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:05.832538Z digest=sha256:12953a8096f7ac255b51b241ecd0472ff965ed1813b36c2006ff5a88609f2898

Observation be0e6ff2-0a63-45fa-9211-af5ce307f4b5 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:08.681194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:05.912097Z digest=sha256:a1feeb31b9a038d3cbc796060d4973501f4fe7bce80d2c8ebf47c780b60fd9b4

Observation e981fc05-c3e7-4b6e-9485-436e038bc44c · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Ego4d: Around the world in 3,000 hours of egocentric video

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.012064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.012064Z digest=sha256:27d9ff5d4926eca5ac6588d5827ae0fa8444824ff6a0a7e3774c9ebac9bf6d98

Observation 30cce918-10a6-4013-896b-4d5db85bf45d · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous mul- titask learning.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Bdd100k: A diverse driving dataset for heterogeneous mul- titask learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:08.514759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:06.141445Z digest=sha256:e673876162a8a98d7028ee5976fb648557a4038384d77ed3271f647d63fdc6fc

Observation 519aae23-4062-4797-a4bb-98b37ebac3b9 · outbound

This paper cites Sharegpt4video: Improving video understanding and gener- ation with better captions.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Sharegpt4video: Improving video understanding and gener- ation with better captions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:08.330549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:06.246206Z digest=sha256:5c1197182853e30b1b2ca9e97ca3885e5df6a5278b4167e788c1a2ea0d58a560

Observation 01bdbe61-9ce9-47c0-a549-2bbcf545fcdb · outbound

This paper cites VidGen-1M: A Large-Scale Dataset for Text-to-video Generation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.335779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.335779Z digest=sha256:2b21988310337bd51de1e7f0686f8ff4371b71f6a51183702e6096456a9ed5fc

Observation c83f4534-ca87-450a-b9dd-c589c8e2b750 · outbound

This paper cites ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.426328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.426328Z digest=sha256:78f3c2bc11d66d30dac7ac0ca7a1acaf67cc9de3a4bb33a2ff1d28632b59e00d

Observation 95346c4f-8a5d-4ab4-8130-d501d820c0f4 · outbound

This paper cites Finevideo.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Finevideo

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:08.165287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:48:06.573272Z digest=sha256:da433d5d5ce9d5e3d80852e49d8f4efaee9035c4e83d38b0ce10ea3828825a46

Observation 51bb15f3-bdcb-4d6f-8df0-9038d79f7a30 · outbound

This paper cites LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.660518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.660518Z digest=sha256:21f198f3eba9d952738619df982e178821e843bffd130e9f5287a1c5ab1dc684

Observation c4a451c7-54ff-418e-b871-b1b8cb46351d · outbound

This paper cites Qwen2.5-VL Technical Report.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Qwen2.5-VL Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.771636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.771636Z digest=sha256:9f95bd544bcce3e0ad8be47620e263419a8be1eef5fe5671c9f2a65ddcd053e9

Observation 25fef6ac-c440-4e1f-a169-4b430eaf73a4 · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation Video instruction tuning with synthetic data, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:06.879949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:06.879949Z digest=sha256:33e810d8ecb7d01f362585210d7852bf2c0b0b4fbf39e3642450779c835385b5

Observation b8f94370-a633-46b5-aad4-a98a9265096e · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation NVILA: Efficient Frontier Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:07.044736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:07.044736Z digest=sha256:2eb83d73ace5d01fd873972a2dc1ed83d1fa89a956b094a38e9cfba8342f5bde

Observation c90acecf-d0fb-429c-9fcf-a05ef4a08058 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 45

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:48:07.215828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:07.215828Z digest=sha256:fd98f08b35227e30f395bf910cd8c69cea854f11e16bae76fa87221069c28cbe

Pith citing papers

Observation b44abbe4-452e-4abe-804d-ad4bb0f2d803 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.756229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.756229Z digest=sha256:e066806693c4fa53fff86d41f8ba8f3232c57cd483b7866f7c0aca6f5bf660aa

Observation 8f9c5f80-2681-4dc3-aed7-67647c00322f · inbound

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models cites this paper.

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:13.314128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:13.314128Z digest=sha256:a81eca1f1cec509a9d397c2a35e3f94902af16e3ef57d63bac6c04aee03e4cbd

Observation 567f4ded-f209-43b9-981a-2df0fcec0af6 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.547812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:6b5064a89be3d329f07c3e18d88669b907d68b768fe622c7c59673665f7a155f