Pith. sign in

Paper Citation Record · LEDGER

VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2406.11303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11303 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:43.788047Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:26:57.184700Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df47b175-ed30-402b-9ad1-4104542abff3 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.342097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:4ae0596d146f21d132f2f9f02e9259e8d0adde55ce242767601e9f24e63e79bc

Observation 29aaa453-09ca-4625-84b3-d9457014c874 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.788047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.788047Z digest=sha256:3cee0a0d7ab97f77aa63d6eb98ae8e03e159fe491573636e28df3d7bd3d0a5ae

Observation fc66876d-c16b-4d4b-977f-7b397612597f · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.131871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:01.131871Z digest=sha256:b4ead16cadb939c26767bda71e8940621ea9f54bdb7919d531eb4ebe3bf2f977

Observation 7514bf63-6fd6-4d1f-91bb-5fac0532564c · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.738031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.738031Z digest=sha256:0be820351cad6635fe902b81a487673f926864352546d3b78954fde3a2a9cf94

Observation f4e62474-1e58-4261-b56c-11dbe3e6481f · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.866771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.866771Z digest=sha256:467cbe87d02b6cc6b989dbeed12a1553183c94b7e4e5d4ca02bed7ea9f8a58c9

Observation 6f5a593c-f7dc-4764-85f5-da4346e18514 · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.943565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.943565Z digest=sha256:2f3dbb6b2764546f0dc34558c0daa6a77847b0593a0b7e8080ac3578982ca56b

Observation 408e7f61-f69b-46cd-97ec-da0311486d57 · inbound

VUDG: A Dataset for Video Understanding Domain Generalization cites this paper.

VUDG: A Dataset for Video Understanding Domain Generalization VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.535033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.535033Z digest=sha256:269d73b5cb79f5aff30b714114ba68f6493797f67ce5c89baf05f9d5614b2d07

Observation f72b07e1-89ec-4e63-a23d-54c6a082e19c · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.739001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:4c0146633e12027ccbdbf36dccbeb2230d8cb7362b5b27ce35ddffe07b074a6a

Observation 72944cf2-74c2-4dd2-91fa-073fcf1672fa · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.901902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.901902Z digest=sha256:40717b5534310419db283e3c6ac84ef0bff3e382ef7978fdb9be9b9101bd1c25

Observation fd3247c8-5278-4a34-86de-a9e8723a234c · inbound

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models cites this paper.

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:51.290582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:51.290582Z digest=sha256:fd1bce31165cffae6bfe634ff7547a3efa91a9c10045c173120eb78ca4ddf664

Observation 7d166b95-e89a-4822-9a73-8309863d4e2a · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.159644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.159644Z digest=sha256:3a5dcff4c680904b003b261d4b8a4455848cda3aecdf2b2ae5874a74e76a9538

Observation 4880c9fe-dc2e-4103-adad-2d91bd29d674 · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:15.574763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:15.574763Z digest=sha256:43cadf40bfa2e74a246134040886d4cbe8571d159fc85afac96ec64fe437997b

Observation 4f18c017-fa11-49b1-9ff5-5a40e6d0cc64 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.406514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:ae7d1f64835b15005d627e895f1c35d75f84b610812992d81e9d01ee95102195

Observation c02b375f-471e-495f-ba2d-dd8d1b12bdb0 · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.051487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:9f8616d833c74181743a7949b2620327baf51edb03baa2ddbab969ddc575e563

Observation b7b3ffe2-8116-4a87-b31b-d0d2b8287e22 · inbound

VidMsg: A Benchmark for Implicit Message Inference in Short Videos cites this paper.

VidMsg: A Benchmark for Implicit Message Inference in Short Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.085934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:25:06.594946Z digest=sha256:52c0f28b4de4f2bb2b91ac0828c4f44f81d72f205d6562363da1f9ef7ff5a85c

Observation fa3a6b1c-d6de-40ef-bfcd-2f77df218b05 · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.186187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:b7022629017c07649d456e02256087c220d94830745b69a43aabbe6613f8b07f

Observation 7ad99eda-7280-48fb-a4c6-8582ee1d7c21 · inbound

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation cites this paper.

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.750946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T00:53:52.724902Z digest=sha256:43f05678583070c4062f5ff4461933597d17fbf7b8bb67587f62a18c6d9f401f

Observation e059e240-3680-4a48-8baa-dbd830e98b72 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.981358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.981358Z digest=sha256:4cc533b8e682d90593667f1ff964e6604f5967d7a290dc492311018b07d965ab