Pith. sign in

Paper Citation Record · LEDGER

Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2305.13903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.13903 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:49:34.791982Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:21:00.904134Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4707c3c2-93dc-4bb6-a532-5c211b47e5bb · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.169118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:3138d2b89fe479e5a4341a7f2c00d9fd7b2e9dd3f309116fcdfb391224baa035

Observation d0be8ea1-5ec6-4bec-ae78-8ae0d3253522 · inbound

A Survey of Hallucination in Large Foundation Models cites this paper.

A Survey of Hallucination in Large Foundation Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:21:00.906448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T15:21:00.778049Z digest=sha256:bbb028e03a6a00872aa5a45a3d80c073ba7958832947171801dfcf5e97a6d728

Observation 6feb924b-87b2-4017-bae1-b50e6550acd2 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.132781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:1fa97e4eb6bd15de2c5eef88a4b9117b6427d025c06a37de30dbb00bcc411c7e

Observation 34ee85ce-7b5b-4e31-ad35-2200692b9b20 · inbound

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models cites this paper.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.791982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.791982Z digest=sha256:e8f3aa63608ca172e1f6ebf663bb66a9277797e9f64e6e657f0b0034cf02fcf7

Observation b440622a-30ba-44b6-94be-850c0c53cbba · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:57.194938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:57.194938Z digest=sha256:d8d6b5bc812a5ca6f3ee9f014696fca6ea62dd634a340f04cc3d613b0a06ba04