Pith. sign in

Paper Citation Record · LEDGER

A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2305.03347.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.03347 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:47:18.006542Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1b30f322-f650-49d0-a240-162c68223b5d · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.077258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:af0604d8c8a06ca6726c14aa1f1e7204dc2376cee75229f3b9cf0ac7b5de7010

Observation 846a8b7a-08f2-441b-9d14-1c632b03e629 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:58.033290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:f9cdf6739f60a68d2e7db2792ca51e811dbb8dc02fdc541ecad8cba7940e41b8

Observation ccbfb08d-01a7-4aaa-a73d-43af0a70c66c · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:18.006542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:18.006542Z digest=sha256:cfcb25b56efac5f27d891fbee9837c0ff31f7a7548d8f31505b37a526eebbd37

Observation 84d842cb-7de8-4da8-85f6-acb4aa60efda · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:18.171697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:4180bb469184f4c2d2facb7395fbc84e60eac1d6f09d00b44f3c71b8f4723716

Observation 6b3390e2-ee80-41d8-ab2a-5550fc490cd4 · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T15:22:19.502624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:93572b187dddad286e254407e864ba201b5561d9c993e772cebee78f872c5d3f