Pith. sign in

Paper Citation Record · LEDGER

A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2305.03347.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.03347 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:47:18.006542Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1b30f322-f650-49d0-a240-162c68223b5d · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.077258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:63c071284388116ba221af7b83f52641ac0567ecad55ca34de819f2cea05e3ec

Observation 846a8b7a-08f2-441b-9d14-1c632b03e629 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:58.033290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:dbfa146e2e5f552b4b95197f6729629008afb5bcbbfa51805500c827f5cac44e

Observation ccbfb08d-01a7-4aaa-a73d-43af0a70c66c · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:18.006542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:18.006542Z digest=sha256:84a8458b9d8f7bd91e3f65a0feb7088fed0726cf93fdfa29767af57c2f2e17a0

Observation 84d842cb-7de8-4da8-85f6-acb4aa60efda · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:18.171697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:9a22d347061abecb2fc0ebb9b1e029b1838e4df6d334ca2183192027000fe237

Observation 6b3390e2-ee80-41d8-ab2a-5550fc490cd4 · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T15:22:19.502624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:a7e6dc04c1f576d6338bff8182445031cb12b298e42ba10e570e67d8e7a0612b