Pith. sign in

Paper Citation Record · LEDGER

Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2402.19479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19479 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:36:00.558029Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:34:37.714710Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ca86c59-c43e-452d-80e4-ad8e62ff0f4b · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.716860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:13eb56d749923476ffc1355aa7c052b646e3267e2b4e20bcdc159ba1f1d00d79

Observation 24360d87-83b0-4eab-a520-3c852ae868df · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 178

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.788975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:e632d94910fd2bf79b1ae63580fa870b4188188b652d933f65c02db7cc70db94

Observation 1fe78a7c-2475-41bd-9421-a07ef0a8d439 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.814669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:944dca8dcd35b1327882d7139e28a34c3ac0470b117b0ea21cd6c517c1d7caf4

Observation 3617f228-9ea2-488b-ac75-bc512ca23935 · inbound

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile cites this paper.

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T16:36:00.558029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:36:00.558029Z digest=sha256:e5eb0e5df5e2fa148b0e349aaeafbfefb137a4ea6058a03cfe0c7ba4760cd731

Observation 9a4bf5ea-8a1e-4037-9f18-93bd8d8d0591 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.739467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.739467Z digest=sha256:6a9bc77e5efbb940e35ca6ceba776291885ae89c2bbc21640d303c52c52370d9

Observation 4ae97d87-06df-42d8-83df-1ca90836f6c6 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.668995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.668995Z digest=sha256:739a2a013802257b51fb585dac2aa03228f69bed1c23dcfc3a157246f697d425

Observation 8489af4d-554b-4286-a8a1-1344536ca1b2 · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:105c4686e37c9ac783058021f7d8cea21bdcb85a96edb8f69c373ef77a1f16a3