Pith. sign in

Paper Citation Record · LEDGER

Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2402.19479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19479 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:00:03.445687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:34:37.714710Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ca86c59-c43e-452d-80e4-ad8e62ff0f4b · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.716860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:3fdc19069efaa64fdb8aa26c1c63a76899c2660e1e52fd9fc6de3d1222445df7

Observation 24360d87-83b0-4eab-a520-3c852ae868df · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 178

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:32.788975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:71458a8fae2a15a3b5fe0b7188d9f20cb426f5c822aee3eb6962785d69732743

Observation 4aee5e9b-9b41-47ce-9798-4aa1d06a11e9 · inbound

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation cites this paper.

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T21:04:21.812059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:04:21.812059Z digest=sha256:836262859b746ca0749b07c9d652bbe68192ffd7940992a344454fe8c4af54aa

Observation 3c826ae8-d066-4145-a3ae-3be54bc84222 · inbound

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection cites this paper.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.512869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.512869Z digest=sha256:5a39ff849517415063aebfaf272b189271a4d8b8699d424736c779964af27a69

Observation ed936f29-a602-4f0f-b38b-b6275cdaa178 · inbound

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation cites this paper.

MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:53:11.579911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:53:11.579911Z digest=sha256:2fc0bf10515d393e49a25f2c5f2ebac8bb8adb39d6354e471022e9b70e872ad5

Observation 3f8fb43b-7709-4f52-b0a4-a0b2a3109792 · inbound

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks cites this paper.

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:08.757039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:08.757039Z digest=sha256:9073c2283c4b6680c90863240b8e038e7f242af204812481e506232540726b89

Observation 85085582-55ce-400c-b4df-279fb124da4a · inbound

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption cites this paper.

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:21.903106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:21.903106Z digest=sha256:cd60537c4a6eb512bcb63b992e137f6979b68158a4ff2e5ec945abd91944d604

Observation 6e768d2f-3099-4da4-bac7-1f6faba5f696 · inbound

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models cites this paper.

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:56.190388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:35:56.190388Z digest=sha256:6d7706e72252f6781f0ece91471c0f5ff325549918040b7b268e8d43ecddf130

Observation 1fe78a7c-2475-41bd-9421-a07ef0a8d439 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.814669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:4f7ac390dd53b1ed560b4b14c182948ce51bc8826901f6e503a37e58eb29dd17

Observation 8bac707c-32be-4615-8705-715a5b0d3b89 · inbound

On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices cites this paper.

On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T10:46:29.404444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:46:29.404444Z digest=sha256:ac948682270db40e16a18b5b526dbb81be04b5a6d1791accc1d7ddecedce1d31

Observation 3617f228-9ea2-488b-ac75-bc512ca23935 · inbound

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile cites this paper.

Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T16:36:00.558029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:36:00.558029Z digest=sha256:bb64f740490e5f65d05b193690c27128be3afcdaaa906ff2c14aff21d218fed2

Observation 9482bb42-850e-4de0-9b9e-9238f6446b3b · inbound

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension cites this paper.

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:00:03.445687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:00:03.445687Z digest=sha256:a2481ac570434920115bac381b92ce8218b655c7821e4bed1e752e46a7934015

Observation 9a4bf5ea-8a1e-4037-9f18-93bd8d8d0591 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.739467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.739467Z digest=sha256:6ca7c7645183fcd7791f774e8cdf2e7d849ced0c495c7db5cbde2c2ccd006d01

Observation 4ae97d87-06df-42d8-83df-1ca90836f6c6 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:24.668995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:24.668995Z digest=sha256:e9b7b3e1aeee6f88962de2551398af9bc3cebe9dda2722b962ed01d1ccb8dc36

Observation 8489af4d-554b-4286-a8a1-1344536ca1b2 · inbound

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning cites this paper.

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T22:18:56.115559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:18:56.115559Z digest=sha256:92c52773101f56dca364ca46e23a8003fc7e8ceeb4426f04fed2e7f54bcf36c2