Pith. sign in

Paper Citation Record · LEDGER

Vript: A Video Is Worth Thousands of Words

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.06040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06040 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:10.639605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T04:02:43.573437Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9138f91f-6504-4749-a4bc-281226eec477 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Vript: A Video Is Worth Thousands of Words

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.209131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:b2325cd263bbeaf347d2390910da350856d6cd394751628400d55b5d700d9371

Observation 717bf914-3f9a-417d-8baa-b96a6588e2d0 · inbound

Open-Sora: Democratizing Efficient Video Production for All cites this paper.

Open-Sora: Democratizing Efficient Video Production for All Vript: A Video Is Worth Thousands of Words

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:01:51.831582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T12:01:51.366667Z digest=sha256:f49d72a61cfdf0b322bcce8bfde54085e3a60d293a9ac067d9688431314ea321

Observation 5cfa7a2c-3401-4e26-b69e-c5672311255c · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Vript: A Video Is Worth Thousands of Words

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.578826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:b860d4ae03fc21e823066c132e962697a1fe086e1f7aa94a0c7f5f3c523b69b9

Observation 50d8304b-f818-4452-a6c0-04aa645e7471 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling Vript: A Video Is Worth Thousands of Words

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.810354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:2375d8230287b981aa5f2d7ed997ec2a92d689156c50ed48bf52d35faa08cc7e

Observation 494bda89-ca82-47cf-91b1-46a129d9f1df · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Vript: A Video Is Worth Thousands of Words

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.787095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:de2a818b588ebc609d9aedf9c9bb401d74d7ad23a749b58468c9d1b63c7e5153

Observation 049d0006-7368-460b-a253-3cfe11c2066f · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Vript: A Video Is Worth Thousands of Words

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:10.639605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:10.639605Z digest=sha256:2b14a9590141048ed2fb3274f2ba7481d062f19f208830d660f95ecfb9f778d5

Observation ae47c4a9-105c-400c-bcc5-1530b30d7e8f · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset Vript: A Video Is Worth Thousands of Words

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:56.117027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:56.117027Z digest=sha256:e140ad370bdd58fdfa42fed109accf670492cd46abf1087d656d29bb55021f98

Observation aec39a04-ce61-4d56-92ae-d28ea6287923 · inbound

Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis cites this paper.

Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis Vript: A Video Is Worth Thousands of Words

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:37.099558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:37.099558Z digest=sha256:230c2c3d6aeb38ee65120bb2f8cb31a625d0ab5ecccfde01a864adf4bcabf18a

Observation a1cf3dde-b5ac-47fb-af1c-7cd6dcf65ebe · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Vript: A Video Is Worth Thousands of Words

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.975944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.975944Z digest=sha256:d37a5a45f003854f40c3c5c2ac0cfe2de106ad11d06f338775fb7653f2d21901

Observation 696b1e40-7be9-4046-870a-3400740071da · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vript: A Video Is Worth Thousands of Words

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:42.744544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:42.744544Z digest=sha256:5bdc3d502d0ea3c802f44d030986a2d41dc576e3d4681eff7227295ce6e0c475