Pith. sign in

Paper Citation Record · LEDGER

VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2109.14084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.14084 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:07:31.643197Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.380967Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe34e4e4-4f16-43b4-b379-342a81ff67a2 · inbound

R3M: A Universal Visual Representation for Robot Manipulation cites this paper.

R3M: A Universal Visual Representation for Robot Manipulation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:26:53.979248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T13:26:53.843613Z digest=sha256:3936323e96dbf71e97aeb6b46bc875c8a572d52d27f3c1ebadc70536e5033d64

Observation 72761ef6-7d17-4890-8c76-6cabcc30f73d · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.608579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:721d7f8e3a46428ce5df03f26bd9a531b7c0bbc67e56182d694b40d318ba4d2c

Observation 1eaf6c1f-7f3f-4dbd-8fe1-793a40eb840d · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.372393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:43171c38731dde0087122811969c69d80157d51af4cad7478ee33b4583221ec2

Observation b9b1d4cd-aa91-48eb-b583-b569e1e5952e · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.675840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:648fc95388711b1f2631add66f94da7eb9f0f237b1099d46eef7b6c176fc2377

Observation eed293b8-bc7f-4447-a1a7-672174934fd9 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.554543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:0ab5c2057f2ab68c0c567c3d2567bc289a8403a6e49c2918e215faf6107dfb79

Observation f4791b68-3222-428b-9824-0255dc411a5c · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.117146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:162fe9e860e925835eba4ab9a209625a6aa7614fcb798c76bd77c4935edef797

Observation 9d4cba46-2e95-48a7-8e42-e8a936162640 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 292

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:24.082290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:19fe201fbb5367926cbd93ca7ce956413c3d729de7dbef756e0879c2c6416821

Observation 56f2d5c3-e0e1-47d8-a3ce-3dc8cb32a4d8 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.250901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:377ba64426ac4e5a93b2a04b375e38f80a0fd8bb0a8175acb26854254ffdfaa9

Observation 61529d6c-d42b-45a7-b6ba-b13560387e97 · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:10:15.136859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:cc4ec23663ed43cac7ff6edcada2dc610b3ba5abdc42bd575b7cf4576db910db

Observation 4b2d807e-fe8f-405d-9432-e9dd4bc960fb · inbound

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation cites this paper.

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:31.643197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:31.643197Z digest=sha256:7ab96f322357660d287da3cc03358bc8e51eb7461551cfe7212d7a82879ac547

Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · inbound

Implicit Counterfactual Learning for Audio-Visual Segmentation cites this paper.

Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.193705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.193705Z digest=sha256:70da2cf0f87433fce341599466e0438e4ef29f7bb3a2013d8f701ecfbfc314b1

Observation 90e63cc7-bf95-4ea5-aa75-fa1289e5b836 · inbound

Group Relative Augmentation for Data Efficient Action Detection cites this paper.

Group Relative Augmentation for Data Efficient Action Detection VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:40.226378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:40.226378Z digest=sha256:fe79893dd719b4eaa290cac748b8e8b2edf616e2c118d0e4d6f6127892d312ee

Observation 2335abde-285a-4cac-8447-b1a2c9cbfa25 · inbound

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models cites this paper.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.517124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.517124Z digest=sha256:d7fe6719f23a633269013a68f1bb3b8789b45282187c2676b5391fcc5b389cb2

Observation 330c954d-ded8-446f-a47b-22ddfd931cff · inbound

Adversarial Video Promotion Against Text-to-Video Retrieval cites this paper.

Adversarial Video Promotion Against Text-to-Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.166881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T00:05:07.182361Z digest=sha256:23491e5b4b4fdd7a953034c83bb027e859b5239fae29701c81b5bc303f145c33

Observation d6b50866-ab14-445b-a0f7-1b0e75e07e7a · inbound

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications cites this paper.

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:58.248635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:58.248635Z digest=sha256:ac0d9b984201467ba110adef85357ae6a6821c2897943c102d620e24f53834b0

Observation 85192c1e-362f-4e51-b9a8-049c42e00291 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.187507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.187507Z digest=sha256:d94ec914457e09268714fb4bad0b238f9279ce3bbe9204e0021855acc025ab1d

Observation 9761cf3d-85dc-4d2b-94c8-8d97a3f83eff · inbound

Calibrated Multimodal Representation Learning with Missing Modalities cites this paper.

Calibrated Multimodal Representation Learning with Missing Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:10:22.692713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T22:08:07.217659Z digest=sha256:79cdc841f0d08e44922b2990930452a258ebec9da69a3dfd7253312cb146d29a

Observation a30d6193-0b77-4a33-8a28-fdcf02fa13f9 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.335434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:17180dd1f74966cc5b694b30567255c30c0367a78670275257bde61036ddf69c

Observation 51e029eb-ce1e-4431-9dbe-a30d6a95526c · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.921271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:cdfef71811b8c67c4ba96bebb634b71352eb1d29cfbc56dd011973c0a1e15293

Observation 9d82ce30-3dc5-4cb9-8455-d148e1611781 · inbound

Learning ORDER-Aware Multimodal Representations for Composite Materials Design cites this paper.

Learning ORDER-Aware Multimodal Representations for Composite Materials Design VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:40:14.550097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:36:44.656655Z digest=sha256:84193d1cc91a99603e23b9a11d7cd69e07713e83bcc62bae0daf439a60cf8c9a

Observation 98a3f7aa-d2a1-49c0-a7f6-9169612aa447 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.570008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:44c9e0ad844385639281e7311261896a136e95dbfb9f63b2742c289c562c1af6

Observation e5989aae-07c8-4169-a9bf-75f9fcec2bff · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:04b856800fdf9bb4be0613dc2ee273b85da01662795959a3190d6013c2230da4

Observation 79bc04fd-414f-4150-a5c8-710534ff86cb · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:b881d942c65cf59ce114dcecf317b09e9d3538dc352f1cdd4545102ae19cfd48

Observation c3f96828-e72e-4f38-953b-6c85b0fefb02 · inbound

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing cites this paper.

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:53.943204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:44:52.768021Z digest=sha256:8734e615004ff29add4e825f9fcb9597f7596f8df8f5cbec17cd3b6e5d19570c

Observation a286e855-f8a7-4198-b58a-1c9456c2bb14 · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.326399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:2849d6c111916edcc3837fd3a940a407868b09f6d12daacbf19d4f3f6fa2cdcf

Observation bf843ac4-63bf-4441-91ca-748db06036a3 · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:21.931008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:07:21.203513Z digest=sha256:b51e1f5a3c73c13eb6fa9d5a88f79175c26e89a0b43bf9ada5ca35f1eafcd550

Observation ee162317-38b2-4cf0-8ca7-b03ee279b6d7 · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T20:05:20.702092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:05:20.702092Z digest=sha256:b465b76f03d2a07eb952b530fa2b8f287cc7db25e8eb8728ad3daf56650f3bcb

Observation e05b8581-ffea-4cf1-8fd9-ae8478f83ced · inbound

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation cites this paper.

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.280521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:04:39.623342Z digest=sha256:f2f461036719f4128eee1edf81b026d2fbd923e3f3f2687fbb201a9fe7947fc1

Observation 6d4afa95-e025-4962-a74a-fc809fa40d2b · inbound

Multimodal LLMs under Pairwise Modalities cites this paper.

Multimodal LLMs under Pairwise Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:39:40.686881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T05:37:56.792564Z digest=sha256:2a655a3d90de563c96a174cc98047f5ecbe29c86d69a8594d2acb73994b8d9e5

Observation b15f4643-7da8-4008-bbbf-485042c8b020 · inbound

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset cites this paper.

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:16.925691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T15:29:51.424572Z digest=sha256:29bec56238f9a488f8532018d2295a9c135a0fe187f0e906c95adf98e73166e1

Observation 8acc79c7-1e61-4a3c-85fb-d94956ca9fc6 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.382403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:d2d4a371a8150329efe0c85011e97d847f5b59b4ab1c424aef15b41df48fcddf