Pith. sign in

Paper Citation Record · LEDGER

VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2109.14084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.14084 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:35:56.167427Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.380967Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe34e4e4-4f16-43b4-b379-342a81ff67a2 · inbound

R3M: A Universal Visual Representation for Robot Manipulation cites this paper.

R3M: A Universal Visual Representation for Robot Manipulation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:26:53.979248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T13:26:53.843613Z digest=sha256:a5e399e9ce96274ee543026df422eb42c2b819a184a798ed4485dd3675614bb8

Observation 72761ef6-7d17-4890-8c76-6cabcc30f73d · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.608579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:0ba8976ddd2de0d4066a1154539b36c9de8e118fbc95afb98bae4d220881fd60

Observation 1eaf6c1f-7f3f-4dbd-8fe1-793a40eb840d · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.372393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:0c96c828324403e5fcd0f85d091f2cf6173803d30fe9be0ed35e855a5dd760e1

Observation b9b1d4cd-aa91-48eb-b583-b569e1e5952e · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.675840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:faf7524ad0da9196b16613118436e160c3a8887d58279ae65a281d7c32737bfe

Observation eed293b8-bc7f-4447-a1a7-672174934fd9 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.554543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:fefa436876f4ed5434ac38d5ed8ba994d996885c1169b9277947e14edd57f82b

Observation f4791b68-3222-428b-9824-0255dc411a5c · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 177

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:27:59.117146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:e0f3f7bebbea12de5bd80327b83edb765357f3c717e4f4809b1b7189025e868f

Observation 9d4cba46-2e95-48a7-8e42-e8a936162640 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 292

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:40:24.082290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:e55acfeb7dce4f5196ac9274c3b65d3859c00f2febaed1d41b9596306b65338f

Observation c26ebaa9-647b-4a15-810e-50452acb4fc8 · inbound

AstroM$^3$: A self-supervised multimodal model for astronomy cites this paper.

AstroM$^3$: A self-supervised multimodal model for astronomy VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:23:18.872569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:23:18.872569Z digest=sha256:3d6058f7e13b1c6ed2cb617664020addcff458d011ddfd5efba59a1659b02e41

Observation 31beffb7-5a7b-4c4d-b56c-48ca0150bb0e · inbound

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos cites this paper.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.293896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.293896Z digest=sha256:c0e72abb9e35d3bd1f28b73ea943b989108dede32c214e246fb60916462d15a5

Observation dff4cb1b-ef3f-4c37-bf19-14488e7ebbf4 · inbound

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning cites this paper.

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 165

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:00.887668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:00.887668Z digest=sha256:102273baeeff0568206733bc4483fca8d27c5edf28b4e3935679de41869b8f7e

Observation 7a3a0449-81e3-42c2-b0c6-761ff8b8693d · inbound

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training cites this paper.

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T05:27:14.887650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:27:14.887650Z digest=sha256:25a28e0484e4d1008d68cc4953c51a328474e874a05fe9d3797e54297ea9e79f

Observation 1cefe960-c1c1-43bf-81ce-9009b494db8c · inbound

FIction: 4D Future Interaction Prediction from Video cites this paper.

FIction: 4D Future Interaction Prediction from Video VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-12T04:55:59.918193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:55:59.918193Z digest=sha256:369ca1d639b140357850eb621cdf8c467739bcac6834b3b36e85258b8d9c25a2

Observation af4a49fd-e24f-4298-a8e9-0d1986020330 · inbound

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning cites this paper.

GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T18:40:22.188066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:40:22.188066Z digest=sha256:229a2fcc02d74ac2284be03211f8052c8add9dc7070cbe31d2e330c3d0f2fc4c

Observation d8ac9601-57a2-44d3-a7d9-3c12dc0ab831 · inbound

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey cites this paper.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.909659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.909659Z digest=sha256:0e7a048dd9cfac6ad747137fc1a15f0f3b47ce04f5348512be8cf1a4b16d069f

Observation 8e7b85d8-1c8e-4ac3-bc9b-b781c8f8a429 · inbound

Automating the Search for Artificial Life with Foundation Models cites this paper.

Automating the Search for Artificial Life with Foundation Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:11:36.126915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:11:36.126915Z digest=sha256:2e14ea4fa0841d0bc5a536c50d298ae6f71e9df4b1a595a055b9f27ad166819c

Observation c638963a-d613-4833-8a32-50babbfec5b3 · inbound

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions cites this paper.

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:03.386398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:34:03.386398Z digest=sha256:b08573b98ef89561364ede37b6b3c399afb97f13902c6d7d41acda4f848e3f02

Observation c5f78117-c6cb-4e9f-b6c9-0ca77317fd44 · inbound

MCRL4OR: Multimodal Contrastive Representation Learning for Off-Road Environmental Perception cites this paper.

MCRL4OR: Multimodal Contrastive Representation Learning for Off-Road Environmental Perception VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:57:58.788470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:57:58.788470Z digest=sha256:71131e25875e71615764787d1ec876c0ee23d90837cee6c1b556b78f9e5f0702

Observation 257e0663-6362-4aa0-9f7b-3bc6d4e64084 · inbound

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos cites this paper.

BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-09T23:03:34.671442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:03:34.671442Z digest=sha256:fd4d33eeb6c388a2934467b0cd6784b16f778c13fbe1f86b4c5c015a6852b791

Observation 543902d4-b4fd-450d-a4e2-9555554887b0 · inbound

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation cites this paper.

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:50:57.952377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:50:57.952377Z digest=sha256:e9607820dd6fea03a2aedd631bd0235e43068a727b3149d47fa4cb5a10477eea

Observation 48be1796-333d-469a-afc5-bb86bdcd62a5 · inbound

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners cites this paper.

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T04:37:09.828306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:37:09.828306Z digest=sha256:7bc2df9d9f809a3025060b6f93c2e899c453b515fa1243dbf90bc72780e5f665

Observation 8f45b970-b8b4-4627-9524-83ad999afd2a · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.199614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.199614Z digest=sha256:1e0cf29cf3c4de080ac9568466ec747c8ae0499af4919a335cdff0107517ab65

Observation 56f2d5c3-e0e1-47d8-a3ce-3dc8cb32a4d8 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.250901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:517640c9dd263072295bf4c824405671b6d92efe103c49a444ddbcbdf5b8bc13

Observation 052cd404-e35f-47d8-b9a3-ced47168b153 · inbound

AdaVid: Adaptive Video-Language Pretraining cites this paper.

AdaVid: Adaptive Video-Language Pretraining VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.167427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.167427Z digest=sha256:ec5d452ec3efef42822bb20d1a1838096f86dcd87ea4d1f0b946b7614f0f9a28

Observation 43426e7b-abff-4d55-821e-d9b778167300 · inbound

Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare cites this paper.

Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:54.433543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:08:54.433543Z digest=sha256:b0df43efafeb13b13fbc78cb3dcfae3da0db012e59e1a8be3e6cfbda7b5d9915

Observation ef308635-7236-46cc-87c2-e4783267ee32 · inbound

Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis cites this paper.

Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:18.170294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:18.170294Z digest=sha256:3fceb4695652ae988327779c949626ce655e36af484ee1453f78422f6b80e8db

Observation 1d03ffe8-d8c3-4558-9a22-81ccdd0cfdac · inbound

Aligning Multimodal Representations through an Information Bottleneck cites this paper.

Aligning Multimodal Representations through an Information Bottleneck VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:53.900356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:43:53.900356Z digest=sha256:b7a53af35dd656f7a344d3c90286dc5c3382f9005d8336c5b14dcbe0ae5b4a3c

Observation bc0f11ee-3143-4498-8238-3cd7ceef2180 · inbound

Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection cites this paper.

Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:16.872047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:20:16.872047Z digest=sha256:27094a4fbbb64e8ab82e3713d072bd88a21df812c08117f6263d6a96a74c9d38

Observation bbecc0ac-db34-4231-bf89-b3a95d52059a · inbound

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding cites this paper.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.906801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.906801Z digest=sha256:305689ab7d4ceeee3783e809561c8d8bce4c09818fd602699bdc2a669c1c1ce7

Observation a2caf74b-d8a6-43da-894a-6d4a97a04b8b · inbound

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning cites this paper.

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:00.365272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:00.365272Z digest=sha256:cb493133a2cdc92cdd7aefdefe85d427fb58f7a0c38c7eb4b60f26ace014911c

Observation 41ba384c-a049-4610-a21c-093d0006747c · inbound

Bridging Brain with Foundation Models through Self-Supervised Learning cites this paper.

Bridging Brain with Foundation Models through Self-Supervised Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:35.139511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:47:35.139511Z digest=sha256:4a1a73a8198b3d9295efbfccb47a963125072250c418314709e1ddc8d941d316

Observation 61529d6c-d42b-45a7-b6ba-b13560387e97 · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:10:15.136859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:47c7967872ca4d2d8175a992609d3078cc57a7a6be199870ff67d43a8034bfc1

Observation 0ff0bbb6-eb8a-41d8-ba58-e1eb879b040d · inbound

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition cites this paper.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:56.989600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:56.989600Z digest=sha256:1b527ae8771fdfdea87a6b5c3b66568a0ebf9031f872c2ecdd3660a5216e1d31

Observation 4b2d807e-fe8f-405d-9432-e9dd4bc960fb · inbound

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation cites this paper.

Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:31.643197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:31.643197Z digest=sha256:fc4baf9bc0da8efb6fc0cd8729eb54aec78f5df4f7e60e1dc3ef598eda6fd6b5

Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · inbound

Implicit Counterfactual Learning for Audio-Visual Segmentation cites this paper.

Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.193705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.193705Z digest=sha256:bdc2fd4dcf9412e3cdaad8b99bde90082e961feeb0bb5067437af27ac05ca255

Observation 90e63cc7-bf95-4ea5-aa75-fa1289e5b836 · inbound

Group Relative Augmentation for Data Efficient Action Detection cites this paper.

Group Relative Augmentation for Data Efficient Action Detection VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:40.226378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:40.226378Z digest=sha256:42c3e985f266556e15a908cee09809f260263f4ddb29d67970e2f00a82da5ab5

Observation 2335abde-285a-4cac-8447-b1a2c9cbfa25 · inbound

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models cites this paper.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.517124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.517124Z digest=sha256:0d67914fea2668971e74fea399a026d33697da1d12f13a99a2ef19e461e10342

Observation 330c954d-ded8-446f-a47b-22ddfd931cff · inbound

Adversarial Video Promotion Against Text-to-Video Retrieval cites this paper.

Adversarial Video Promotion Against Text-to-Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.166881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T00:05:07.182361Z digest=sha256:6045bdcc3ab25631b582259c7cdf9ee11a114fb4839403a465e822b261df6550

Observation d6b50866-ab14-445b-a0f7-1b0e75e07e7a · inbound

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications cites this paper.

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:58.248635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:58.248635Z digest=sha256:9e7c9ea390d99062e1d753130f6ce2e38eee88a73231f930f94c6d3ac01a3b26

Observation 85192c1e-362f-4e51-b9a8-049c42e00291 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.187507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.187507Z digest=sha256:17ea860d804e83e3723dd7cc41958f094a4cf2adf1f0c04496dc1613e64aa343

Observation 9761cf3d-85dc-4d2b-94c8-8d97a3f83eff · inbound

Calibrated Multimodal Representation Learning with Missing Modalities cites this paper.

Calibrated Multimodal Representation Learning with Missing Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:10:22.692713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T22:08:07.217659Z digest=sha256:d11a8fe5feab67107d56d4549a1fc6e984af4961aba586968ee7a0db9029dc85

Observation a30d6193-0b77-4a33-8a28-fdcf02fa13f9 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.335434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:88b702b39ab2b77bbe7d42ff7b98af711a88958d0608c30b5e722085e9ed657a

Observation 51e029eb-ce1e-4431-9dbe-a30d6a95526c · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.921271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:13f905862051c48a29c546feef61f688e6943a16f1798a04d5bfce095d2aee32

Observation 9d82ce30-3dc5-4cb9-8455-d148e1611781 · inbound

Learning ORDER-Aware Multimodal Representations for Composite Materials Design cites this paper.

Learning ORDER-Aware Multimodal Representations for Composite Materials Design VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:40:14.550097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T14:36:44.656655Z digest=sha256:e5d3b453d423673b9199fc82bab8d2ef09834060bd8270affe707597f7ab2666

Observation 98a3f7aa-d2a1-49c0-a7f6-9169612aa447 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.570008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:c1bb7e658f721d8e370fd1930db4c369b25f4924c318fae299375807c8a9ed5d

Observation e5989aae-07c8-4169-a9bf-75f9fcec2bff · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:a4e90723e5339c52cee4dba26d9391443131d4157f22cf32a2729e0f33dc0fa1

Observation 79bc04fd-414f-4150-a5c8-710534ff86cb · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:1dfdaf9526f8c082a6b9c25f481678950ecab51333145dd8c7d44c7cf194733e

Observation c3f96828-e72e-4f38-953b-6c85b0fefb02 · inbound

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing cites this paper.

DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:53.943204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:44:52.768021Z digest=sha256:6e83aa6906dcff49451e53d3eeee7514c540c903f0fa393ea2baf07cf343c9af

Observation a286e855-f8a7-4198-b58a-1c9456c2bb14 · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.326399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:f039b71c0d89467b4c3f69ea2bc9e9a435213330074586bea0e21799b814a4d1

Observation bf843ac4-63bf-4441-91ca-748db06036a3 · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:21.931008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T12:07:21.203513Z digest=sha256:8a85b07bead2b40ec44eb7c050f81d76a2e4e66c3f49e7ba7ce22bebb54068b9

Observation ee162317-38b2-4cf0-8ca7-b03ee279b6d7 · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T20:05:20.702092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:05:20.702092Z digest=sha256:43ace0b72e08baf73acfd60f09211e52a6d2bea257952b732aa713323406412d

Observation e05b8581-ffea-4cf1-8fd9-ae8478f83ced · inbound

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation cites this paper.

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.280521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T20:04:39.623342Z digest=sha256:83c16b2eabd194a448843f0a099f64ecc5886980ecb788cfc65cfa7255e4060a

Observation 6d4afa95-e025-4962-a74a-fc809fa40d2b · inbound

Multimodal LLMs under Pairwise Modalities cites this paper.

Multimodal LLMs under Pairwise Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:39:40.686881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T05:37:56.792564Z digest=sha256:f4aadabc13a921c608943d6a440afa6074cb4c000e0c8705f706545254c7fe37

Observation b15f4643-7da8-4008-bbbf-485042c8b020 · inbound

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset cites this paper.

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:16.925691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T15:29:51.424572Z digest=sha256:fbe8a7b980e19277b94a3f88786fd47d0f80aaec3e73c38ee9cfcde9d11f4de9

Observation 8acc79c7-1e61-4a3c-85fb-d94956ca9fc6 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.382403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:0eb16edc8a3a55fb2e6ee14fdf28e80eb5ff1c5bfacdf2879f2fbdc56934952f

Observation e10c871d-faef-49c8-bac8-ad08c41aec11 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.133036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.133036Z digest=sha256:5b6e6cd1205cd2facef184bc60e94507aa376f4fb6e2089105bc894c2d27c840