Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2109.14084.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:07:31.643197Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:39:37.380967Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation fe34e4e4-4f16-43b4-b379-342a81ff67a2 · inbound
R3M: A Universal Visual Representation for Robot Manipulation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 72761ef6-7d17-4890-8c76-6cabcc30f73d · inbound
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1eaf6c1f-7f3f-4dbd-8fe1-793a40eb840d · inbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b9b1d4cd-aa91-48eb-b583-b569e1e5952e · inbound
VideoChat: Chat-Centric Video Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eed293b8-bc7f-4447-a1a7-672174934fd9 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f4791b68-3222-428b-9824-0255dc411a5c · inbound
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d4cba46-2e95-48a7-8e42-e8a936162640 · inbound
Revisiting Feature Prediction for Learning Visual Representations from Video VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 292
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 56f2d5c3-e0e1-47d8-a3ce-3dc8cb32a4d8 · inbound
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61529d6c-d42b-45a7-b6ba-b13560387e97 · inbound
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b2d807e-fe8f-405d-9432-e9dd4bc960fb · inbound
Enhancing Remote Sensing Vision-Language Models Through MLLM and LLM-Based High-Quality Image-Text Dataset Generation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · inbound
Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e63cc7-bf95-4ea5-aa75-fa1289e5b836 · inbound
Group Relative Augmentation for Data Efficient Action Detection VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2335abde-285a-4cac-8447-b1a2c9cbfa25 · inbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 330c954d-ded8-446f-a47b-22ddfd931cff · inbound
Adversarial Video Promotion Against Text-to-Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6b50866-ab14-445b-a0f7-1b0e75e07e7a · inbound
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85192c1e-362f-4e51-b9a8-049c42e00291 · inbound
Video Understanding by Design: How Datasets Shape Video Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 224
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9761cf3d-85dc-4d2b-94c8-8d97a3f83eff · inbound
Calibrated Multimodal Representation Learning with Missing Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a30d6193-0b77-4a33-8a28-fdcf02fa13f9 · inbound
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51e029eb-ce1e-4431-9dbe-a30d6a95526c · inbound
Adapting MLLMs for Nuanced Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d82ce30-3dc5-4cb9-8455-d148e1611781 · inbound
Learning ORDER-Aware Multimodal Representations for Composite Materials Design VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 98a3f7aa-d2a1-49c0-a7f6-9169612aa447 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5989aae-07c8-4169-a9bf-75f9fcec2bff · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bc04fd-414f-4150-a5c8-710534ff86cb · inbound
CoVR-R:Reason-Aware Composed Video Retrieval VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f96828-e72e-4f38-953b-6c85b0fefb02 · inbound
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a286e855-f8a7-4198-b58a-1c9456c2bb14 · inbound
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf843ac4-63bf-4441-91ca-748db06036a3 · inbound
DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee162317-38b2-4cf0-8ca7-b03ee279b6d7 · inbound
DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05b8581-ffea-4cf1-8fd9-ae8478f83ced · inbound
Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d4afa95-e025-4962-a74a-fc809fa40d2b · inbound
Multimodal LLMs under Pairwise Modalities VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b15f4643-7da8-4008-bbbf-485042c8b020 · inbound
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8acc79c7-1e61-4a3c-85fb-d94956ca9fc6 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.