Pith. sign in

Paper Citation Record · LEDGER

CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2106.11097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.11097 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:09.738221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T04:48:54.299822Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d75c229d-e26a-4636-9a92-0421d108cc6f · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.683764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:23032dfbc1a68e0a491459d0ec0bc139747338820b91900cd8d7b20dbbd32efe

Observation 4a47bb5f-7b6b-42c4-8b69-3b555bf773fd · inbound

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels cites this paper.

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:48:54.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T04:46:18.542521Z digest=sha256:5c24d62d72a5a8f7cc76c71044f4a459097a626fc03387d37937f140e2142dc4

Observation b372a1b7-fd12-4298-83af-85acfe7bdb65 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.297596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:1dd5e925151202ca4f79265232573766b6d600c33643e1899be1077ea94795ac

Observation 06498fb3-0c1d-4cf3-921d-2ca0b4ee5cfb · inbound

A Mathematical Perspective On Contrastive Learning cites this paper.

A Mathematical Perspective On Contrastive Learning CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.738221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.738221Z digest=sha256:787d672fb8c2f4225488996a0025b35c18581ce098d98967b9fb25cdfb72d6c6

Observation ddea426d-f18b-44c0-99df-70fe5494a31e · inbound

Learning Speaker-Invariant Visual Features for Lipreading cites this paper.

Learning Speaker-Invariant Visual Features for Lipreading CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:57.052410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:57.052410Z digest=sha256:93e53708640960abcfc5086607828543eda27cd6f2af869be09f0b455ec5cf99

Observation 475e19fe-8987-4049-afcf-545444f1c113 · inbound

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval cites this paper.

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:36:27.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:36:27.471308Z digest=sha256:b61dc7e8cb376dee41fd2bc0b58d13bbf9cf14864b40463c0c121c0b5de59fba

Observation 8adc1884-5841-4d92-9623-872b8dd0785b · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.766732Z digest=sha256:f359bafd1043e25ba23d58bec04ccc79a13c38a8e485881db06c5d201edf9350

Observation 18390aa0-4e10-4ff7-94b5-3a14c7a0bbbc · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.914340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:2f4c30a4fce28cd32cf6ecd965ec1f95a3702b2339788e30dff8ca3774b2d6d3

Observation 7bc0539d-8372-4464-8dcb-3e49fe2c6d7f · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.665420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:5034c97310710ce04eb2d72fdf8f3cc5ccf1a7fda9a92fc9abcd21647c819edb

Observation 15cdf844-dbd2-4316-8418-c8d476152921 · inbound

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation cites this paper.

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.314300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T20:04:39.623342Z digest=sha256:d4a55d7de2dffbe4792d7a6b8ce73bab3bd31b1f9facfeea339c78764c280f35

Observation e7f17c4f-79fc-4248-98ae-e47fbf162147 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.640974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:7381e09e165c61ce51a3a3704e0adf5b00db6304f28c2c9efb5f9b2be3819e08

Observation 599e28ca-8d4e-46ce-9590-96fdadc8c36b · inbound

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval cites this paper.

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T04:19:16.924621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T04:19:16.924621Z digest=sha256:dd0e1b68b3d49a08dbb1cfc3b6cd7b2b121642454d9e688a7d55bd2c2e5f82dc

Observation 8de72bdb-d372-4f2c-9b39-7a4f0850d79a · inbound

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification cites this paper.

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:00.942499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:01:00.942499Z digest=sha256:c388437ce18f1d1815b575314ccfb86e5f0c4b9a20eff0bdf7e3abf8549ce84d