Pith. sign in

Paper Citation Record · LEDGER

CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2106.11097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.11097 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:09.738221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T04:48:54.299822Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d75c229d-e26a-4636-9a92-0421d108cc6f · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.683764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:7cda251de9d0b3de45c4afd6910cc61447eb26913a5771fec8fd4775e007cfe1

Observation 4a47bb5f-7b6b-42c4-8b69-3b555bf773fd · inbound

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels cites this paper.

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:48:54.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T04:46:18.542521Z digest=sha256:8139eb21595f21329cb3142788d28b53694b005e6d4ce7a3666666bf547e2ad5

Observation b372a1b7-fd12-4298-83af-85acfe7bdb65 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.297596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:a0cb6feabd9ae6ce473319b37cf4a28673c079a55835bc0adea25e9a4ad4cf35

Observation 06498fb3-0c1d-4cf3-921d-2ca0b4ee5cfb · inbound

A Mathematical Perspective On Contrastive Learning cites this paper.

A Mathematical Perspective On Contrastive Learning CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.738221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.738221Z digest=sha256:787d672fb8c2f4225488996a0025b35c18581ce098d98967b9fb25cdfb72d6c6

Observation ddea426d-f18b-44c0-99df-70fe5494a31e · inbound

Learning Speaker-Invariant Visual Features for Lipreading cites this paper.

Learning Speaker-Invariant Visual Features for Lipreading CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:57.052410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:57.052410Z digest=sha256:93e53708640960abcfc5086607828543eda27cd6f2af869be09f0b455ec5cf99

Observation 475e19fe-8987-4049-afcf-545444f1c113 · inbound

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval cites this paper.

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:36:27.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:36:27.471308Z digest=sha256:b61dc7e8cb376dee41fd2bc0b58d13bbf9cf14864b40463c0c121c0b5de59fba

Observation 8adc1884-5841-4d92-9623-872b8dd0785b · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.766732Z digest=sha256:f359bafd1043e25ba23d58bec04ccc79a13c38a8e485881db06c5d201edf9350

Observation 18390aa0-4e10-4ff7-94b5-3a14c7a0bbbc · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.914340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:1871e55b9dbc55ab1547ed32d87bb56188f6c51c80c8233d75da14575543af5f

Observation 7bc0539d-8372-4464-8dcb-3e49fe2c6d7f · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.665420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:94aa7cf0ad16984d996ea0e0dd541d2968929b69bae069c87e9243021fe89a4b

Observation 15cdf844-dbd2-4316-8418-c8d476152921 · inbound

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation cites this paper.

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.314300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T20:04:39.623342Z digest=sha256:51a30a1ebc2115b88dacbb64b4d1d0b250ecb3bfef3f9da0f8bb124fcfdeb6fa

Observation e7f17c4f-79fc-4248-98ae-e47fbf162147 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.640974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:330b8123aa337dfc9901fb5cf01960c533acf8bb79d5f74b2f7cefc8bcfbf1eb

Observation 599e28ca-8d4e-46ce-9590-96fdadc8c36b · inbound

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval cites this paper.

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T04:19:16.924621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T04:19:16.924621Z digest=sha256:dd0e1b68b3d49a08dbb1cfc3b6cd7b2b121642454d9e688a7d55bd2c2e5f82dc

Observation 8de72bdb-d372-4f2c-9b39-7a4f0850d79a · inbound

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification cites this paper.

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:00.942499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:01:00.942499Z digest=sha256:c388437ce18f1d1815b575314ccfb86e5f0c4b9a20eff0bdf7e3abf8549ce84d