Pith. sign in

Paper Citation Record · LEDGER

CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2104.08860.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.08860 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:23:18.768419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:09:15.255843Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7e639819-48cb-481d-bca9-afea9da9f0da · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.789565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:8b0629b6a4d63a2556c4a8acd08d1f9df46f7fea8b929e446c5fb2b2f67e7400

Observation d552b34c-5ce0-4e56-8486-79dd1ee6230b · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:20:20.353154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:f2456103ad5ff10280fa421086aeac102f260e7f8960b4d67e77e77551646983

Observation e45785c1-8ef1-41aa-a520-d9f87f4335e9 · inbound

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels cites this paper.

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:48:54.297379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T04:46:18.542521Z digest=sha256:6680689503ce0e3d9d8242f9466cf5be32b17b7b874f8833e7ab93af21ae5ce3

Observation a0f7aa7c-4b58-4382-9d42-9f1779120058 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 210

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.141307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:72c7b8c24ba51477c11b42635cc6b3c68837051e2881ac93abedd5ac04861cfe

Observation 3a1d64b6-eb39-477a-98f1-62d3e1a5c280 · inbound

AstroM$^3$: A self-supervised multimodal model for astronomy cites this paper.

AstroM$^3$: A self-supervised multimodal model for astronomy CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:23:18.768419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:23:18.768419Z digest=sha256:16795460d2c538a5bc311f30bf0f3bbaceea71537157f38b2c0937f8efc16fbf

Observation 544c7b6c-8643-477f-b6c9-ee214541a8b9 · inbound

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning cites this paper.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.988340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.988340Z digest=sha256:7840212253b2c3bd3c77ce923a3c4b38ceff39f15b2dd45f1d125bf0aedb636d

Observation 8ea95b71-15e6-45ac-a3b7-82a473567ab1 · inbound

Needle: A Generative AI-Powered Multi-modal Database for Answering Complex Natural Language Queries cites this paper.

Needle: A Generative AI-Powered Multi-modal Database for Answering Complex Natural Language Queries CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:13:50.127973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:13:50.127973Z digest=sha256:eaf2c88f7ac05bd1c8a52ded088d1856b727a0954e9da5261b5884fbd01aa3a8

Observation 0d5d01f8-1b03-425e-bab1-55dc2bbe0c20 · inbound

MADGEN: Mass-Spec attends to De Novo Molecular generation cites this paper.

MADGEN: Mass-Spec attends to De Novo Molecular generation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:20:45.980215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:20:45.980215Z digest=sha256:67e0f3b62fdda3fce6f779759e80ec46444460bb8e6558943848560d432f1bfb

Observation 92c26dc6-4b83-4553-a6ca-f3662ab335d6 · inbound

Vision-Language Models Do Not Understand Negation cites this paper.

Vision-Language Models Do Not Understand Negation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.149164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.149164Z digest=sha256:b3b4a4019810a22d953d273dbbc115510dba0e5b4b5fd643127c534bd4e42502

Observation 4a36bfaa-22be-4fb5-92c1-7335d9074463 · inbound

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval cites this paper.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.747047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.747047Z digest=sha256:3000ff3458f92f10f19ae4ba8ef892e035a977ac6d73e5e71bdd81f117d21683

Observation 7b6c544f-d6ff-45c9-ac0f-04e709b8b115 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:26.509966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:26.509966Z digest=sha256:53924adea85d2cb52aa0295b77ad87f567fca1ae666cabc9e019231538a2d21c

Observation 17b3a461-14f7-47f8-a8ff-fbd05e9e0b22 · inbound

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges cites this paper.

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:04.406744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:04.406744Z digest=sha256:a21090fbcb3b15bf9f105bcafdeca1cea10e02fa73f18db4749db7caaa9155c3

Observation 97af0fe0-60ea-4c3a-b795-75a52cc6c807 · inbound

Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking cites this paper.

Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:18:52.176751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:18:52.176751Z digest=sha256:83f669c3cd660d088d303fca40cf757573134e5b962e0139fbbd23153dd91f76

Observation 499b8bba-365c-49d0-80ec-fbbeb4520bba · inbound

Regularizing Subspace Redundancy of Low-Rank Adaptation cites this paper.

Regularizing Subspace Redundancy of Low-Rank Adaptation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:36.569663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:23:36.569663Z digest=sha256:be4d3706f50d88e0e146e45397805aaf52ec8350dafa0eb76efc13530968f30b

Observation a0f7ebcc-fc1a-409f-a5c6-ebf1340f57dd · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.236365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.236365Z digest=sha256:2eefac3d99a446e6e1ac974daee9d740d0727eac75fa4f0778e29f19ca6cd042

Observation f26efacf-4cd4-448b-b021-45248dea36ad · inbound

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval cites this paper.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.954552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.954552Z digest=sha256:0460814b437b64e37bc44da3663582f5cc698ba5dc7270fffca237fc56d759b2

Observation 9ce8db31-0a39-4f83-822f-0b39afdb2979 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.841437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:2c2bdffd9a08cef13f42e0d194c59a7728f1e926dd20b34ca4835925b0f2e1aa

Observation b2eb13b6-a277-40e2-a154-7f6a4c60b214 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.054638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:637ace63a9861341bec5192da69d736885960dff3e9ca5b61d45d8ed27a1ec43

Observation 586e7a31-488a-496e-9a38-222c98c78703 · inbound

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation cites this paper.

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:15:58.956061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:45:36.306400Z digest=sha256:1f3d3c0c5ca483c5232e11c50644a87a72873bbf9eef762650ef50ceeda18134

Observation 5a9f84e7-0509-4cf8-a276-f9ec489e4730 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:10.396147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:1e8b9e4e1e52d93c6cb62029ba56d9c78b95b68f65d354f916e329cc67b47f70

Observation e8c95ba4-e1e7-4552-aedb-8b882c9f613f · inbound

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search cites this paper.

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.810751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T13:39:20.777029Z digest=sha256:b3917865c1ef5a611c00a5bf06cf6e4636e659e04fbad28408b7c58beadab9b7

Observation 87cfe60d-3ef6-44fe-9c5a-546387cc0ee6 · inbound

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations cites this paper.

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:28.874868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T07:20:46.651636Z digest=sha256:46a5e0589850dc448630884650eded06376ba417daca6b3b52f3f98a03ca8f52

Observation ba1939d0-95c7-4c34-8bab-9027a636fa86 · inbound

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations cites this paper.

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.073515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-14T21:52:31.915302Z digest=sha256:eb0a82d11b50272829f1f4f33edba2fd1d98efb65a739d859844e9cf4b414572

Observation 9764b215-1560-48d0-ab4e-fc4cabaa5b5c · inbound

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation cites this paper.

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.561178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T18:08:50.574960Z digest=sha256:0087ef208f3dd1d61c0a1e1c216aad787e14beaa289ec25963c72e684bac8559

Observation d78b6606-e4de-4570-a2a5-daa99bb65ab7 · inbound

Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins cites this paper.

Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.526835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:26:49.843308Z digest=sha256:1e1542fac69cd252a9c3eca9681ce00fec83774e80c341982fda84060eb86caf

Observation 52547ef6-7fe1-4c88-a9a0-ab204e70d174 · inbound

LARE: Low-Attention Region Encoding for Text-Image Retrieval cites this paper.

LARE: Low-Attention Region Encoding for Text-Image Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:09:15.257933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T21:25:12.373068Z digest=sha256:ba4eb138d017aaf18f2fc3dd9399e6c6c27e7b54d4b2f42e2b36efdd296865e0

Observation 76321b28-b6b5-4d49-8a86-e750b1d2fd32 · inbound

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement cites this paper.

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:17:07.335570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T15:09:47.855795Z digest=sha256:8b95a999978e96f83a482531619e7088cb8046194d94554169b0a478e45c3593

Observation eb3295a6-4ac8-4ff7-b996-ca452ecc1e35 · inbound

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing cites this paper.

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T08:55:44.783824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T08:55:44.783824Z digest=sha256:fd9f136dacb1342572602e7dbb8cc147b8d8278dfb9f4fcf5cb7c2d690e763d6

Observation 559a8468-8043-4bd0-ad4b-54f04ec750ab · inbound

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data cites this paper.

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T14:53:56.693464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:53:56.693464Z digest=sha256:48d9662ad1cccbfe027ab0cf541669b4b48d7ae8451ccd2bc8a7f4700f62af67

Observation 371c535c-6c7c-4f15-b28d-535a3c791952 · inbound

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification cites this paper.

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:00.937550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:01:00.937550Z digest=sha256:8558c7c0d369f0fcc995426887987dde2f52923cea2f65b814d4452b61b9a887

Observation 879e6ef0-36cd-460b-ab67-452bc62fe40c · inbound

Trajectory-aware Cross-view Geo-localization with Sequential Observations cites this paper.

Trajectory-aware Cross-view Geo-localization with Sequential Observations CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.353507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.353507Z digest=sha256:41112e2f89f880a1ada203766edd025c9e5de17869d696dac832b1f3d110eaed

Observation 482c4a35-b131-4ba4-a023-80cb7c0e3616 · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:30:56.884700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:30:56.884700Z digest=sha256:9f4a53f7fb9ce8fc6f78bcbb6c71fc80ad30ab2774e410527e2c7c29bc31575e

Observation 3c406d12-39be-46f6-a23c-b57b58c50c7c · inbound

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition cites this paper.

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:39.899176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:02:39.899176Z digest=sha256:9db819c6430b8147686f9fb88307385381d5b718d0e622625fe22aea9e3293bb