Pith. sign in

Paper Citation Record · LEDGER

CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2104.08860.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.08860 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:15:12.259822Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:09:15.255843Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7e639819-48cb-481d-bca9-afea9da9f0da · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.789565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:212eb5e7fc7387ccd8074a3617a509635bc77ea6dae6368aca3c0e68425da35d

Observation d552b34c-5ce0-4e56-8486-79dd1ee6230b · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:20:20.353154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:c9e011bdbb6cf2b45c9571121511389014685544e70b783b0c1fe89c6df5eb31

Observation e45785c1-8ef1-41aa-a520-d9f87f4335e9 · inbound

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels cites this paper.

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:48:54.297379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T04:46:18.542521Z digest=sha256:9183b621a227e9959db43fe6a83f0c318e0b7ebe36a33f9b5e04ad0299a5fde8

Observation a0f7aa7c-4b58-4382-9d42-9f1779120058 · inbound

LLaVA-Video: Video Instruction Tuning With Synthetic Data cites this paper.

LLaVA-Video: Video Instruction Tuning With Synthetic Data CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 210

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:33.141307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T23:20:32.330351Z digest=sha256:d0a0357167e2049b3d5d6e59b697b14aec4bc3a49a08fc840a6551ab273aaaae

Observation 3a1d64b6-eb39-477a-98f1-62d3e1a5c280 · inbound

AstroM$^3$: A self-supervised multimodal model for astronomy cites this paper.

AstroM$^3$: A self-supervised multimodal model for astronomy CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:23:18.768419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:23:18.768419Z digest=sha256:e1bf79576687fbf53d46ca5bfebbde65555cc381b7cb3f10c9e5f7b8f610fffd

Observation 544c7b6c-8643-477f-b6c9-ee214541a8b9 · inbound

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning cites this paper.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.988340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.988340Z digest=sha256:5b8fd38b038c8c8b5852992506c256472563dfecaaef3bdfe59eab8c6c3fb5fb

Observation 8ea95b71-15e6-45ac-a3b7-82a473567ab1 · inbound

Needle: A Generative AI-Powered Multi-modal Database for Answering Complex Natural Language Queries cites this paper.

Needle: A Generative AI-Powered Multi-modal Database for Answering Complex Natural Language Queries CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T05:13:50.127973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:13:50.127973Z digest=sha256:9ae7373933b7d9d4730210801af6523041a6c68abb236ea2864bf286ad92115b

Observation 0d5d01f8-1b03-425e-bab1-55dc2bbe0c20 · inbound

MADGEN: Mass-Spec attends to De Novo Molecular generation cites this paper.

MADGEN: Mass-Spec attends to De Novo Molecular generation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:20:45.980215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:20:45.980215Z digest=sha256:702858787a71595e45007f9a4bb5a7cc60411a4bbbb89461931eda97bb99f63a

Observation 92c26dc6-4b83-4553-a6ca-f3662ab335d6 · inbound

Vision-Language Models Do Not Understand Negation cites this paper.

Vision-Language Models Do Not Understand Negation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.149164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.149164Z digest=sha256:efb23dbc0bc224b72c92399bde3d26b5bf9be0b0bb951ddb23402bfbc5e53a9c

Observation 4a36bfaa-22be-4fb5-92c1-7335d9074463 · inbound

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval cites this paper.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.747047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.747047Z digest=sha256:c9634dc4d8d2d71fc2be4e54a2db3b423ef0495653e9d9ad02d7e01e3b88d328

Observation 7b6c544f-d6ff-45c9-ac0f-04e709b8b115 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:26.509966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:26.509966Z digest=sha256:4c97b73d497fccf6355b654bb09b293a28d7ade150a9077ef7e6347854e21533

Observation 17b3a461-14f7-47f8-a8ff-fbd05e9e0b22 · inbound

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges cites this paper.

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:04.406744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:04.406744Z digest=sha256:1fb07b1601290f1a75efc519dfa5b9c83ddd2ac77f40ef651dddce5916bc8e18

Observation 97af0fe0-60ea-4c3a-b795-75a52cc6c807 · inbound

Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking cites this paper.

Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:18:52.176751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:18:52.176751Z digest=sha256:1e2a224d82c049206fc1b9c91cbcdda703c412b3baa67d0e961537a4668e08eb

Observation 499b8bba-365c-49d0-80ec-fbbeb4520bba · inbound

Regularizing Subspace Redundancy of Low-Rank Adaptation cites this paper.

Regularizing Subspace Redundancy of Low-Rank Adaptation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:36.569663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:23:36.569663Z digest=sha256:4e68515e85d1ce33baf6c28247d6be3216a0653a25436bb6c7d24a522c68d956

Observation 14da950d-3ede-4d13-a8c1-66fb2b4f2b5a · inbound

CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter cites this paper.

CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:15:12.259822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:15:12.259822Z digest=sha256:0941034cad5be18fca46f63e4201d36b3b38ff20d8158cb89debd1c8f00ac46b

Observation a0f7ebcc-fc1a-409f-a5c6-ebf1340f57dd · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.236365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.236365Z digest=sha256:58c7652db2753764a767d3a2a952cbe472adb93900a2f92d9bfd02000c14aa7d

Observation f26efacf-4cd4-448b-b021-45248dea36ad · inbound

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval cites this paper.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.954552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.954552Z digest=sha256:bda9bf68be0ae55eccfc316c98f45db1ce5cdc68c2adb22d1089f785ac694048

Observation 9ce8db31-0a39-4f83-822f-0b39afdb2979 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.841437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:2b40cbad7b3d1f81a38cf64f863089f5e40901e16ae13b0b0396ed42cc7e5867

Observation b2eb13b6-a277-40e2-a154-7f6a4c60b214 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.054638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:f60c4aa7f41a91384437f7007de1823233f3c27e16cad2d24b364a907f046ebf

Observation 586e7a31-488a-496e-9a38-222c98c78703 · inbound

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation cites this paper.

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:15:58.956061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:45:36.306400Z digest=sha256:a05a240f2b44316c6b1d54235049f40da6c1ce92796a6f2c4551f98357b20460

Observation 5a9f84e7-0509-4cf8-a276-f9ec489e4730 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:10.396147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:e16f1de4d1f2649e33a0610c8c0b1ac354a5a8ce1d5aea6b981aace29d823fd7

Observation e8c95ba4-e1e7-4552-aedb-8b882c9f613f · inbound

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search cites this paper.

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.810751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T13:39:20.777029Z digest=sha256:ab3c7dc6b243a71a9fb0b832f7e3303e267424e2d6283d01b67d4dbb155f52f4

Observation 87cfe60d-3ef6-44fe-9c5a-546387cc0ee6 · inbound

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations cites this paper.

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:28.874868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T07:20:46.651636Z digest=sha256:2bf5a0676997692fc524a528bcc6dcb02e898e485b5f198a17b05faae8b5efa6

Observation ba1939d0-95c7-4c34-8bab-9027a636fa86 · inbound

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations cites this paper.

Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.073515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T21:52:31.915302Z digest=sha256:911c0c4c4b8376dbe7a19799937066a8324c27938066e1a146e145844f8ac069

Observation 9764b215-1560-48d0-ab4e-fc4cabaa5b5c · inbound

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation cites this paper.

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.561178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-29T18:08:50.574960Z digest=sha256:597594152b9a9ae1125e29d4abc99f467780628d0ec72c4cc86bca688dc22e7d

Observation d78b6606-e4de-4570-a2a5-daa99bb65ab7 · inbound

Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins cites this paper.

Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.526835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T03:26:49.843308Z digest=sha256:dc021d70f56bb49eca3576c4d7840851ab6532de7b29cddd2e37d9e7332db8c3

Observation 52547ef6-7fe1-4c88-a9a0-ab204e70d174 · inbound

LARE: Low-Attention Region Encoding for Text-Image Retrieval cites this paper.

LARE: Low-Attention Region Encoding for Text-Image Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:09:15.257933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T21:25:12.373068Z digest=sha256:43e2f2db0118946fb92ff2b530417cccc1adc52419bad5c9e6c0e6b9d5209d0f

Observation 76321b28-b6b5-4d49-8a86-e750b1d2fd32 · inbound

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement cites this paper.

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:17:07.335570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-02T15:09:47.855795Z digest=sha256:e3bccdcfe9be65556da23e08739e4c98119e4c7232e6a5202e4695164ebf3b81

Observation eb3295a6-4ac8-4ff7-b996-ca452ecc1e35 · inbound

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing cites this paper.

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T08:55:44.783824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T08:55:44.783824Z digest=sha256:4a6bdace9153f32c0c6879407b2df2416210e7047a94f6923264dba03b8111ec

Observation 559a8468-8043-4bd0-ad4b-54f04ec750ab · inbound

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data cites this paper.

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T14:53:56.693464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:53:56.693464Z digest=sha256:7085186204a812aa279e1b7b14b112df1ea341389d28aebf51bdcf03573c5f82

Observation 371c535c-6c7c-4f15-b28d-535a3c791952 · inbound

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification cites this paper.

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:00.937550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:01:00.937550Z digest=sha256:df217fa39d443144bce19def396b05f4e543de742583d553d9125f17de0a5dea

Observation 879e6ef0-36cd-460b-ab67-452bc62fe40c · inbound

Trajectory-aware Cross-view Geo-localization with Sequential Observations cites this paper.

Trajectory-aware Cross-view Geo-localization with Sequential Observations CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:14:55.353507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:14:55.353507Z digest=sha256:ad4e90cbdea433d0431c8eff516c697189904bc02a5d0cd72fc2d9de9c4abf77

Observation 482c4a35-b131-4ba4-a023-80cb7c0e3616 · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:30:56.884700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:30:56.884700Z digest=sha256:05c2fef487c88dd74ed8c2a8ce448c0b61f5f62da168ad12032c644873a1ee0e

Observation 3c406d12-39be-46f6-a23c-b57b58c50c7c · inbound

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition cites this paper.

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:39.899176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:02:39.899176Z digest=sha256:be84c4749e1e664ed3ad8553d09c02868965a68a27721db29260aa4d6225ad89