Pith. sign in

Paper Citation Record · LEDGER

Efficient Transfer Learning for Video-language Foundation Models

As of 13 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2411.11223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11223 v4

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:51:32.331261Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c05c3e24-0778-469d-a368-8a74ab018d5c · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset.

Efficient Transfer Learning for Video-language Foundation Models Quo vadis, action recognition? A new model and the kinetics dataset

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.941389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.157689Z digest=sha256:1108ea35ae5b8c147d475764084e5f7340b26b12c1547a7a47db538187fda47f

Observation bc68abc1-ba72-4fbc-83c8-1793b0d665b7 · outbound

This paper cites A Short Note about Kinetics-600.

Efficient Transfer Learning for Video-language Foundation Models A Short Note about Kinetics-600

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.163043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.163043Z digest=sha256:08878b66a8f02199d92e68a5e50c398e97d0b04950dc91389fb0f0a6571b268b

Observation 06def9af-1602-4a13-b661-b199d33f22f8 · outbound

This paper cites Conditional Prototype Rectification Prompt Learning.

Efficient Transfer Learning for Video-language Foundation Models Conditional Prototype Rectification Prompt Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:51:32.477848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.167655Z digest=sha256:2fe8d09fc642e7d53ea077b8191de102c4ed635f97274d4dd2dc7862e7cd2f92

Observation 84871c5d-bd36-4ce6-9ab6-e5971deade99 · outbound

This paper cites Elaborative rehearsal for zero- shot action recognition.

Efficient Transfer Learning for Video-language Foundation Models Elaborative rehearsal for zero- shot action recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.927182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.172404Z digest=sha256:44668b3c3c2d348c63c87e46b8a809dc03ba5bb53d0683e9cb111985ebc2b68e

Observation 02df26ee-5fed-4ccd-af74-e1c2748f3e0c · outbound

This paper cites Adaptformer: Adapt- ing vision transformers for scalable visual recognition.

Efficient Transfer Learning for Video-language Foundation Models Adaptformer: Adapt- ing vision transformers for scalable visual recognition

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.914675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.176661Z digest=sha256:fc08c0c21b66e69c5fd6bcb9cdb2d3f3547391d2cee1581ae9f8b2b66ce2a863

Observation dd0a75b9-e3e4-428b-b984-d6f95203b989 · outbound

This paper cites OST: refining text knowledge with optimal spatio-temporal descriptor for general video recognition.

Efficient Transfer Learning for Video-language Foundation Models OST: refining text knowledge with optimal spatio-temporal descriptor for general video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.901484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.180797Z digest=sha256:394eff103baa14e169364199519d6620eb6de20867b6ac3457b210958def413b

Observation 64169dcd-7543-40cd-b53a-7a7859339e53 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Efficient Transfer Learning for Video-language Foundation Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.185399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.185399Z digest=sha256:cda9a4637202bf64518bb36ff6008a6ee6c083b9caaced34589de07a8021e837

Observation a73a580d-1cfc-47bd-8984-ce9645fcb818 · outbound

This paper cites BERT: pre-training of deep bidirectional trans- formers for language understanding.

Efficient Transfer Learning for Video-language Foundation Models BERT: pre-training of deep bidirectional trans- formers for language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.887646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.189867Z digest=sha256:8718c28ad8d6e66b59bd37fa604d4f366186228a67f2ca3bffd3ed923d3de666

Observation 51e2781d-4dfb-4305-991d-d954ceaa4437 · outbound

This paper cites De- coupling zero-shot semantic segmentation.

Efficient Transfer Learning for Video-language Foundation Models De- coupling zero-shot semantic segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.873866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.194005Z digest=sha256:bbfd0416933d95cf3564049cd99bba5c7d41048672bc20d29036fa97c5dcad28

Observation d24d709d-cf16-48ae-8cd8-637ddf5e4aca · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Efficient Transfer Learning for Video-language Foundation Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.198261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.198261Z digest=sha256:1f7a4a925bae101bad2aa29724b176cf30f101c87ece71d2ce77ace1b2711060

Observation c89b2d0f-cf04-4968-b43c-cda9938ac3f0 · outbound

This paper cites Zero-shot and few-shot video question answering with multi-modal prompts.

Efficient Transfer Learning for Video-language Foundation Models Zero-shot and few-shot video question answering with multi-modal prompts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.851728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.202528Z digest=sha256:a724b29f3c1e273520274022924f1a0d0246352af7d360c961b7b8aa3de9252d

Observation ef64be30-d0f7-4ecc-a849-f314d3d0eb7d · outbound

This paper cites Promptdet: Towards open-vocabulary detection using uncurated images.

Efficient Transfer Learning for Video-language Foundation Models Promptdet: Towards open-vocabulary detection using uncurated images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.838213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.206377Z digest=sha256:98951d1dde503fbce4b0a0e37f533d22aef59406ffd8a6d289e101f77f6b107d

Observation 4f0f3511-9d45-4e68-9ea5-bb245398a014 · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense.

Efficient Transfer Learning for Video-language Foundation Models The ”something something” video database for learning and evaluating visual common sense

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.825448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.209926Z digest=sha256:aed5d65356d5f39085c8a6000181217adbcf27dea816b6aa946028ab67d0f66d

Observation c30ee68f-84f5-41f6-a264-a0002c0a3b52 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level perfor- mance on imagenet classification.

Efficient Transfer Learning for Video-language Foundation Models Delving deep into rectifiers: Surpassing human-level perfor- mance on imagenet classification

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.811000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.213285Z digest=sha256:358d0ad2e4fb21a0d0b1d414896421cb403d3145c17be2d43403d410ce8a7128

Observation d44bca27-8f3d-493a-9aff-f5a3b43d41e0 · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

Efficient Transfer Learning for Video-language Foundation Models Activitynet: A large-scale video bench- mark for human activity understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.797187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.217131Z digest=sha256:f5076cb0a9c21e404bf71b45cd1cec383bd4cb9dd317465442329a5c2142e209

Observation a0b527de-e7ec-445c-aa64-96004d7203a7 · outbound

This paper cites Parameter-efficient transfer learning for NLP.

Efficient Transfer Learning for Video-language Foundation Models Parameter-efficient transfer learning for NLP

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.783218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.220676Z digest=sha256:d8f7da054824547345d1268445a3bf2000c66f97e469185e541287e5db7a05b7

Observation fb9340b1-89c0-49ca-8d43-cd58e4858111 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Efficient Transfer Learning for Video-language Foundation Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.768964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.224073Z digest=sha256:15bb193b23d421cf97445f2d50114512e2d026e4e7668758941cc8c55f0f518d

Observation 0009498b-7b35-422f-88fa-d868da4253bd · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Efficient Transfer Learning for Video-language Foundation Models Prompting visual-language models for efficient video understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.756022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.227996Z digest=sha256:b5e9646542d4c127c0e0762a45353ebab3e4eb2cb06c98bc27e0d563ce4fe334

Observation 2233691b-1772-495b-b8bc-66f9d17a9519 · outbound

This paper cites Khan, and Fahad Shahbaz Khan.

Efficient Transfer Learning for Video-language Foundation Models Khan, and Fahad Shahbaz Khan

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.742319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.231907Z digest=sha256:54287fe73a421624c2ba2303d9ca878410cb2d9fd7ba90d0d5c36e1cbd509cdb

Observation 6fe6aa39-a558-40da-93cd-f0564bc07fdb · outbound

This paper cites Poggio, and Thomas Serre.

Efficient Transfer Learning for Video-language Foundation Models Poggio, and Thomas Serre

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.728632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.236217Z digest=sha256:9794559c28c4b10d4f867a22893298c1f103a2d3822d29b31f3ef3451e51c7c8

Observation 5f8e94e0-9fba-466a-a96c-785bfdd9c4c4 · outbound

This paper cites an unresolved cited work.

Efficient Transfer Learning for Video-language Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:51:32.714590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.240182Z digest=sha256:5be84611d715ba98a16bf4830d9ba974c557dacc71bdbcd4eb59faccc303e9b4

Observation b1f5c81a-c436-4ae8-8a69-dea6eea4028c · outbound

This paper cites Decoupled weight decay regularization.

Efficient Transfer Learning for Video-language Foundation Models Decoupled weight decay regularization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.244229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.244229Z digest=sha256:2f618f611fc64948e1a2af3cef9e7986629c77f227f3ed86778ed92193ae2eea

Observation dcdf8144-d4b5-431f-8236-67c2c8e8532a · outbound

This paper cites an unresolved cited work.

Efficient Transfer Learning for Video-language Foundation Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:51:32.693148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.248199Z digest=sha256:468ba1d1771d253dc2a2cbd0f51d5bdba77306a86ed6c1fefab5811204d5609b

Observation e7c95845-99dd-4121-a314-3af3bea6996b · outbound

This paper cites Expanding language-image pretrained models for general video recognition.

Efficient Transfer Learning for Video-language Foundation Models Expanding language-image pretrained models for general video recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.681209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.252190Z digest=sha256:28958f6763de7d9163c02d1b2ca53853caa743d2d37f5f8a1ce7167714fd56c7

Observation f2ba8a1d-2690-4d1e-93e9-4de8e9658d3f · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Efficient Transfer Learning for Video-language Foundation Models St-adapter: Parameter-efficient image-to-video transfer learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.667757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.256194Z digest=sha256:f9b1943336156b675e945b751798c795cd9d434b6393202b3d1ef3eccf53009d

Observation 2a9380f7-b780-4151-9c3f-689c1e994bc9 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Efficient Transfer Learning for Video-language Foundation Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.260261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.260261Z digest=sha256:d2ea764806bcf69afdbbdef4b8d62d5417fc55714cb066c723fc6d922423246d

Observation 374f26f8-07b6-48d9-a5d2-1cd86c548d40 · outbound

This paper cites Haupt- mann.

Efficient Transfer Learning for Video-language Foundation Models Haupt- mann

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.653674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.264647Z digest=sha256:645f57deaf7b39f007d2588f232edb7afa5748812d565df704d2afbd17b1ea85

Observation 0d756e45-12dc-4f09-a5e3-5c1ee58985c3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Efficient Transfer Learning for Video-language Foundation Models Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.268714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.268714Z digest=sha256:c92a6b9d6b8973b3cda88fe6b060364bbb36986bfe07f61820c8bb35b24ad798

Observation 8a18f7be-87a1-4fa1-90bf-1c4c2c66c01d · outbound

This paper cites Khan, and Fahad Shahbaz Khan.

Efficient Transfer Learning for Video-language Foundation Models Khan, and Fahad Shahbaz Khan

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.631293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.272633Z digest=sha256:f4fb449d3bce37febdd3cd55d995e55734bccb1a6aad12a1862ba5cf7b5d69ac

Observation cd0f600c-53c6-4a02-aefc-ffce390002b5 · outbound

This paper cites Consistency-guided prompt learning for vision-language models.

Efficient Transfer Learning for Video-language Foundation Models Consistency-guided prompt learning for vision-language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.617576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.276907Z digest=sha256:6d52cef91529b2d4bcc4d9c8b43d12869aec73c6654296f2ed74d6eeee113037

Observation a9605f1f-4c3d-4b1e-aad5-a329a9c4cddb · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Efficient Transfer Learning for Video-language Foundation Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.280808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.280808Z digest=sha256:934f92628ba95bdfcaf29308d36fce8a6dc09270d0fdce5863360db8a07d49a1

Observation f79f76ee-461f-4f68-9051-2391cd078b98 · outbound

This paper cites Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Efficient Transfer Learning for Video-language Foundation Models Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.604232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.285071Z digest=sha256:f5d3dd48fe3f54a271c00b4814c21b63ff0d1f3ccc490feb1983b0b8b1058606

Observation 51ae5c49-5551-4193-8894-c2e01e346428 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Efficient Transfer Learning for Video-language Foundation Models Representation Learning with Contrastive Predictive Coding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.288909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.288909Z digest=sha256:1b742d1289343a5b25eedd73ccc2e0f18c14f696bad421a7e4526df42cef9b7e

Observation f257c5ef-fb14-4fb8-b53a-c31a223251dc · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Efficient Transfer Learning for Video-language Foundation Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.293134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.293134Z digest=sha256:ab885ae0e488e58454fb1f27479d30355612023bffdac7366d77127f8afc4c48

Observation c447dbe6-346e-4474-ac13-16f21c6f1b55 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Efficient Transfer Learning for Video-language Foundation Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.297314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.297314Z digest=sha256:545377b47c31eb47e4af67173ffc3769ab156d4e383058eeef7472b3cbf680d0

Observation 29ffe277-2538-41a7-b97b-982aa7d11776 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

Efficient Transfer Learning for Video-language Foundation Models Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.590644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.301670Z digest=sha256:2f7efdc2da719e0c77ef9214118e6d03011fdf00d09a4b02480429352e4e31b1

Observation 2fac2da3-01b3-467d-99af-5091d5ddbd49 · outbound

This paper cites Khan, Fa- had Shahbaz Khan, and Mubarak Shah.

Efficient Transfer Learning for Video-language Foundation Models Khan, Fa- had Shahbaz Khan, and Mubarak Shah

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.576314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.305641Z digest=sha256:2f62d47411e3663056880395a2c0af7f140fa5b22bf1e5c01ea0f5ed9c2f2967

Observation 17a54940-769a-48e2-9215-5c7565ea1cb1 · outbound

This paper cites Open-vclip: Transforming CLIP to an open-vocabulary video model via interpolated weight optimization.

Efficient Transfer Learning for Video-language Foundation Models Open-vclip: Transforming CLIP to an open-vocabulary video model via interpolated weight optimization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.561506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.309484Z digest=sha256:6ea1d3e0def33f88df6e00fd4590e1f346351ae8bd37c703b3c585087692e459

Observation 5a0174dd-4fa7-477b-a97e-061346b1add6 · outbound

This paper cites CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching.

Efficient Transfer Learning for Video-language Foundation Models CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.548119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.313381Z digest=sha256:fa7d179931cae1bd8662575f08ebed21ab8d7e9a314bda96179469051d146eab

Observation f6e9fd4c-d2e2-428e-b48b-7ef8972cf0ed · outbound

This paper cites MMA: multi-modal adapter for vision-language models.

Efficient Transfer Learning for Video-language Foundation Models MMA: multi-modal adapter for vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.534607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.316907Z digest=sha256:55bf2829a5b20ee220818af08c251b0f48fa63ce1399bdfac75a0e5864bcbe55

Observation 61b81311-cc84-437b-91f9-c8e6c963e082 · outbound

This paper cites AIM: adapting image models for efficient video action recognition.

Efficient Transfer Learning for Video-language Foundation Models AIM: adapting image models for efficient video action recognition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.520390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.320603Z digest=sha256:d22ac788deb538d49e7df7b570a87ab64e4ec88a6f7b79940d560997514710b9

Observation 409cab93-66aa-40f5-8c71-e07c99cb1ee7 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Efficient Transfer Learning for Video-language Foundation Models Florence: A New Foundation Model for Computer Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.324016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.324016Z digest=sha256:cd654e8d7b5f4e7cf812de69a55d8ccf339a5660f2928b4114b7165f852d0a10

Observation 93e7e4cb-bf08-40b8-9098-95b03725c978 · outbound

This paper cites Tip- adapter: Training-free adaption of clip for few-shot classifica- tion.

Efficient Transfer Learning for Video-language Foundation Models Tip- adapter: Training-free adaption of clip for few-shot classifica- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.506560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:51:32.327769Z digest=sha256:0beb98e303a330b81f5be8ad3399b9a1db4db62e56e2a16625be49d62c15628f

Observation 7844c60d-ab71-4bc1-a07b-e8d458cbb83e · outbound

This paper cites MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer.

Efficient Transfer Learning for Video-language Foundation Models MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.331261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.331261Z digest=sha256:1c211b4da8fc3c478d9a80d598f48e261143afd61017a706a2e9ce90134045f4

Pith citing papers

No inbound Pith citation observations are available.