Pith. sign in

Paper Citation Record · LEDGER

Efficient Transfer Learning for Video-language Foundation Models

As of 13 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2411.11223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11223 v4

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:51:32.331261Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c05c3e24-0778-469d-a368-8a74ab018d5c · outbound

This paper cites Quo vadis, action recognition? A new model and the kinetics dataset.

Efficient Transfer Learning for Video-language Foundation Models Quo vadis, action recognition? A new model and the kinetics dataset

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.941389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.157689Z digest=sha256:309c521f8cce3e639dc144d9ecd077b024aa87842c8d1616aff83e7edecd9b18

Observation bc68abc1-ba72-4fbc-83c8-1793b0d665b7 · outbound

This paper cites A Short Note about Kinetics-600.

Efficient Transfer Learning for Video-language Foundation Models A Short Note about Kinetics-600

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.163043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.163043Z digest=sha256:08878b66a8f02199d92e68a5e50c398e97d0b04950dc91389fb0f0a6571b268b

Observation 06def9af-1602-4a13-b661-b199d33f22f8 · outbound

This paper cites Conditional Prototype Rectification Prompt Learning.

Efficient Transfer Learning for Video-language Foundation Models Conditional Prototype Rectification Prompt Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:51:32.477848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.167655Z digest=sha256:e44687719c186013dda0ad5fb66b29bc0cdc0322a17bcfe8d853da458fdd1762

Observation 84871c5d-bd36-4ce6-9ab6-e5971deade99 · outbound

This paper cites Elaborative rehearsal for zero- shot action recognition.

Efficient Transfer Learning for Video-language Foundation Models Elaborative rehearsal for zero- shot action recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.927182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.172404Z digest=sha256:973e2bb3e27e0c4e56f0f719473981abc4073ae25513ac99fe88ac6ab727a0c3

Observation 02df26ee-5fed-4ccd-af74-e1c2748f3e0c · outbound

This paper cites Adaptformer: Adapt- ing vision transformers for scalable visual recognition.

Efficient Transfer Learning for Video-language Foundation Models Adaptformer: Adapt- ing vision transformers for scalable visual recognition

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.914675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.176661Z digest=sha256:c907da69ee38cb87cf1ab27c62acf1ac779fbf91f7580549ff08ca593f78b27b

Observation dd0a75b9-e3e4-428b-b984-d6f95203b989 · outbound

This paper cites OST: refining text knowledge with optimal spatio-temporal descriptor for general video recognition.

Efficient Transfer Learning for Video-language Foundation Models OST: refining text knowledge with optimal spatio-temporal descriptor for general video recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.901484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.180797Z digest=sha256:d75294ae9045a4cc3d6e23bca2cbabe33432562c2dc2f534e16e505b77336faa

Observation 64169dcd-7543-40cd-b53a-7a7859339e53 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Efficient Transfer Learning for Video-language Foundation Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.185399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.185399Z digest=sha256:cda9a4637202bf64518bb36ff6008a6ee6c083b9caaced34589de07a8021e837

Observation a73a580d-1cfc-47bd-8984-ce9645fcb818 · outbound

This paper cites BERT: pre-training of deep bidirectional trans- formers for language understanding.

Efficient Transfer Learning for Video-language Foundation Models BERT: pre-training of deep bidirectional trans- formers for language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.887646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.189867Z digest=sha256:9527fc09e9484e4d0b10e22d36167903b5a378d875bc523b8df60551b030f50c

Observation 51e2781d-4dfb-4305-991d-d954ceaa4437 · outbound

This paper cites De- coupling zero-shot semantic segmentation.

Efficient Transfer Learning for Video-language Foundation Models De- coupling zero-shot semantic segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.873866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.194005Z digest=sha256:563614c4642b8292addf8dbbdb14b71fa8459a426242ceecb57090ad02ae8450

Observation d24d709d-cf16-48ae-8cd8-637ddf5e4aca · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Efficient Transfer Learning for Video-language Foundation Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.198261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.198261Z digest=sha256:1f7a4a925bae101bad2aa29724b176cf30f101c87ece71d2ce77ace1b2711060

Observation c89b2d0f-cf04-4968-b43c-cda9938ac3f0 · outbound

This paper cites Zero-shot and few-shot video question answering with multi-modal prompts.

Efficient Transfer Learning for Video-language Foundation Models Zero-shot and few-shot video question answering with multi-modal prompts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.851728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.202528Z digest=sha256:9bd1c84c7e52fa7a98947061e650ab64d87f92b925f89a4e747cedb71df59333

Observation ef64be30-d0f7-4ecc-a849-f314d3d0eb7d · outbound

This paper cites Promptdet: Towards open-vocabulary detection using uncurated images.

Efficient Transfer Learning for Video-language Foundation Models Promptdet: Towards open-vocabulary detection using uncurated images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.838213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.206377Z digest=sha256:dbfe9e58940ef3c523ed3e80a4dc4b646b048d6cb2f7570f853cc2ff5434c560

Observation 4f0f3511-9d45-4e68-9ea5-bb245398a014 · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense.

Efficient Transfer Learning for Video-language Foundation Models The ”something something” video database for learning and evaluating visual common sense

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.825448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.209926Z digest=sha256:a5b59e8b31e142d0bf327ea74a0b088024a78692ad99b9866c73b4484a2f7527

Observation c30ee68f-84f5-41f6-a264-a0002c0a3b52 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level perfor- mance on imagenet classification.

Efficient Transfer Learning for Video-language Foundation Models Delving deep into rectifiers: Surpassing human-level perfor- mance on imagenet classification

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.811000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.213285Z digest=sha256:0f590de850cce08a6891784ba002a35f6fdcc62565f79808142c34d8cebe73a1

Observation d44bca27-8f3d-493a-9aff-f5a3b43d41e0 · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

Efficient Transfer Learning for Video-language Foundation Models Activitynet: A large-scale video bench- mark for human activity understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.797187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.217131Z digest=sha256:f3b614f7bf814bee7879b093739fddde7606cbf94dfe81cd25bfdb4a2cab934a

Observation a0b527de-e7ec-445c-aa64-96004d7203a7 · outbound

This paper cites Parameter-efficient transfer learning for NLP.

Efficient Transfer Learning for Video-language Foundation Models Parameter-efficient transfer learning for NLP

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.783218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.220676Z digest=sha256:fba86e0fd1d3e66a098b2f536c3eca968c38b06d5ad838dd650d4831a3d07653

Observation fb9340b1-89c0-49ca-8d43-cd58e4858111 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Efficient Transfer Learning for Video-language Foundation Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.768964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.224073Z digest=sha256:1b1a4252c29692e3db6469da2ec88ea7dac34c608721b15e58f08962500f91b1

Observation 0009498b-7b35-422f-88fa-d868da4253bd · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Efficient Transfer Learning for Video-language Foundation Models Prompting visual-language models for efficient video understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.756022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.227996Z digest=sha256:b686fc70ce183d79bca5e17a027b133d793f46b453fafedde556f616370cd0ee

Observation 2233691b-1772-495b-b8bc-66f9d17a9519 · outbound

This paper cites Khan, and Fahad Shahbaz Khan.

Efficient Transfer Learning for Video-language Foundation Models Khan, and Fahad Shahbaz Khan

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.742319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.231907Z digest=sha256:b1f3732e9ab73c3d6e4e50b7ac9ccada3dc0b2887886924da2c6a077002a468b

Observation 6fe6aa39-a558-40da-93cd-f0564bc07fdb · outbound

This paper cites Poggio, and Thomas Serre.

Efficient Transfer Learning for Video-language Foundation Models Poggio, and Thomas Serre

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.728632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.236217Z digest=sha256:7d0952bc15e0320304c660b9b14b4d07c8b4f194bf352e0fdb43acec2ffd8bf2

Observation 5f8e94e0-9fba-466a-a96c-785bfdd9c4c4 · outbound

This paper cites an unresolved cited work.

Efficient Transfer Learning for Video-language Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:51:32.714590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.240182Z digest=sha256:c758181e26bdf2e03d9742f45cbb23fd69c81b74e99c1b9d62d6fb1781c3b078

Observation b1f5c81a-c436-4ae8-8a69-dea6eea4028c · outbound

This paper cites Decoupled weight decay regularization.

Efficient Transfer Learning for Video-language Foundation Models Decoupled weight decay regularization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.244229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.244229Z digest=sha256:2f618f611fc64948e1a2af3cef9e7986629c77f227f3ed86778ed92193ae2eea

Observation dcdf8144-d4b5-431f-8236-67c2c8e8532a · outbound

This paper cites an unresolved cited work.

Efficient Transfer Learning for Video-language Foundation Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:51:32.693148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.248199Z digest=sha256:a5bf8c99b75747eb9627385370c189b01f10e088d482dc8dc5c46fad24752c13

Observation e7c95845-99dd-4121-a314-3af3bea6996b · outbound

This paper cites Expanding language-image pretrained models for general video recognition.

Efficient Transfer Learning for Video-language Foundation Models Expanding language-image pretrained models for general video recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.681209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.252190Z digest=sha256:0d2c49b4b4e3ee4ca93107639026fa47ca81597545044c72b43712448f9f353d

Observation f2ba8a1d-2690-4d1e-93e9-4de8e9658d3f · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Efficient Transfer Learning for Video-language Foundation Models St-adapter: Parameter-efficient image-to-video transfer learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.667757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.256194Z digest=sha256:7d8a0231400bbb3b27b2ca160e6b86e4839aaf88d46757904b9f49ee917daa29

Observation 2a9380f7-b780-4151-9c3f-689c1e994bc9 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Efficient Transfer Learning for Video-language Foundation Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.260261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.260261Z digest=sha256:d2ea764806bcf69afdbbdef4b8d62d5417fc55714cb066c723fc6d922423246d

Observation 374f26f8-07b6-48d9-a5d2-1cd86c548d40 · outbound

This paper cites Haupt- mann.

Efficient Transfer Learning for Video-language Foundation Models Haupt- mann

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.653674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.264647Z digest=sha256:a20f34bdd46d33e0d4cf370f970cf3ae5faabd0288461658db4fcdc23c7f1a16

Observation 0d756e45-12dc-4f09-a5e3-5c1ee58985c3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Efficient Transfer Learning for Video-language Foundation Models Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.268714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.268714Z digest=sha256:c92a6b9d6b8973b3cda88fe6b060364bbb36986bfe07f61820c8bb35b24ad798

Observation 8a18f7be-87a1-4fa1-90bf-1c4c2c66c01d · outbound

This paper cites Khan, and Fahad Shahbaz Khan.

Efficient Transfer Learning for Video-language Foundation Models Khan, and Fahad Shahbaz Khan

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.631293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.272633Z digest=sha256:ba8dc9b7a7ef777de6a35a74cbc6b886f2cdd2d140af1545e3827c8404b7fe77

Observation cd0f600c-53c6-4a02-aefc-ffce390002b5 · outbound

This paper cites Consistency-guided prompt learning for vision-language models.

Efficient Transfer Learning for Video-language Foundation Models Consistency-guided prompt learning for vision-language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.617576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.276907Z digest=sha256:b40589ca32a51bb38808a50a7b48066c1d6d681777a11c787357fc678b593aaf

Observation a9605f1f-4c3d-4b1e-aad5-a329a9c4cddb · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Efficient Transfer Learning for Video-language Foundation Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.280808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.280808Z digest=sha256:934f92628ba95bdfcaf29308d36fce8a6dc09270d0fdce5863360db8a07d49a1

Observation f79f76ee-461f-4f68-9051-2391cd078b98 · outbound

This paper cites Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Efficient Transfer Learning for Video-language Foundation Models Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.604232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.285071Z digest=sha256:2dc81cb352b9dacb7d2ec094c7c27f2ee6d566a651ad7e4bf9787872c84e4aa3

Observation 51ae5c49-5551-4193-8894-c2e01e346428 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Efficient Transfer Learning for Video-language Foundation Models Representation Learning with Contrastive Predictive Coding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.288909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.288909Z digest=sha256:1b742d1289343a5b25eedd73ccc2e0f18c14f696bad421a7e4526df42cef9b7e

Observation f257c5ef-fb14-4fb8-b53a-c31a223251dc · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Efficient Transfer Learning for Video-language Foundation Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.293134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.293134Z digest=sha256:ab885ae0e488e58454fb1f27479d30355612023bffdac7366d77127f8afc4c48

Observation c447dbe6-346e-4474-ac13-16f21c6f1b55 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Efficient Transfer Learning for Video-language Foundation Models InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.297314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.297314Z digest=sha256:545377b47c31eb47e4af67173ffc3769ab156d4e383058eeef7472b3cbf680d0

Observation 29ffe277-2538-41a7-b97b-982aa7d11776 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

Efficient Transfer Learning for Video-language Foundation Models Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.590644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.301670Z digest=sha256:c69981c5ef767cc17efb5ee0f41ced7e66b459e294949ef0b2542c83167af027

Observation 2fac2da3-01b3-467d-99af-5091d5ddbd49 · outbound

This paper cites Khan, Fa- had Shahbaz Khan, and Mubarak Shah.

Efficient Transfer Learning for Video-language Foundation Models Khan, Fa- had Shahbaz Khan, and Mubarak Shah

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.576314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.305641Z digest=sha256:c21c6791dcf36fb4d126adf16e30691d741927ce4af4bacdd411eb554d96383f

Observation 17a54940-769a-48e2-9215-5c7565ea1cb1 · outbound

This paper cites Open-vclip: Transforming CLIP to an open-vocabulary video model via interpolated weight optimization.

Efficient Transfer Learning for Video-language Foundation Models Open-vclip: Transforming CLIP to an open-vocabulary video model via interpolated weight optimization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.561506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.309484Z digest=sha256:b61bb96118bdb3d8010726d81d8c8e908087f6513e0e2e178121de69d62b9ab0

Observation 5a0174dd-4fa7-477b-a97e-061346b1add6 · outbound

This paper cites CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching.

Efficient Transfer Learning for Video-language Foundation Models CORA: adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.548119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.313381Z digest=sha256:5b7a806789cde01f0a4c315d52d8ab1908f9dc4d707c41dc4b351a1fd71c1ae5

Observation f6e9fd4c-d2e2-428e-b48b-7ef8972cf0ed · outbound

This paper cites MMA: multi-modal adapter for vision-language models.

Efficient Transfer Learning for Video-language Foundation Models MMA: multi-modal adapter for vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.534607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.316907Z digest=sha256:b9fbe4ca3b6ffa6c80e6a190fe3d0dacee6a4e2a014c5dc09cc581393c746561

Observation 61b81311-cc84-437b-91f9-c8e6c963e082 · outbound

This paper cites AIM: adapting image models for efficient video action recognition.

Efficient Transfer Learning for Video-language Foundation Models AIM: adapting image models for efficient video action recognition

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.520390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.320603Z digest=sha256:fb14a99d729a66f66896dc9d5461f79952517e6f95ccf00bf920c48bc0f0e814

Observation 409cab93-66aa-40f5-8c71-e07c99cb1ee7 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Efficient Transfer Learning for Video-language Foundation Models Florence: A New Foundation Model for Computer Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.324016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.324016Z digest=sha256:cd654e8d7b5f4e7cf812de69a55d8ccf339a5660f2928b4114b7165f852d0a10

Observation 93e7e4cb-bf08-40b8-9098-95b03725c978 · outbound

This paper cites Tip- adapter: Training-free adaption of clip for few-shot classifica- tion.

Efficient Transfer Learning for Video-language Foundation Models Tip- adapter: Training-free adaption of clip for few-shot classifica- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:51:32.506560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T18:51:32.327769Z digest=sha256:ac851f22b58caaf201921fcd821ff4650fde566b52b56b90299e3a4d409c1945

Observation 7844c60d-ab71-4bc1-a07b-e8d458cbb83e · outbound

This paper cites MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer.

Efficient Transfer Learning for Video-language Foundation Models MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.331261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:51:32.331261Z digest=sha256:1c211b4da8fc3c478d9a80d598f48e261143afd61017a706a2e9ce90134045f4

Pith citing papers

No inbound Pith citation observations are available.