Pith. sign in

Paper Citation Record · LEDGER

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.18675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18675 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:16:15.771212Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c906201-ce95-4b06-8361-e5bf22231b03 · outbound

This paper cites Leveraging Vision -Language Models for Improving Domain Generalization in Image Classification.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Leveraging Vision -Language Models for Improving Domain Generalization in Image Classification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.363661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.602955Z digest=sha256:5ab8aba042d5ff459701bb9ec1f5aea30cfbed3076d67577abc4b1420c3d5775

Observation b42e8a9c-025b-42ac-a052-645e77104a56 · outbound

This paper cites Are Visual-Language Models Effective Action Recognition? A Comparative Study.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Are Visual-Language Models Effective Action Recognition? A Comparative Study

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.348253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.608943Z digest=sha256:232c19c1523eec9749c9cd5718f1ae57958e6a857c3db27300f7f6728cd80a03

Observation 5091d515-2eca-4f6b-b38b-2a0d3f6841d2 · outbound

This paper cites Vision–language model for visual question answering in medical imagery.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision–language model for visual question answering in medical imagery

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:16:16.333516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.613602Z digest=sha256:d02dd83d0b6e82a49103efae626b0b49a94602501eb9baa3da66c6c1216d3783

Observation c64017ab-cd13-4b9a-a5f9-3f23b1f9ff1f · outbound

This paper cites When deep learners change their mind: Learning dynamics for active learning.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks When deep learners change their mind: Learning dynamics for active learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.317730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.618379Z digest=sha256:53f774b7fce884129a28c998e5e54b58c7a9d51c0f4a765173d13878bb5565bb

Observation 7116ff25-bad7-432e-913f-b73209cc6692 · outbound

This paper cites An Introduction to Vision-Language Modeling.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks An Introduction to Vision-Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.623253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.623253Z digest=sha256:76e1107ff54af1ba5a44734a301df5ff09fb32cc252edb4320b0cdb3371abcab

Observation 5a18d4b3-1509-478e-b4d8-5816046489a7 · outbound

This paper cites PracticalDG: Perturbation Distillation on Vision- Language Models for Hybrid Domain Generalization.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PracticalDG: Perturbation Distillation on Vision- Language Models for Hybrid Domain Generalization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.301352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.628962Z digest=sha256:f130e7b91071162dc9355fa51776d1e4eb26a138b9185446bfc481900ca4038f

Observation 4cb39969-b1da-4a9e-a20f-d0374c62e8fe · outbound

This paper cites Pub - medclip: How much does clip benefit visual question answering in the medical domain?.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Pub - medclip: How much does clip benefit visual question answering in the medical domain?

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.284660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.634590Z digest=sha256:4519656f184b987c4eebe697f53fcde63c9237b208a75be0dc05ac3642570c7c

Observation a01ce183-c5e8-4d54-9b26-c00771e8b294 · outbound

This paper cites Clipsyntel: clip and llm synergy for multi - modal question summarization in healthcare.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Clipsyntel: clip and llm synergy for multi - modal question summarization in healthcare

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.268728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.639482Z digest=sha256:8a392370cff9eb95ddfef4f8514f624040ba7709d9e7e476959c776e0b88d379

Observation 8599834e-4962-45b3-afce-5fc894760d96 · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in action recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learn2augment: learning to composite videos for data augmentation in action recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.253854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.643907Z digest=sha256:87a174bf63248730cae1037c9938cd2033b265d6486d5ea7bd947354335a58e2

Observation 99c5f218-63d8-4483-a43b-e727a90ef1ea · outbound

This paper cites Class-Specific Noise Injection for Improved Road Segmentation.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Class-Specific Noise Injection for Improved Road Segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.239299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.648059Z digest=sha256:b45e8d364d740a67dacfda6a5079a7a2cf61baeb4dc8c7f875dbc961397d5bc1

Observation 9c393a4b-7fef-4d1f-aaf6-e1540c4185f7 · outbound

This paper cites Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.223687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.652509Z digest=sha256:daebb575765b1c5f2fba7c44228f33b0efb35fc7334d712929e526db6191438d

Observation e05f7981-d6fa-474f-97f9-3828fb363c00 · outbound

This paper cites Perturbation- based methods for explaining deep neural networks: A survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Perturbation- based methods for explaining deep neural networks: A survey

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.207039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.657146Z digest=sha256:2464630685ee9b6cdbed3df2fd3e89f3ceebff58aa602328a2e0a49fee0abf64

Observation a1a36eec-d4d2-4eaf-b7f6-39708a0b1ad7 · outbound

This paper cites A dversarial attack on yolov5 for traffic and road sign detection.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks A dversarial attack on yolov5 for traffic and road sign detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.192083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.661596Z digest=sha256:85c82560141524e53a1e7e0388fc16dc227493d5d88f74e5b1a2cebb9e93cae4

Observation 2814c903-1eea-4b62-8ec2-a39b84838e12 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.176666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.666027Z digest=sha256:fc825b17bc2404fc8748a61cfca440767c687a6b4621c41822b16538da51e0eb

Observation e99755b9-7c87-4a02-8f9c-c7eebdc838d3 · outbound

This paper cites Language augmentation in clip for improved anatomy detection on multi-modal medical images.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Language augmentation in clip for improved anatomy detection on multi-modal medical images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.161183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.670475Z digest=sha256:52bfb3810182acff0d5d0a3ba10835689a9285fb908968ef831487ad379961fe

Observation c0a9211b-c188-4843-a68b-8cbbf12409c7 · outbound

This paper cites PALM: Predicting Actions through Language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PALM: Predicting Actions through Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:16:15.971825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.675732Z digest=sha256:f092fd1c5b85e4763864756b8fc6bfa1136e6ca34907cf8b8ecea5c6f0431641

Observation 2a182daa-20e5-425d-a840-eb420f81bad6 · outbound

This paper cites Segment Anything.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Segment Anything

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.680772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.680772Z digest=sha256:264a9c1950f9cda4ed3a67e22862f42e25e9fd855c9d5e0f983d7da9c0b7c004

Observation 6e691084-371f-4eee-9828-7132922ec243 · outbound

This paper cites Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.685735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.685735Z digest=sha256:c3688d8eb9b69e680c02cc064b55f25690967352046fccf273e809b17e44069e

Observation 57251939-f85f-4b06-908c-fd9499ee11db · outbound

This paper cites Enhancing clip with gpt-4: Harness- ing visual descriptions as prompts.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Enhancing clip with gpt-4: Harness- ing visual descriptions as prompts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.145623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.690697Z digest=sha256:d585e8750375a8951da92c82b834215cfc800e432e5a3f59816b3d0a46fdadb9

Observation da2b2fb5-34e3-45a8-afb0-7628970b5995 · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learning transferable visual models from nat- ural language supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.130369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.695366Z digest=sha256:9a77c10196bacee9fdb0984db8902b095358ed1905c2657ab4c826024e60b6d7

Observation 7f484ea5-da3b-4100-8470-099fa43f2cb0 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.699872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.699872Z digest=sha256:bf414cea1d26d5cda47f816a93cae1e0e58c585e367172f165be9dcec3c2bcad

Observation 331a88c7-9971-4d18-8149-35d44671dac7 · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.704561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.704561Z digest=sha256:b389368af8eff8d36f6f6e8693a841779ac6f1e38d46c85824f97cef9c4e75f5

Observation ddf7199f-94ae-4eed-86a3-d5d6845f63a9 · outbound

This paper cites Test-time prompt tuning for zero-shot general- ization in vision-language models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Test-time prompt tuning for zero-shot general- ization in vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.115214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.709612Z digest=sha256:0d6494ae4f3df173ef9d0c81ab7cff1b5be57e7ad5929d9e1c43805aab3bfc64

Observation bf7884b7-6150-40d6-b059-befcade4fa8c · outbound

This paper cites Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.713836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.713836Z digest=sha256:c4253ec26bfe266d1c74b151f1adc1d4c3708a83cb5b5f5eaa19382bf1ff6204

Observation b738c268-606b-4352-b927-1ba254d6912f · outbound

This paper cites Motionclip: Exposing human motion genera - tion to clip space.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Motionclip: Exposing human motion genera - tion to clip space

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.100219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.718612Z digest=sha256:62b0f5b147291ad83a200c165907b1ede7ebd921f2987cfb2d015dbcb783b548

Observation 0b0f2520-2a16-4ca4-b37d-d030f0fcd319 · outbound

This paper cites XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.723157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.723157Z digest=sha256:05f7827684ed0d153884b6f95cec81241eb1ffb4dd26f404e927111c8a92a209

Observation e3929c2c-f003-455f-b096-1c14574783e1 · outbound

This paper cites CLIP with Quality Captions: A Strong Pretraining for Vision Tasks.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP with Quality Captions: A Strong Pretraining for Vision Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.728021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.728021Z digest=sha256:26afabcaf0b22485a562819c5ae69c6496b575b18ddc1da299f60d9c650b2cb7

Observation 315d1def-45df-424e-974c-f6743b83bf3c · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks ActionCLIP: A New Paradigm for Video Action Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.732959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.732959Z digest=sha256:e7a50cb399d7e778ce83ce79e114c44ff6505c35c67bd55e032c7f0ae0dd55ed

Observation 8dbdee37-4579-415f-b08d-0d3b788052dc · outbound

This paper cites Actionclip: Adapting language-image pre- trained models for video action recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Actionclip: Adapting language-image pre- trained models for video action recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.084863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.737740Z digest=sha256:169623705069ee5675c3423127db0c461b4194b90c09f96434cf3076c977a6f0

Observation d57a36ad-9822-466a-87d4-44028a193620 · outbound

This paper cites Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.067933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.742143Z digest=sha256:7c749198758d41f7de81a08e4d2e0bfc206a3f361f0df01e9f142ec76713bfdd

Observation 088b7585-9ecb-44ad-bdeb-4998e8802493 · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.746661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.746661Z digest=sha256:c0620e68567661a4b6eea6181141bd2175f6f13fb3b6da4b2a2717b6486d36f9

Observation ba183e20-0401-4f89-9bb3-aaceb3f15e20 · outbound

This paper cites Revisiting classi - fier: Transferring vision-language models for video recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Revisiting classi - fier: Transferring vision-language models for video recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.051759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.751782Z digest=sha256:f6200d466f488c4de0d9a3572a1bee91a7292e98831d732e2d59398b514bf8fb

Observation e8df1d71-a006-40a7-aed9-07546bda48d5 · outbound

This paper cites Investigating Compositional Challenges in Vision- Language Models for Visual Grounding.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Investigating Compositional Challenges in Vision- Language Models for Visual Grounding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.036514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.756446Z digest=sha256:56f998ddba33c95ea37f4d27149820a9977d4a5743840fedcf33973813ce35fe

Observation 9bd7a2d6-847d-41af-9280-a3edbb1c6d6b · outbound

This paper cites PeVL : Pose -Enhanced Vision-Language Model for Fine-Grained Human Action Recognition.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PeVL : Pose -Enhanced Vision-Language Model for Fine-Grained Human Action Recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.020935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.760727Z digest=sha256:9b9b0323a96f7405938f55fba8a35deee72bdf57e1e1735c85d3a5127c816524

Observation 88089a6f-6aff-4ae9-978b-d96fc226bd36 · outbound

This paper cites Vision -language models for vision tasks: A survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision -language models for vision tasks: A survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:16:16.004907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:16:15.765144Z digest=sha256:896e5e4f971908b594e982a8f325156c1a05cd1ef47cd546362242addd3eb89f

Observation c828b911-0368-4390-9f3a-ef08f5229124 · outbound

This paper cites CLIP in Medical Imaging: A Survey.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP in Medical Imaging: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.771212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.771212Z digest=sha256:e7a32c4d2490108be4da51f2f2ec53aeb124d87f7e01c6ae7fd43e904f513658

Pith citing papers

No inbound Pith citation observations are available.