Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:16:15.771212Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.18675.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:16:15.771212Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c906201-ce95-4b06-8361-e5bf22231b03 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Leveraging Vision -Language Models for Improving Domain Generalization in Image Classification
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b42e8a9c-025b-42ac-a052-645e77104a56 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Are Visual-Language Models Effective Action Recognition? A Comparative Study
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5091d515-2eca-4f6b-b38b-2a0d3f6841d2 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision–language model for visual question answering in medical imagery
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c64017ab-cd13-4b9a-a5f9-3f23b1f9ff1f · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks When deep learners change their mind: Learning dynamics for active learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7116ff25-bad7-432e-913f-b73209cc6692 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks An Introduction to Vision-Language Modeling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a18d4b3-1509-478e-b4d8-5816046489a7 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PracticalDG: Perturbation Distillation on Vision- Language Models for Hybrid Domain Generalization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4cb39969-b1da-4a9e-a20f-d0374c62e8fe · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Pub - medclip: How much does clip benefit visual question answering in the medical domain?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a01ce183-c5e8-4d54-9b26-c00771e8b294 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Clipsyntel: clip and llm synergy for multi - modal question summarization in healthcare
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8599834e-4962-45b3-afce-5fc894760d96 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learn2augment: learning to composite videos for data augmentation in action recognition
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 99c5f218-63d8-4483-a43b-e727a90ef1ea · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Class-Specific Noise Injection for Improved Road Segmentation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9c393a4b-7fef-4d1f-aaf6-e1540c4185f7 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Temporal Modeling Approach for Video Action Recognition Based on Vision-language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e05f7981-d6fa-474f-97f9-3828fb363c00 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Perturbation- based methods for explaining deep neural networks: A survey
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a1a36eec-d4d2-4eaf-b7f6-39708a0b1ad7 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks A dversarial attack on yolov5 for traffic and road sign detection
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2814c903-1eea-4b62-8ec2-a39b84838e12 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e99755b9-7c87-4a02-8f9c-c7eebdc838d3 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Language augmentation in clip for improved anatomy detection on multi-modal medical images
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c0a9211b-c188-4843-a68b-8cbbf12409c7 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PALM: Predicting Actions through Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a182daa-20e5-425d-a840-eb420f81bad6 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Segment Anything
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e691084-371f-4eee-9828-7132922ec243 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57251939-f85f-4b06-908c-fd9499ee11db · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Enhancing clip with gpt-4: Harness- ing visual descriptions as prompts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation da2b2fb5-34e3-45a8-afb0-7628970b5995 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Learning transferable visual models from nat- ural language supervision
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7f484ea5-da3b-4100-8470-099fa43f2cb0 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision language models are blind: Failing to translate detailed visual features into words
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331a88c7-9971-4d18-8149-35d44671dac7 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf7199f-94ae-4eed-86a3-d5d6845f63a9 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Test-time prompt tuning for zero-shot general- ization in vision-language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bf7884b7-6150-40d6-b059-befcade4fa8c · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b738c268-606b-4352-b927-1ba254d6912f · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Motionclip: Exposing human motion genera - tion to clip space
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0b0f2520-2a16-4ca4-b37d-d030f0fcd319 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3929c2c-f003-455f-b096-1c14574783e1 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315d1def-45df-424e-974c-f6743b83bf3c · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks ActionCLIP: A New Paradigm for Video Action Recognition
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dbdee37-4579-415f-b08d-0d3b788052dc · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Actionclip: Adapting language-image pre- trained models for video action recognition
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d57a36ad-9822-466a-87d4-44028a193620 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Incorporating Scene Graphs into Pre-trained Vision-Language Models for Multimodal Open-vocabulary Action Recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 088b7585-9ecb-44ad-bdeb-4998e8802493 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba183e20-0401-4f89-9bb3-aaceb3f15e20 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Revisiting classi - fier: Transferring vision-language models for video recognition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e8df1d71-a006-40a7-aed9-07546bda48d5 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Investigating Compositional Challenges in Vision- Language Models for Visual Grounding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9bd7a2d6-847d-41af-9280-a3edbb1c6d6b · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks PeVL : Pose -Enhanced Vision-Language Model for Fine-Grained Human Action Recognition
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 88089a6f-6aff-4ae9-978b-d96fc226bd36 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision -language models for vision tasks: A survey
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c828b911-0368-4390-9f3a-ef08f5229124 · outbound
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks CLIP in Medical Imaging: A Survey
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.