Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:46:02.319525Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2506.23502.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:46:02.319525Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-17T04:46:34.946640Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T04:49:02.937990Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f4e51e50-4dc3-49c3-ae30-6e3617a91fff · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Incorporating geo-diverse knowledge into prompting for increased geographical robustness in object recognition
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c74f03bd-4c91-43f0-991b-d0ac68c1b878 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning the best pooling strategy for visual semantic embedding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9652da36-73e0-41f2-a322-e5c10925e01e · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Large language models are visual reasoning coordinators.NeurIPS, pages 10–16, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea30dcd6-f499-4062-91a9-6da3efc4fa07 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d9e395-835c-4803-a5cc-711754ad2412 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross-modal graph matching network for image- text retrieval.TOMM, 18(4):1–23, 2022
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2a6fbbc-57a4-43f6-a773-5ec29fdd347f · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Reproducible scal- ing laws for contrastive language-image learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8701963-6001-43a5-a0c7-04f1a90cf92e · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Sim- ilarity reasoning and filtration for image-text matching
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 351997d4-cfa6-4159-8360-6514f64e8570 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fleet, Jamie Ryan Kiros, and Sanja Fidler
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 095d7716-b2e3-4634-bb8a-5ea1bdb37e31 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning semantic relationship among instances for image- text matching
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e664c977-85da-4a1a-a165-cdcda54adc90 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretraining
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dee985aa-e72f-41c5-9cae-5522cf82d76d · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross-modal semantic enhanced interaction for image-sentence retrieval
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b42cf75e-bdf4-481f-bd3a-d4dd5c018bdc · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Hiclip: Contrastive language-image pre- training with hierarchy-aware attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1934b8d-98d5-4a57-bcd2-32ae4b06f868 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching From im- ages to textual prompts: Zero-shot visual question answering with frozen large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5f97260-9444-42ce-a1fb-844f1d922875 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Visual program- ming: Compositional visual reasoning without training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3231f3df-5b52-4677-a74f-d088b652f9e4 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Structure-clip: Towards scene graph knowledge to enhance multi-modal structured representa- tions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2309f4f4-fa2c-41ed-b325-866efec1c4a0 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fineclip: Self-distilled region-based clip for better fine-grained under- standing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c33dae3-0e7e-4537-b57b-8138ce59554b · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Knowledge-aware prompt tun- ing for generalizable vision-language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ae628f9-cb49-4137-9e6e-ec1174499752 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Maple: Multi-modal prompt learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37eeb6a9-2b61-4b18-b572-09e88fe4e44e · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Stacked cross attention for image-text matching
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66a2902f-4696-492e-a048-c47e39ade055 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Action-aware em- bedding enhancement for image-text retrieval
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e7a3d5f-46e1-487c-a0af-9296832a10e9 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning background prompts to dis- cover implicit knowledge for open vocabulary object detec- tion
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54680ab0-30df-4504-99de-ff27fa414cb2 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross- modal alternating learning with task-aware representations for continual learning.TMM, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1abbc369-b536-41d1-8579-bc551a8d61ea · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Image-text bidirectional learning network based cross-modal retrieval
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85adc5a4-ee2a-49da-af69-1ad9ebcfc1f7 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning customized visual models with retrieval-augmented knowledge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cce8655-cbc3-4550-86c7-124630b1c125 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Multi-modal attribute prompting for vision-language models.TCSVT, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71ccba72-23ac-486c-93f6-c97ae5125f4b · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fine-grained visual– text prompt-driven self-training for open-vocabulary object detection.TNNLS, pages 1–11, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 216fa119-48a1-4cea-943f-eef01ebd306f · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Visual classification via description from large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7bfca97-fb79-4275-8ad4-e8578c7cb3fd · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching SCHEMA: state changes matter for procedure planning in instructional videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 492369ea-3c73-4552-ae31-4dc7e546eac7 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Teaching clip to count to ten
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a958523-11f2-4aac-9de5-3c11cf5486b2 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fine-grained image-text matching by cross-modal hard aligning network
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d500885-10b2-4dfe-83df-900a3002a46d · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Dynamic modality interaction modeling for image-text retrieval
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 784facf2-5037-4c1f-8ffa-229826f8e16a · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning transferable visual models from natural language supervision
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df139403-5b87-4590-aef3-278e5f8cfb7a · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Vlc-bert: Visual question answering with contextualized commonsense knowledge
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fec0c54-161a-4dde-ad96-fec20fb6ce54 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Language models are causal knowledge ex- tractors for zero-shot video question answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97d650ee-0f8e-49fa-b45a-415ff64f2dc8 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 306a3b8e-e638-4a3f-873f-ffc8c008123b · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Compound text-guided prompt tuning via image-adaptive cues
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 683ec262-ffed-4b1b-b9f8-1ba75477c251 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Multi-granularity cross-modal align- ment for generalized medical visual representation learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7505452-b647-4910-8541-9e75740df224 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Consensus-aware visual-semantic embedding for image- text matching
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 106832fd-dd91-4d87-98f0-f50eb4dd2bb8 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Vilt- clip: Video and language tuning clip with multimodal prompt learning and scenario-guided optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce17a852-6b73-4383-837c-9ee1589a4735 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Position-guided text prompt for vision-language pre- training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e85078f-a7d3-4f7d-9755-7bc0e018c68a · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching ActionCLIP: A New Paradigm for Video Action Recognition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e788e16-0286-4999-b93e-6f5467d83c86 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross-modal scene graph matching for relationship-aware image-text retrieval
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69cc5f4b-da85-4fc3-9195-af6880442ea6 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Diffusion feedback helps clip see better
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6b2526c-d082-4833-896e-56fb0d19f469 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Balance act: Mitigat- ing hubness in cross-modal retrieval with query and gallery banks
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a085f20a-3d6a-4c66-95a4-c2e1911b43d8 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning hierarchical prompt with structured linguistic knowledge for vision-language models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b308134-caa9-4bed-893e-e343532fe2b1 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Multi-view inter-modality representation with progressive fusion for image-text matching.Neurocomputing, 535:1–12, 2023
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fcd561d1-4d2e-4a45-9c47-0713da403415 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Saco loss: Sample-wise affinity con- sistency for vision-language pre-training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27802d92-26da-4d5c-a7a2-ada570d0ec48 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Visual- language prompt tuning with knowledge-guided context op- timization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c60afda-21e5-4646-838d-6aa3ccc5a0ee · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching FILIP: fine-grained interactive language-image pre-training
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffffa9ae-697c-4519-8c11-e33500c03a49 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.TACL, 2:67–78, 2014
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0c6aea0-a122-448b-a748-51b057bff30b · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Dual-path convolutional image-text embeddings with instance loss.TOMM, 16(2): 1–23, 2020
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dde1ff41-5a67-4728-82f1-f784c595b202 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Large Language Models are Good Prompt Learners for Low-Shot Image Classification
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b13f4e3d-710c-4d46-b2bd-61e6ae75c0fe · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Regionclip: Region-based language- image pretraining
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15d59c97-d7a3-4b17-b0ac-37ee4fb5d7bf · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Relclip: Adapting language- image pretraining for visual relationship detection via rela- tional contrastive learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 757f26c7-3387-439a-ba94-4b7b9d7a6fe0 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching w/o ac- tion knowledge
Reference 336
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d955b46-826e-4e4d-821c-5930fc88c276 · outbound
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Datasets Details Flickr30K[50] dataset contains 31,000 images collected from the Flickr website
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99e15fe9-6af7-4601-9aff-15c5c1ba243f · inbound
Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.