Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:40:10.108333Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2605.25479.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T22:40:10.108333Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 992e04dd-a6b2-41cf-8da8-5358ca445e37 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models On the Opportunities and Risks of Foundation Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0326633f-c968-4b36-a692-9ec17e49978d · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning transferable visual models from natural language supervision,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df8e4d00-85a0-4214-b987-d4cbf8e8b722 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Deep residual learning for image recognition,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01d33bd-a95a-4853-ab78-a7a0e1c130b2 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da35c657-7a4b-443c-9eeb-b42f3f8023bf · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models arXiv preprint arXiv:2511.21772 , year =
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9590a9a0-da35-49a6-af75-364e296a19f5 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Cp-clip: Core- periphery feature alignment clip for zero-shot medical image analysis,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96f39f0-b6a7-4821-b4f4-41845e6e0367 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Vpl: Visual proxy learning framework for zero-shot medical image diagnosis,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69da6986-dd80-449f-b039-214e0afdb685 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Medclip: Contrastive learning from unpaired medical images and text,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea79b72a-9672-4306-a769-6e3b63c12818 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Pros: Prompting-to-simulate generalized knowledge for universal cross-domain retrieval,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 969637b0-7bf3-4095-ba78-55a833cff838 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Depro: Domain ensemble using decoupled prompts for universal cross-domain retrieval,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5c23bd-f9bd-449a-aa6d-c135d760bb0e · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e4813c8-df96-4405-b3ce-000cb89a74f5 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Building a multi-modal spatiotemporal expert for zero-shot action recognition with clip,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e48ed2-d29a-488c-98ff-cea4fcbaf4a5 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Leveraging temporal contextu- alization for video action recognition,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c421381-d968-4398-83b6-1f90b8ada3c1 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa52433-aac7-47e0-bcf1-ab042fc13da8 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning to prompt for vision- language models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776412d6-ced3-4dfa-bd60-f038ea57b9ec · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Conditional prompt learning for vision-language models,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd06bec-ab04-4f22-931e-382546746f42 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Visual-language prompt tuning with knowledge-guided context optimization,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667653dc-a60b-4db9-8c62-f584b430ad86 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Self-regulating prompts: Foundational model adaptation without forgetting,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf2aba9-ac7e-4592-8a79-764c4ed1aae8 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Maple: Multi-modal prompt learning,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c56974-9c0d-4f0a-a209-2f77b7676fd6 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Mmrl: Multi-modal representation learning for vision-language models,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45e2041-2883-462b-ba96-11665a76d92b · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4e49347-c519-4420-bec0-5335364eef30 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Tip-adapter: Training-free adaption of clip for few-shot classification,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e48f86-a9f7-4bab-9625-855813514b7f · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Clip-adapter: Better vision-language models with feature adapters,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be2556d-ed72-4e07-b8c2-428effb1ffd5 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Mma: Multi-modal adapter for vision-language models,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143b27d9-b305-4486-8d0d-9395b9317ca3 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Scaling & shifting your features: A new baseline for efficient model tuning,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9500fe27-4700-422b-9ffb-6395f60786c7 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Multi-modal interactive agent layer for few-shot universal cross-domain retrieval and beyond,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bdd98b8-33db-4113-8150-bc6d4bc125db · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Vilt: Vision-and-language transformer without convolution or region supervision,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee89f22-6100-4c4f-a9e3-8c750374d758 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Align before fuse: Vision and language representation learning with momentum distillation,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fc24b0a-1085-4894-b559-21291b11214b · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Vlmo: Unified vision-language pre- 12 training with mixture-of-modality-experts,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8ef49b5-f256-458c-803a-e1f016c51fc7 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918bb84d-aa35-4c01-ac7c-2490ad2e7b3b · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Visual instruction tuning,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce677725-b034-4345-b9c7-fa2419d11132 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Parameter-efficient transfer learning for nlp,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71f8aa0-7ab6-4d4b-8fee-e2c6b9fdbea2 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Lora: Low-rank adaptation of large language models,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 425f9359-f853-4436-82e9-33a02a482345 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Bitfit: Simple parameter- efficient fine-tuning for transformer-based masked language-models,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b473cbc-a176-46bb-a819-62e483f4938d · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning with enriched inductive biases for vision-language models,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8df2e91-4aa0-4156-bdcb-9d7bdfbdf19a · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Not all features matter: Enhancing few-shot CLIP with adaptive prior refinement,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89facad9-e895-4d0b-8d3e-bf529d32b59d · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Task residual for tuning vision-language models,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f16d64-8ac6-4804-8dd3-e1fc83ecf811 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aeb8e29-5cfe-420a-8794-f1ae7b2e4fb3 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Attention is all you need,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e726c4f-339a-4614-92ee-5bbd4926af4f · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Imagenet: A large-scale hierarchical image database,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97044348-37d3-4c02-b202-f0184dc999ab · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dbd24bb-d811-4cde-9136-1ca21837bd01 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Cats and dogs,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b61579f-c24b-4ac7-8b97-960bf17f1d20 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models 3d object representations for fine-grained categorization,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297235c9-f9ee-498f-892d-e6222d7d72e9 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Automated flower classification over a large number of classes,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68427a17-e17e-4f05-9d07-2fa07e4aa935 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Food-101–mining discriminative components with random forests,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a36021cc-78b3-4fbd-b227-1ff06103faa5 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Fine-Grained Visual Classification of Aircraft
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e307de1-8908-43ea-bb42-6567c367a4a4 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Sun database: Large-scale scene recognition from abbey to zoo,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd93f93-afd6-4296-9950-8474a5971ef4 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models A dataset of 101 human action classes from videos in the wild,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67c31fa7-9e9a-4a62-a735-ebe3044f8d7b · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Describing textures in the wild,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b3f018-d004-4c2e-8f60-5f2c907040d9 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7231f43e-3f34-44f2-9c5c-2ec0440fc988 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Do imagenet classifiers generalize to imagenet?
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4039c3-9bc3-4b15-a7c0-04526ebec405 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning robust global representations by penalizing local predictive power,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3abb92-4ba5-4039-8d86-196336e1b920 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Natural adversarial examples,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f0708fb-a9dd-465e-adc4-ce6e83aa0c4d · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models The many faces of robustness: A critical analysis of out-of-distribution generalization,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77935810-de96-4a9a-af3c-9c8fbe8bfd9a · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Moment matching for multi-source domain adaptation,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b884223-0559-4206-a4f4-da0dbef047a6 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models The sketchy database: learning to retrieve badly drawn bunnies,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98fca510-8bfe-41c0-9329-0a04adad9a81 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Deep sketch hashing: Fast free-hand sketch-based image retrieval,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a97fd1-cb6b-4d14-958c-851981d987f6 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models How do humans sketch objects?
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e52f92-a3d4-42b9-a688-3db858ae88e9 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Sketchnet: Sketch classification with web images,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d16e8f-d448-4d2f-b631-d9d9f437d7a7 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Tcp: Textual-based class-aware prompt tuning for visual-language model,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10695c7a-8441-4fb1-a90f-135e14042b69 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Divergence-enhanced knowledge-guided context optimization for visual-language prompt tun- ing,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11581c72-621b-4718-905a-054609fc1c18 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Bi-modality individual- aware prompt tuning for visual-language model,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 467015b1-0524-463c-a7d1-c9d3c6ab69cc · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Frequency-based comprehensive prompt learning for vision-language models,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf539b8-0013-4ede-8433-6f91f6322b66 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Hierarchical cross- modal prompt learning for vision-language models,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada916bc-52bc-4f23-9856-33ffeda88a51 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Promptkd: Unsupervised prompt distillation for vision-language mod- els,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426fef9a-c597-4e1a-8c49-f946d3ddf1b8 · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Adapt- former: Adapting vision transformers for scalable visual recognition,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7133d924-1d48-4059-9803-1da5da9c7cfe · outbound
MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Visual prompt tuning,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.