Pith. sign in

Paper Citation Record · LEDGER

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2605.25479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25479 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:40:10.108333Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 992e04dd-a6b2-41cf-8da8-5358ca445e37 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models On the Opportunities and Risks of Foundation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.463778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:7a78a35819ec88b5b7c4650bd1357e8dede5c2d9ef86e5aa874bde0e09ad66ec

Observation 0326633f-c968-4b36-a692-9ec17e49978d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:a9d49514cd8cd3f00865cafb07d2f93bd5aca0ae31a2db79eab57de7dcdfe296

Observation df8e4d00-85a0-4214-b987-d4cbf8e8b722 · outbound

This paper cites Deep residual learning for image recognition,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Deep residual learning for image recognition,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:7e34c6dd7422613f796dcb57e3e5c1ce69f2533eb75dae720c1d0d3a03a02d1d

Observation b01d33bd-a95a-4853-ab78-a7a0e1c130b2 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:aa98630628f04ae6ab6f8ee9e60c20081246a6be2db4592eb73af4eeee4eac0f

Observation da35c657-7a4b-443c-9eeb-b42f3f8023bf · outbound

This paper cites arXiv preprint arXiv:2511.21772 , year =.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models arXiv preprint arXiv:2511.21772 , year =

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:44:01.457361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:97d5daf4d9b33508093db38700c6b844ce23dd7a9a3516e39c147ec06e272a64

Observation 9590a9a0-da35-49a6-af75-364e296a19f5 · outbound

This paper cites Cp-clip: Core- periphery feature alignment clip for zero-shot medical image analysis,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Cp-clip: Core- periphery feature alignment clip for zero-shot medical image analysis,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:f7aeb863402482ef7d97fda5c2b229f980f2c330a522e58d985db6f59397bb11

Observation a96f39f0-b6a7-4821-b4f4-41845e6e0367 · outbound

This paper cites Vpl: Visual proxy learning framework for zero-shot medical image diagnosis,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Vpl: Visual proxy learning framework for zero-shot medical image diagnosis,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:947f37527f6fdd240b0255e0dac4ffb31cba31e4b7a0bf4c6aa32500999dc469

Observation 69da6986-dd80-449f-b039-214e0afdb685 · outbound

This paper cites Medclip: Contrastive learning from unpaired medical images and text,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Medclip: Contrastive learning from unpaired medical images and text,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:96afefd787c3a26600f887778f006bc5b6f5a32a58441e1fb9be2dc2ae23dbf1

Observation ea79b72a-9672-4306-a769-6e3b63c12818 · outbound

This paper cites Pros: Prompting-to-simulate generalized knowledge for universal cross-domain retrieval,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Pros: Prompting-to-simulate generalized knowledge for universal cross-domain retrieval,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:e5cc568628eadc4062acfb5c23741f547f06e6f1687072895571f19e5d47cf0e

Observation 969637b0-7bf3-4095-ba78-55a833cff838 · outbound

This paper cites Depro: Domain ensemble using decoupled prompts for universal cross-domain retrieval,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Depro: Domain ensemble using decoupled prompts for universal cross-domain retrieval,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:85b941d2107604ffd6c95850fff359d77e814007ddfeb7ed1c8035019e2ecb9d

Observation cf5c23bd-f9bd-449a-aa6d-c135d760bb0e · outbound

This paper cites Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:44:01.450876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:289c871cc9046cb3ccdf8615b499943e526efc9b31fcff1d41d16bb2157ee3e1

Observation 0e4813c8-df96-4405-b3ce-000cb89a74f5 · outbound

This paper cites Building a multi-modal spatiotemporal expert for zero-shot action recognition with clip,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Building a multi-modal spatiotemporal expert for zero-shot action recognition with clip,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:7827a6f191474036ea2c000ec883951710cfef3f77003af5b415afe3ff8612d0

Observation b2e48ed2-d29a-488c-98ff-cea4fcbaf4a5 · outbound

This paper cites Leveraging temporal contextu- alization for video action recognition,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Leveraging temporal contextu- alization for video action recognition,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:57263549216fa6863aa03478f24b156718940e9061ea98cfbf775ec6fe508663

Observation 5c421381-d968-4398-83b6-1f90b8ada3c1 · outbound

This paper cites Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:fbcfa6d620200522135d08db2ea8740bca4783c9ad2d144ec71efabff41a8337

Observation 7fa52433-aac7-47e0-bcf1-ab042fc13da8 · outbound

This paper cites Learning to prompt for vision- language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning to prompt for vision- language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:6e68cc5a3a9549cc32c3509fe91b36e4fa46e89bdeb9d1931eaf57b256e99b23

Observation 776412d6-ced3-4dfa-bd60-f038ea57b9ec · outbound

This paper cites Conditional prompt learning for vision-language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Conditional prompt learning for vision-language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:d6b702e915d9219488e01fed156f060c204c5a835aed1d31996460e201d45f1f

Observation fbd06bec-ab04-4f22-931e-382546746f42 · outbound

This paper cites Visual-language prompt tuning with knowledge-guided context optimization,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Visual-language prompt tuning with knowledge-guided context optimization,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:06f840eac75ec15d5f82c4776caed2a3fc51d624bb9f343a6f7bed199ff44dd6

Observation 667653dc-a60b-4db9-8c62-f584b430ad86 · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Self-regulating prompts: Foundational model adaptation without forgetting,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:fcacf6e58bc820cec4c200fd21bd4690135530c422490453b923ed4ce58d3c55

Observation 9cf2aba9-ac7e-4592-8a79-764c4ed1aae8 · outbound

This paper cites Maple: Multi-modal prompt learning,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Maple: Multi-modal prompt learning,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:c01776189eb5330068763b850b9b951d63cf822560e760ddd7ac9fb576430f9a

Observation d6c56974-9c0d-4f0a-a209-2f77b7676fd6 · outbound

This paper cites Mmrl: Multi-modal representation learning for vision-language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Mmrl: Multi-modal representation learning for vision-language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:403b8991a90659ec5bb10b61776fc0431f290c90db7f4d0295b6b3fe6bb9fa23

Observation b45e2041-2883-462b-ba96-11665a76d92b · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.454414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:45d405900dba8e3ee6bd773ad6677db67beaea6ca1d247e850c990096eef6c28

Observation a4e49347-c519-4420-bec0-5335364eef30 · outbound

This paper cites Tip-adapter: Training-free adaption of clip for few-shot classification,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Tip-adapter: Training-free adaption of clip for few-shot classification,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:16d8f18729546108eda787a90e574b6479e8dd0c8c6bfeb7a502b68d9b0d537d

Observation d8e48f86-a9f7-4bab-9625-855813514b7f · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Clip-adapter: Better vision-language models with feature adapters,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:40b5310e4a3b201333d560b854ce60b0f14769eb6f0955f4e8e20572f1f1e95d

Observation 7be2556d-ed72-4e07-b8c2-428effb1ffd5 · outbound

This paper cites Mma: Multi-modal adapter for vision-language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Mma: Multi-modal adapter for vision-language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:80ab21316671a19922e211fa73661f5e09e73b841ff94fc5d78f1174bd338f55

Observation 143b27d9-b305-4486-8d0d-9395b9317ca3 · outbound

This paper cites Scaling & shifting your features: A new baseline for efficient model tuning,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Scaling & shifting your features: A new baseline for efficient model tuning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:d4546296672cc9ac1f35bea598032f329fefe0f1d8f2a87a7b0f1884aabc655d

Observation 9500fe27-4700-422b-9ffb-6395f60786c7 · outbound

This paper cites Multi-modal interactive agent layer for few-shot universal cross-domain retrieval and beyond,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Multi-modal interactive agent layer for few-shot universal cross-domain retrieval and beyond,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:254bbaec12dc7e12076dec8572de912e7aba0d9f499a6d82e9a0ce4c443e1364

Observation 2bdd98b8-33db-4113-8150-bc6d4bc125db · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:5d6f11e36fbcf6000d1ae4555317be4af0dde60d56d76acd52796db0bf365919

Observation bee89f22-6100-4c4f-a9e3-8c750374d758 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Align before fuse: Vision and language representation learning with momentum distillation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:0f50f9720582c243b81bf13f38a0d5a21a56b78ef9dbdcb19388992cac9d2673

Observation 3fc24b0a-1085-4894-b559-21291b11214b · outbound

This paper cites Vlmo: Unified vision-language pre- 12 training with mixture-of-modality-experts,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Vlmo: Unified vision-language pre- 12 training with mixture-of-modality-experts,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:e579cf37a96500c07eaa1100d7550e45acf89860dd3093452c3f065e669c182f

Observation d8ef49b5-f256-458c-803a-e1f016c51fc7 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:c84b087baa17e811752f942d98b8885276fc366aa9b6667dda26f9746df7c3e7

Observation 918bb84d-aa35-4c01-ac7c-2490ad2e7b3b · outbound

This paper cites Visual instruction tuning,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Visual instruction tuning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:eaab174f98dd1bc7cd276b28d8fd09ddce723ac3f14da9cf571f85df2234dfc5

Observation ce677725-b034-4345-b9c7-fa2419d11132 · outbound

This paper cites Parameter-efficient transfer learning for nlp,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Parameter-efficient transfer learning for nlp,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:30da494e0e588e47ac3a664e9dd78ea2a70af151a68505d566d28375e8f68739

Observation d71f8aa0-7ab6-4d4b-8fee-e2c6b9fdbea2 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Lora: Low-rank adaptation of large language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:87925708db743b2fb7a558a068aa8d93a7c47ecdb043e855cf2887ede6b6b956

Observation 425f9359-f853-4436-82e9-33a02a482345 · outbound

This paper cites Bitfit: Simple parameter- efficient fine-tuning for transformer-based masked language-models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Bitfit: Simple parameter- efficient fine-tuning for transformer-based masked language-models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:96d1b7f6972b2b2fd30794eefb65f78f77f30de0a1b8e8433e8b0c2ef5fc872b

Observation 1b473cbc-a176-46bb-a819-62e483f4938d · outbound

This paper cites Learning with enriched inductive biases for vision-language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning with enriched inductive biases for vision-language models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:720e4b1efcb8ac708070a67f4188812d464efbabfa757ecf331636c82fcbb576

Observation e8df2e91-4aa0-4156-bdcb-9d7bdfbdf19a · outbound

This paper cites Not all features matter: Enhancing few-shot CLIP with adaptive prior refinement,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Not all features matter: Enhancing few-shot CLIP with adaptive prior refinement,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:daeb086bee7d89a346d6699dc8ba1b3a1cb6bbf399cfc8d4b8fdddcd71b158ee

Observation 89facad9-e895-4d0b-8d3e-bf529d32b59d · outbound

This paper cites Task residual for tuning vision-language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Task residual for tuning vision-language models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:3f601add236deca707fb58aa5280c2b2d317d8eb3d6c7985e3c0ff6c2b79aad6

Observation f8f16d64-8ac6-4804-8dd3-e1fc83ecf811 · outbound

This paper cites Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:4456d728d5d1557f7bcd019fa46cc6dc09ed3475abca56142710dc237bded167

Observation 4aeb8e29-5cfe-420a-8794-f1ae7b2e4fb3 · outbound

This paper cites Attention is all you need,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Attention is all you need,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:26a7692f2102b08b43060508a62396535636e2b68700d3be5efa6cc4e2175873

Observation 6e726c4f-339a-4614-92ee-5bbd4926af4f · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Imagenet: A large-scale hierarchical image database,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:d6422a6f4ae7f16bb06fea37677061f0919ec7101041018aba16e7ac1147c0d9

Observation 97044348-37d3-4c02-b202-f0184dc999ab · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:dde9d2d8919793c23d1927e761741d0d27bdba9180ee94519f0f73b7fdf20c40

Observation 2dbd24bb-d811-4cde-9136-1ca21837bd01 · outbound

This paper cites Cats and dogs,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Cats and dogs,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:22094febee6e32f623ebf166162ebd53da7eba710462b6ea7e618ba73e3ba32f

Observation 6b61579f-c24b-4ac7-8b97-960bf17f1d20 · outbound

This paper cites 3d object representations for fine-grained categorization,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models 3d object representations for fine-grained categorization,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:d52f26c7c58f285f547fa4e4ac13fb0199da27ce217958d93bf4e735380a6ed7

Observation 297235c9-f9ee-498f-892d-e6222d7d72e9 · outbound

This paper cites Automated flower classification over a large number of classes,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Automated flower classification over a large number of classes,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:906dceb85579e862c31984a17bced25812d0586ff2d9c552aa5a07c291c69a5b

Observation 68427a17-e17e-4f05-9d07-2fa07e4aa935 · outbound

This paper cites Food-101–mining discriminative components with random forests,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Food-101–mining discriminative components with random forests,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:86c0eda05e39b269dba75d56a8bc5db211748ac4514cd82ecda7b953b23c66eb

Observation a36021cc-78b3-4fbd-b227-1ff06103faa5 · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Fine-Grained Visual Classification of Aircraft

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:44:01.447672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:f3d9c87fa980988ef7f8fdf2158fff628114a37986e8835b604f8e9bf712da05

Observation 0e307de1-8908-43ea-bb42-6567c367a4a4 · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Sun database: Large-scale scene recognition from abbey to zoo,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:02efed666267138414defb47ab3b028c90bc72e8aa70946d26193f3be99902f4

Observation 3fd93f93-afd6-4296-9950-8474a5971ef4 · outbound

This paper cites A dataset of 101 human action classes from videos in the wild,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models A dataset of 101 human action classes from videos in the wild,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:51afe897fa022fd94d85eea82b4ee489c2ff5520f3955b7d964daa314344b788

Observation 67c31fa7-9e9a-4a62-a735-ebe3044f8d7b · outbound

This paper cites Describing textures in the wild,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Describing textures in the wild,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:5b74a2100c842a476da8e3737c222755355008044da0b5ee1fb41d3b4741acb3

Observation 63b3f018-d004-4c2e-8f60-5f2c907040d9 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:2123019b49f59c150c9ae6703c1b59349cf495eba5637371f537d933e8569647

Observation 7231f43e-3f34-44f2-9c5c-2ec0440fc988 · outbound

This paper cites Do imagenet classifiers generalize to imagenet?.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Do imagenet classifiers generalize to imagenet?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:e2d6015a0cfec76044917fcfb556227c01b55263f5bf6ce254a8ee0f83dcc7e1

Observation ed4039c3-9bc3-4b15-a7c0-04526ebec405 · outbound

This paper cites Learning robust global representations by penalizing local predictive power,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Learning robust global representations by penalizing local predictive power,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:55bb43dc57bf39688ef66a82fd7e3f730fad2a8a7341a6aa31de388b7fff93d9

Observation 1f3abb92-4ba5-4039-8d86-196336e1b920 · outbound

This paper cites Natural adversarial examples,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Natural adversarial examples,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:6785a3b0910c76257534295a4c7d6748a8121faac96c33f692c816166ce48961

Observation 5f0708fb-a9dd-465e-adc4-ce6e83aa0c4d · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models The many faces of robustness: A critical analysis of out-of-distribution generalization,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:c68978349fd5bf3679c968c683b06dab3545a271b5e79b7e2038bf568db5e57f

Observation 77935810-de96-4a9a-af3c-9c8fbe8bfd9a · outbound

This paper cites Moment matching for multi-source domain adaptation,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Moment matching for multi-source domain adaptation,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:c8ffc332739799777120385b5c5b27ab0da884d1b3c204727c95906ca8bb0084

Observation 3b884223-0559-4206-a4f4-da0dbef047a6 · outbound

This paper cites The sketchy database: learning to retrieve badly drawn bunnies,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models The sketchy database: learning to retrieve badly drawn bunnies,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:af88e99fdf10439e13e43ef72440e7d73b5ae00e265242f63001f4afec8dcca8

Observation 98fca510-8bfe-41c0-9329-0a04adad9a81 · outbound

This paper cites Deep sketch hashing: Fast free-hand sketch-based image retrieval,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Deep sketch hashing: Fast free-hand sketch-based image retrieval,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:daf695651da79b46a0d98caba672a637b95f1b2d71bde1793ac1cc164c89d2f6

Observation 61a97fd1-cb6b-4d14-958c-851981d987f6 · outbound

This paper cites How do humans sketch objects?.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models How do humans sketch objects?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:72c75ee55f7936156f3c4135903763d1ada85d7b3a39294c2c93613a25133586

Observation a5e52f92-a3d4-42b9-a688-3db858ae88e9 · outbound

This paper cites Sketchnet: Sketch classification with web images,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Sketchnet: Sketch classification with web images,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:d305ccc652d37fd0c5a0d81e911d173bbb1d49a7ba442b6a1b47287403a1123d

Observation b0d16e8f-d448-4d2f-b631-d9d9f437d7a7 · outbound

This paper cites Tcp: Textual-based class-aware prompt tuning for visual-language model,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Tcp: Textual-based class-aware prompt tuning for visual-language model,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:fa5f55e7349cbe7218a4926ae96a17c2cbe18cb043071cb50e666071456b8276

Observation 10695c7a-8441-4fb1-a90f-135e14042b69 · outbound

This paper cites Divergence-enhanced knowledge-guided context optimization for visual-language prompt tun- ing,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Divergence-enhanced knowledge-guided context optimization for visual-language prompt tun- ing,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:cb60b5d6c8212f8b9d1425e1e0dc6f02fd6be4612dcfaf004f7f68d06c964cf8

Observation 11581c72-621b-4718-905a-054609fc1c18 · outbound

This paper cites Bi-modality individual- aware prompt tuning for visual-language model,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Bi-modality individual- aware prompt tuning for visual-language model,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:6d756a476dbdd4b5cabb598bf3f39f5294a71c93b010c9217c7ded80f096eaff

Observation 467015b1-0524-463c-a7d1-c9d3c6ab69cc · outbound

This paper cites Frequency-based comprehensive prompt learning for vision-language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Frequency-based comprehensive prompt learning for vision-language models,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:8378bc77cbfbb8a3514d844e17c904d9772d4a57c27f9f97a2f9ad45eae47de3

Observation 4cf539b8-0013-4ede-8433-6f91f6322b66 · outbound

This paper cites Hierarchical cross- modal prompt learning for vision-language models,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Hierarchical cross- modal prompt learning for vision-language models,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:ad845144a289bea58b6711999786dfdc9e228b9092d03f114d1ccb49cad414b1

Observation ada916bc-52bc-4f23-9856-33ffeda88a51 · outbound

This paper cites Promptkd: Unsupervised prompt distillation for vision-language mod- els,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Promptkd: Unsupervised prompt distillation for vision-language mod- els,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:1f008d80ce33c66e08f3d643150f0938316e419c312517179fedaa27fbed786e

Observation 426fef9a-c597-4e1a-8c49-f946d3ddf1b8 · outbound

This paper cites Adapt- former: Adapting vision transformers for scalable visual recognition,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Adapt- former: Adapting vision transformers for scalable visual recognition,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:fa2521d98b3b0393b9b639c1ad418d220ab18bec06822be00913835e3c4a14ab

Observation 7133d924-1d48-4059-9803-1da5da9c7cfe · outbound

This paper cites Visual prompt tuning,.

MAIL++: Multi-Modal Bi-directional Agent Layer for Vision-Language Models Visual prompt tuning,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-29T22:40:10.108333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:40:10.108333Z digest=sha256:f589de009db1ebd0ac78eb1d427bcb659efce6370eca422e7abbaf8242f220b3

Pith citing papers

No inbound Pith citation observations are available.