Pith. sign in

Paper Citation Record · LEDGER

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2507.08410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08410 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:27:47.245370Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact3
  • verified fuzzy43
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9156d484-00be-4473-948a-c2cad2f6b3c9 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.747091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.747091Z digest=sha256:d23a03e87622e2864024976a4a0beb12aa5e62a4da4ec77000ba4400d39a70d5

Observation 2f4deb52-f0f0-4b76-8427-1491df391ee8 · outbound

This paper cites Why are Visually-Grounded Language Models Bad at Image Classification?.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.828449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.828449Z digest=sha256:817aa0dc85ee60d3474d4bc088cce4647a548156e307becad032423bf6054864

Observation 13029e43-a1b6-49e4-bafc-42fb665b7bfb · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Eva: Exploring the limits of masked visual representation learning at scale,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.858025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.858025Z digest=sha256:0de413aea1e14840b98138ee9ceafe3437baf86835dda44f4db68e4492acc579

Observation 39aa69af-70b4-4a4f-8183-f1198ec0dbeb · outbound

This paper cites Conditional prompt learning for vision-language models,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Conditional prompt learning for vision-language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.862663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.862663Z digest=sha256:03a738ddecb8171e89430badcf1e0bc880ea7ac006bc629b059286546c91b16b

Observation 46276c91-e114-4954-aa13-5a3a05cb3661 · outbound

This paper cites Learning to prompt for vision-language models,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Learning to prompt for vision-language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.867601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.867601Z digest=sha256:a3c09655b4e525748f812663c9b85e3505e80e123ab3dac7eef95f581de634c7

Observation e285e8a3-1e6c-46ea-beb8-b17c138e9631 · outbound

This paper cites Maple: Multi-modal prompt learning,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Maple: Multi-modal prompt learning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.871905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.871905Z digest=sha256:d99ee455de208e5494f411091707db899a31c405ecdeff9c4bc4fdcb25c9089b

Observation a6fc409e-cd32-491d-b6b6-0a58406c2ab0 · outbound

This paper cites Tcp: Textual-based class-aware prompt tuning for visual-language model,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Tcp: Textual-based class-aware prompt tuning for visual-language model,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.876620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.876620Z digest=sha256:31ee4188f37b589e9437b888ba88cc5cea10038ace9d1d7c4cfe2cc41d02c6e5

Observation d1b1b93a-719a-4ddf-8302-3d48af181791 · outbound

This paper cites Dept: Decoupled prompt tuning,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Dept: Decoupled prompt tuning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:53.308142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.881369Z digest=sha256:3f8a8525c921ef1ec5b621087e5679c029d10ef0b554750a9409e8187cf854c1

Observation 1e11ad71-9bdb-44c6-96e9-9c28f9822ee4 · outbound

This paper cites PromptKD: Unsupervised Prompt Distillation for Vision-Language Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models PromptKD: Unsupervised Prompt Distillation for Vision-Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:27:47.628086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.885780Z digest=sha256:851f23f7bddc2ac5cd9b79e9879d0b21abedade164bcd8aeb9c3c8e3e75a2f91

Observation ff40152d-8d67-4828-ba13-342a9de8c0cb · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Self-regulating prompts: Foundational model adaptation without forgetting,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.890805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.890805Z digest=sha256:045c5400622975995f8f85331c291da6bd79d2e1754d6be26e5597c47ecccff6

Observation 6bf813c6-152a-4bcf-9010-af13b40d0145 · outbound

This paper cites Consistency-guided Prompt Learning for Vision-Language Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Consistency-guided Prompt Learning for Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.895300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.895300Z digest=sha256:6dcaca72f26de94cacd2d392668c6bac242d80f60c9130879ed55c58a34946fa

Observation beabbd0c-6a93-456d-884e-253b41e88a75 · outbound

This paper cites Large Language Models are Good Prompt Learners for Low-Shot Image Classification.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Large Language Models are Good Prompt Learners for Low-Shot Image Classification

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:27:47.588821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.900106Z digest=sha256:794a61fa1f6087e82249367ef8892d10b5b9496531c772cd35bd1dc2bae8ee24

Observation b0972253-29f7-495d-b171-ee7e2656d84a · outbound

This paper cites Improved zero-shot classification by adapting vlms with text descriptions,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Improved zero-shot classification by adapting vlms with text descriptions,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:53.132833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.904781Z digest=sha256:974d8a23b769b56bad2808666565bd87739fb1313d8eecac3de746eb697ae79d

Observation b51b309f-f8ee-4a08-aa8c-88b0b494d25e · outbound

This paper cites Bilateral adaptive cross-modal fusion prompt learning for clip,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Bilateral adaptive cross-modal fusion prompt learning for clip,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:53.023998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.909676Z digest=sha256:556bed661309401db7be2d3d0b9d03f8911e4d6553e639dad147cbb4c5dd0c86

Observation 238e3490-284d-4bb2-961f-47d8d62720ee · outbound

This paper cites Unified Vision and Language Prompt Learning.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Unified Vision and Language Prompt Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.914058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.914058Z digest=sha256:24e06fdf40b0b01e8e2bca0c594cfc03114e441dc2249471c3c873d3c7694df2

Observation df39cd68-4808-4b93-84a3-376673743d68 · outbound

This paper cites Visual prompt tuning,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Visual prompt tuning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.918620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.918620Z digest=sha256:606744e18372b5a6cd70e3b4b3e8e3d0fe0dd578d305dd5e1bea787979c41110

Observation 8de8356e-06a5-4660-bfaa-dcb20f95c1c9 · outbound

This paper cites Enhancing clip with gpt-4: Harnessing visual descriptions as prompts,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Enhancing clip with gpt-4: Harnessing visual descriptions as prompts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.914404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.922850Z digest=sha256:7f4bbb506294f94bab78165b2dc8fa838dd457687f940308176e71aba01e7fe2

Observation 3bd84a38-a290-439d-98ea-de7097bf09bf · outbound

This paper cites Prompt distribution learning,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Prompt distribution learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.819718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.927549Z digest=sha256:aa14eb4d497125d0140a1b1988ee2e274759391d0e166908a4f01b70fc5eb130

Observation 746ad6f6-bf5f-4710-a27b-837d46252273 · outbound

This paper cites Prompt-aligned gradient for prompt tuning,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Prompt-aligned gradient for prompt tuning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.688397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.931748Z digest=sha256:8346914f91d0c04070f466634d367d2e8527ea0cbe89728140ee6ce960182a99

Observation ce721eb8-229b-4418-bd7f-ddf5a5307220 · outbound

This paper cites Visual-language prompt tuning with knowledge-guided context optimization,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Visual-language prompt tuning with knowledge-guided context optimization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.589453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.936059Z digest=sha256:2e951ddfb3f36844eb7c702ee394d438e41516d8a36f0e290ea8171f18f85377

Observation bd6d3c8b-1362-4dc4-9d32-66c8364755c7 · outbound

This paper cites Gradient-regulated meta-prompt learning for generalizable vision-language models,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Gradient-regulated meta-prompt learning for generalizable vision-language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.463189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.940199Z digest=sha256:851a4c107d0ca8bfdc60fb026c130d96a98b3a050cab27c255df4573868fcfbc

Observation 10ae3ddd-ef59-4170-a0e7-6ca883cffe58 · outbound

This paper cites Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.944447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.944447Z digest=sha256:63852a96e0e0da59f2c4c671c097b9fb571d44af30f181dbfe77fc04c0407227

Observation 9a37ab55-322c-41fb-aa5d-67c6ecef6e1f · outbound

This paper cites Prompt Learning via Meta-Regularization.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Prompt Learning via Meta-Regularization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:27:47.531123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.950350Z digest=sha256:6ec6b3c8d4a774d99255c7321a224c65612986f9ff4459e1d26cf7acbc710c68

Observation 220f93bd-0fb4-4ca8-9531-dfae0fce96f5 · outbound

This paper cites Extract free dense labels from clip,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Extract free dense labels from clip,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.954920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.954920Z digest=sha256:018c0e027e4e4fdd03b2300c8bb87febd02e3a6ba86142863ea2c101f7a3ba21

Observation 30df0cf1-407e-485a-b5cf-7dffa62d7dc8 · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Maskclip: Masked self-distillation advances contrastive language-image pretraining,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.347774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.959853Z digest=sha256:576eb887ed621503593ecb5c7f2254990da18afb39b5ccc5ad065a5bc5130480

Observation 656ca6ad-35e3-44cb-8f23-7b855be7738e · outbound

This paper cites Scaling open-vocabulary image segmentation with image-level labels,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Scaling open-vocabulary image segmentation with image-level labels,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.187378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.964096Z digest=sha256:143f1d5d559d6d536f3acb2a223e2844ffd2335cb57d1d934aa48774dba05285

Observation 91fa11dc-3e8a-4aa1-a9ca-edfed4077d34 · outbound

This paper cites Clip-actor: Text-driven recom- mendation and stylization for animating human meshes,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Clip-actor: Text-driven recom- mendation and stylization for animating human meshes,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:52.049810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.968647Z digest=sha256:a63b47da637b6a26fc0ec5c33000cc7d4fd719d464f59e06f92beea8a8ac4359

Observation 0a060441-440b-47f6-a6c8-5f8ac15e06e3 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Motionclip: Exposing human motion generation to clip space,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.926665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.972909Z digest=sha256:b590244a7f0918544d9ea5ec4bea2372c4d510e1b2bbb05107427b29e4668a34

Observation 4eecf22f-9870-4ade-b773-d53cbb9d45ec · outbound

This paper cites Lerf: Language embedded radiance fields,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Lerf: Language embedded radiance fields,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.794526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.977678Z digest=sha256:0f5608a5bd0447cad3bbed53b2a3d61ed058d30fa48d316f9a640b498a4fba9d

Observation 02d23dc6-b701-44f5-ada5-0e86715b5721 · outbound

This paper cites Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.672169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.982662Z digest=sha256:0de6b93fc4aa6c7bc51620f6c4f66266194f36c3b9a3bf869f6c68ba12d2f495

Observation 23b94d3d-e2e8-4244-8bbb-6e32f8929bca · outbound

This paper cites Zero-shot text-to-image generation,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Zero-shot text-to-image generation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.986836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.986836Z digest=sha256:5f41cf393ce25be33180e40f2f7c1081b1a60883b2dc6885d1a1273d3f533a93

Observation 35d1dcec-ae34-4f33-bb68-d42d1619d1db · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models High- resolution image synthesis with latent diffusion models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.991834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.991834Z digest=sha256:44e172e5b1ed5d2b549f9827c8755aa799a26a0ed7b388445e378a4dff9a0f05

Observation 6abaeceb-78db-4f94-a602-10c36459ebab · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.558422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:46.996424Z digest=sha256:e873fb0a8daa458cb6025b5e5575bdeddcc27ee51dd97253dea1d504833e9729

Observation dc50ac75-8e87-4e05-b27d-bdd526939afd · outbound

This paper cites Learning multiple visual do- mains with residual adapters,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Learning multiple visual do- mains with residual adapters,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.422176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.000463Z digest=sha256:379f5abfed8b267c0169ed384891bcbb65c8ee6d9913e1fcbd1af94abe90b137

Observation df3884a4-e868-4a13-b3f5-779d240ad034 · outbound

This paper cites Task residual for tuning vision-language models,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Task residual for tuning vision-language models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.004756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.004756Z digest=sha256:936814b2f712dcd3a76324cdd21222f9d3d080351f5eb2ba036faab48cecb619

Observation 1049f89e-7bdc-42f0-842a-40ef4abe2cfb · outbound

This paper cites Graphadapter: Tuning vision-language models with dual knowledge graph,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Graphadapter: Tuning vision-language models with dual knowledge graph,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.287187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.009320Z digest=sha256:e7d3ba5e580c5e15ca0f2f91c5b49b1b9640a299eab30f117aa665ee3315c449

Observation 1eda3c94-5ad5-4f4d-be25-959d22145c73 · outbound

This paper cites Tip-adapter: Training-free adaption of clip for few-shot classification,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Tip-adapter: Training-free adaption of clip for few-shot classification,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.184126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.014458Z digest=sha256:39eb2cd0939b3ac285d38d61210bdc2e987f91e1291c75a18dea9c2894c32325

Observation 1abc598d-c2f6-4616-ad26-a737a108eb74 · outbound

This paper cites Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:51.062183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.020022Z digest=sha256:cd228d3de7986522739d422d89da2a2e8e78bddfe0b38544aff0bf572b76d478

Observation 9f914995-6d80-4a1f-9ae2-225555a4edfd · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:50.934926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.024837Z digest=sha256:9771d2d9111b2d674c6b374a1648dd18de2973561b9809d3dea4276a4f90f7da

Observation a0da0837-f247-4fa4-961c-15e60e0e2d10 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Clip-adapter: Better vision-language models with feature adapters,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.029881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.029881Z digest=sha256:d65b74eebf2f88fa080a4884ee6790c9dbc922059696f3525f69e3f7f9300a6f

Observation c4e96674-b6bb-4465-8a7b-65835d098491 · outbound

This paper cites Visual instruction tuning,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Visual instruction tuning,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:50.825910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.034432Z digest=sha256:ffe2948788789d3d28b1918861c568e94d284109066af35cc56840f804cbec4d

Observation ef0bc9f0-d96f-414c-8020-02d97cf13d3b · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:50.620422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.040924Z digest=sha256:1faf4882af70cb9b57ab174a30d1fd5453b0db20bc7a0cb0810bfc9a89944536

Observation 28ae536f-a7b7-4040-a6a4-6e8d39f811f1 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.045952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.045952Z digest=sha256:3beaeba555ad047e9aa3564af4e53863fe3294f7e4fff2d579682ea75e41f86f

Observation fb9e63f8-2e1f-419d-83b6-5ea2e4e0e2bc · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.051880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.051880Z digest=sha256:a4b8d9d5839942c48f6b08c01013e035873e237344d7e76a9cf9e1c4e42e97f4

Observation b6fbb7e0-9f7a-4505-a77c-8edc06cd4e2c · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.056475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.056475Z digest=sha256:0d625c8d7556a5c7aeefd7ce83b87a83ed40178f85f5f2f2ca8eab12ef39ede4

Observation db69644f-4f1c-46dd-892f-a28717e3872e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.061316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.061316Z digest=sha256:aa429336e92ea8a8208a5327afdda7fc831220c53bae603bab754f9142a3c08c

Observation 4ea8671e-27eb-4810-9f07-15b354a3fbe0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.067589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.067589Z digest=sha256:7fcb048c6b04ac1420fee843a89e7f1cf948ac362d36a6946716041dbbbd0ba6

Observation df2878d6-2386-40f8-a3be-cf412156b8f0 · outbound

This paper cites Language models are few-shot learners,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Language models are few-shot learners,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:50.483391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.072777Z digest=sha256:7182c54eb7a02a1455064d4a795c2f3395a137fbb75f575484429baee2c6752f

Observation e21620ed-086e-4534-9205-d2ba8e2d569d · outbound

This paper cites GPT-4 Technical Report.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.077647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.077647Z digest=sha256:0eca2a9cc62b26c13d652973cd490140587500a95b013a2e96de2d5c137b2d10

Observation da45060a-6976-45f4-af72-cf3165bc0341 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Lit: Zero-shot transfer with locked-image text tuning,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:50.324687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.082212Z digest=sha256:aee0f22d8b6ed762fe7505aea5dfa10092a1917ac163f97ac20daff51feb1a18

Observation 35d68594-d399-4b80-9387-415d325d1bc9 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.087432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.087432Z digest=sha256:c1e1174e2ac4f7bed58abb9b5e3512384079c033693673c4d8541f38b457f3df

Observation 9f233f02-4d06-4e35-acbb-1f53b16ebe0c · outbound

This paper cites Compound text- guided prompt tuning via image-adaptive cues,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Compound text- guided prompt tuning via image-adaptive cues,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:50.217897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.091790Z digest=sha256:db03469b2586784acea957def56802d6ca3d6b20e80aad4da951b2a6ed7a001f

Observation e3af5be0-e7e6-4bbf-ad16-cf73bf0f6d94 · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classification,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models What does a platypus look like? generating customized prompts for zero-shot image classification,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:50.121431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.095841Z digest=sha256:bfc2b306d080e2256e3eddccc19e6f260a560a888c584ede5ddd156e0239a6f7

Observation 058280ef-4874-4509-bcbf-2194a2a50869 · outbound

This paper cites Knowledge- aware prompt tuning for generalizable vision-language models,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Knowledge- aware prompt tuning for generalizable vision-language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:49.956199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.099862Z digest=sha256:b81a069c86c0d5fb267071de2df92f3d82aecfc2de6e0cad6d97b5ca4aadf848

Observation 1d8720a8-38d4-4b3a-8478-00c46d016f16 · outbound

This paper cites Visual Classification via Description from Large Language Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Visual Classification via Description from Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.103905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.103905Z digest=sha256:e63a5b4e104cd2025898b0ebab2d4c5c4a4982af3bb1e06b08eb5d45b2a547c7

Observation 75668555-2309-4fa6-bb4c-f2118b5d3851 · outbound

This paper cites Learning concise and descriptive attributes for visual recognition,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Learning concise and descriptive attributes for visual recognition,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:49.775912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.108072Z digest=sha256:7967df2182d4bd851b2388b96a8507f66787bf2bbe58fceaca0458eb5e2bdcc1

Observation 34ada827-974f-4242-ba9b-9fd737e8d60b · outbound

This paper cites Language in a bottle: Language model guided concept bottlenecks for interpretable image classification,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Language in a bottle: Language model guided concept bottlenecks for interpretable image classification,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:49.621856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.112614Z digest=sha256:c45a58676b8736d794660f397ced7eb06777d8dab198ec424411de2911a65c03

Observation 50027c03-64e6-4f04-9744-23d2afa54312 · outbound

This paper cites CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.118712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.118712Z digest=sha256:af0ff447fe4ac73214cb3328f4720421748f883a30e6662a478cd541e932cb21

Observation c6b5550b-2977-42f3-9ad0-d2c49cf999a3 · outbound

This paper cites Enhancing clip with a third modality,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Enhancing clip with a third modality,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:49.445333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.123856Z digest=sha256:d6da66c8f278946403722664aca3bb6a438316d68bb4d742d4a027dc4f810133

Observation ec1c2b0c-570f-417f-8122-21e4f8930b80 · outbound

This paper cites Democratizing Fine-grained Visual Recognition with Large Language Models.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Democratizing Fine-grained Visual Recognition with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.128776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.128776Z digest=sha256:693e5ce639a767221bf63de5cf367e229a9badfcf9f75dc7cc59acabfd07856e

Observation 47bfcda6-c303-4016-97c9-7eb8c868765a · outbound

This paper cites Fine- tuned clip models are efficient video learners,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Fine- tuned clip models are efficient video learners,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:49.292751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.133211Z digest=sha256:0eaa29e004f97ef437cdfdd0d2964730c4d756b05235e93756aed369afd41aad

Observation 14e7dd2a-a6e4-44ef-83af-2c44af5d7ca2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Imagenet: A large-scale hierarchical image database,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.137318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.137318Z digest=sha256:320d51b99eb9f65a9efb5ba2310fbb105345997fe0557094a3cf6a8b61e2812d

Observation 475c5e0c-cfcc-4dfa-988d-dfcd5af801cf · outbound

This paper cites 3d object representations for fine-grained categorization,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models 3d object representations for fine-grained categorization,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:49.180107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.141493Z digest=sha256:3ba9c9958a555992f8af847cb5971ae500c4be5f8bdebc3104d3e3d81cbcb965

Observation 9857915e-5ab4-4c8a-8e06-b24ec7478460 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.145588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.145588Z digest=sha256:c5adac63390791af4ec89ab16ad60b21b4bf05cfe50d97a9e333bd6125c446b3

Observation 15ca188b-e6c5-41de-ac53-07fcd31661af · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.150373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.150373Z digest=sha256:84911bf7958cd2806969062a58bcb74272fb367eefa21c06e8a3eeaa6869e276

Observation 0a4b130c-298f-4fa9-8c3b-a67156111844 · outbound

This paper cites Automated flower classification over a large number of classes,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Automated flower classification over a large number of classes,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:48.995277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.155160Z digest=sha256:7e89d0794480d8fcc254458d58954801d82b950e47125518ee8999c80fd1f52b

Observation cfb0970c-d0ef-4a1e-9a4e-cdb51eac85ba · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Sun database: Large-scale scene recognition from abbey to zoo,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:48.855739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.159461Z digest=sha256:ed50a5c40f5250f6e2e9b14eb6f3a23e41ea05b44cdfdc29e7a319ad4445c1b2

Observation cb1de294-5148-4188-80ce-d383d634eec2 · outbound

This paper cites Describing textures in the wild,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Describing textures in the wild,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.164738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.164738Z digest=sha256:c9eeb4d59e93e14b5873bdacd5c808b916ddbf47610f451337456e7e40be0cfb

Observation a67ce6cd-e736-4acf-ab61-68a8ea592866 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.169601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.169601Z digest=sha256:74912c82532fda858945c7b3d1d94da0f38a26935a4c01b713c900d967b69287

Observation e59cc4cc-4243-4123-a5ec-3092aa3d8d3a · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Fine-Grained Visual Classification of Aircraft

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.173763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.173763Z digest=sha256:109d8a46360c0a3923cebd5d9e6d5afa541a9b9860fcd0096ba6d1dc3526dc53

Observation b3085d30-3809-40f2-8b98-d7c610e63ed1 · outbound

This paper cites Cats and dogs,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Cats and dogs,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:48.630071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.178390Z digest=sha256:e61a04a62ce1c51524d3fd5d01a3756f253b0bef4382e8b929c980ff333c392f

Observation 909a32ad-7675-47a7-b735-0ae79adae21f · outbound

This paper cites Food-101–mining discriminative components with random forests,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Food-101–mining discriminative components with random forests,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:48.360216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.182947Z digest=sha256:92df0c3a00815caedf250a5dc169e0a191cb7eab128ea9686fd5f84336a9faca

Observation 72fc57ec-122e-432a-9988-717872c231ea · outbound

This paper cites Natural adversarial examples,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Natural adversarial examples,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.187502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.187502Z digest=sha256:45bf0208fc879fe43d4ac394543143de30a0bf7b33d44f70fb714f4128e07302

Observation cc9ce98a-bbfc-4d28-8fe6-d2c3c6a19bb2 · outbound

This paper cites Do imagenet classifiers generalize to imagenet?.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Do imagenet classifiers generalize to imagenet?

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:48.217435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.191888Z digest=sha256:627f2cfc0a79c0a8f86c8102cb1e1e0c12e4ad76286af5e8e16898ad73fc9fd9

Observation 65100d47-d282-481d-8bd0-386ace54c4d1 · outbound

This paper cites Learning robust global representations by penalizing local predictive power,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Learning robust global representations by penalizing local predictive power,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:48.073216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.196049Z digest=sha256:fe8af480a8d9902716851b0beae290c2ad64cbbd931148d28f823b2dfdc570d4

Observation 3f9b258a-8bca-4d7d-a7a9-79e5628ee05f · outbound

This paper cites Decoupled Weight Decay Regularization.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Decoupled Weight Decay Regularization

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.200213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.200213Z digest=sha256:951fb7a310de80f6a7630c35c6bd6b2d839fa1e8e11fd6bc5e10a6634182c57c

Observation ae1a09a8-1962-4d48-a8aa-019755f1e3df · outbound

This paper cites VL-Mamba: Exploring State Space Models for Multimodal Learning.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models VL-Mamba: Exploring State Space Models for Multimodal Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.204582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.204582Z digest=sha256:27f6736f7a8912944020bfd1e58237271876519c8651aba9f46872c1b7c6e49f

Observation fbef7d83-68a3-4d39-aaeb-6e6f10e51d9c · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Grad-cam: Visual explanations from deep networks via gradient-based localization,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.208713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.208713Z digest=sha256:9aeb5951d5a5cff6592fed08cd68046ef4b019473f081d84747d05708c3ed176

Observation c553cb0d-8a20-4659-9ea2-1ba1c4a0a9f7 · outbound

This paper cites Efficiently scaling transformer inference,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Efficiently scaling transformer inference,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.212991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.212991Z digest=sha256:4b9337e044852a158fb5b81d11783c207a9c7e4262e6251292a1cbb11edf0915

Observation 5a4c514e-fda3-44e4-acaf-7f7137416859 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Transformers: State-of-the-art natural language processing,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:47.920034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.217153Z digest=sha256:0177e5175a70925844b985e791d6491f2bdaaf950b876d57532c85deeffa2b69

Observation 934244d5-50ea-42da-a768-6dbe291ce291 · outbound

This paper cites Modality- consistent prompt tuning with optimal transport,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Modality- consistent prompt tuning with optimal transport,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:47.791480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.222518Z digest=sha256:803c623d10dc0c67a5d0961603d2d44f4dc113f716038ae106d3af13174b6717

Observation c1c080de-e9d9-4f7d-8580-e4b6fec02d7c · outbound

This paper cites Hierarchy- aware interactive prompt learning for few-shot classification,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Hierarchy- aware interactive prompt learning for few-shot classification,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:47.761692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.226756Z digest=sha256:4983e0cf2f74cf4bf2e22395c80cffeb695d434ded47b9fb43d7caad14b88e3a

Observation b8cd94c3-82c5-4186-96e5-eb32b176a837 · outbound

This paper cites Language-driven visual consensus for zero-shot semantic segmentation,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Language-driven visual consensus for zero-shot semantic segmentation,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:47.737604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.232099Z digest=sha256:865cb7d942c007f2bdcd4715da9ac00b794317a5a9e37781331be2a9db6fbebe

Observation 0626ab68-3e72-4287-8bd3-3356f62aaa44 · outbound

This paper cites Pedestrian attribute recognition via clip based prompt vision-language fusion,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Pedestrian attribute recognition via clip based prompt vision-language fusion,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:47.711586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.236731Z digest=sha256:f32b0bb539b61cfe97c1f8d890cc705fb865b4d6ca9d69a244cb1816640227dc

Observation d0247ef0-75df-4c82-9424-44f458c37556 · outbound

This paper cites Understanding and mitigating overfitting in prompt tuning for vision-language models,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Understanding and mitigating overfitting in prompt tuning for vision-language models,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:47.240681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:47.240681Z digest=sha256:07a61a315ba5f93a4df7409151cf9cd3d4a55c4194174b8d58f8ad7846d11d9e

Observation 536fa8b1-0817-4ccd-873a-84eb3ba6cabb · outbound

This paper cites Clipood: Generalizing clip to out-of-distributions,.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Clipood: Generalizing clip to out-of-distributions,

Reference 86

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:27:47.672562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:27:47.245370Z digest=sha256:40ff507417d8103e8b1b46e370ded564034b3e6872e9e795bad4ba1e105121cf

Pith citing papers

No inbound Pith citation observations are available.