Pith. sign in

Paper Citation Record · LEDGER

Learning to Prompt for Vision-Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2109.01134.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.01134 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:09:12.400739Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:07.768237Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2607
arxiv_reference, observed 2026-07-04T20:00:07.768237Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68a3ef13-421f-44f3-a3a6-9243823bc9d8 · inbound

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion cites this paper.

An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion Learning to Prompt for Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:08:55.541350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T18:08:55.311069Z digest=sha256:4f2dc5431b8a6dbd0f0bb4c1e26ac47093ae7c862cb0dbfaa5c34ca22e886b11

Observation d7d90f01-0bf6-4483-be44-a1e77b96d2e4 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space Learning to Prompt for Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.452379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:908244585ef726fe00f136e6b496b7ad44c76f6799372bad5a1bd8a32b325a65

Observation 57e163cd-281f-4af1-b3db-a4b1ea8f8487 · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey Learning to Prompt for Vision-Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.334206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.334206Z digest=sha256:f38f9b945cb152bf52adcca2b9f0eda2c7c5372dfb8c7117b6e594ba4be915ca

Observation 4f0bb4d4-8c45-4fd6-874b-f3d6569f2a3b · inbound

Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin cites this paper.

Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin Learning to Prompt for Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:12.400739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:09:12.400739Z digest=sha256:20eb4e0e28ff4c7b56cbdc42f35f72be5e734ab86d41ce4b511bb7a029c8e056

Observation 2b551534-5774-43ca-bc29-26542cf35cca · inbound

StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection cites this paper.

StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection Learning to Prompt for Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:29.866668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:43:29.866668Z digest=sha256:d59eb0ad0c6733fe1981b82cefbbfc1d34225dedd10323b4ebce06de453d65bf

Observation 6495820b-8c45-4a35-966d-7ba4266044cf · inbound

SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction cites this paper.

SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction Learning to Prompt for Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.412320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.412320Z digest=sha256:3dc2fbb2894e41b176998d87b6ff0868e3d6d872238ec4113a40b383e1f9287d

Observation 52bc2d64-fc76-4e3f-8c65-74ef080be272 · inbound

$\Delta \mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization cites this paper.

$\Delta \mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization Learning to Prompt for Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:14:08.549417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:14:08.549417Z digest=sha256:f7d2d653382a7c20b127c5a65c74211f517802bcf7db16d0f9b7e13dd71123ce

Observation ef9fd95f-1da4-4d7d-af66-76f365e15dc0 · inbound

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion cites this paper.

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion Learning to Prompt for Vision-Language Models

Reference 12

Resolution
verified exact
doi, observed 2026-05-16T20:18:23.191256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T20:15:50.210484Z digest=sha256:a1e07aa43af2464bfd525b10654fa6f442c54bbf35137a6863f5bc9a0c8c3090

Observation 3d304dad-0597-4e3e-a9a5-8a963d34d842 · inbound

ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification cites this paper.

ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification Learning to Prompt for Vision-Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T04:45:20.976440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:41:38.183119Z digest=sha256:cf11253a287ad40206501944f23282dbc23bac326c42488ff32f16862d398ae0

Observation a5c6457d-c08f-41be-a024-fc9916c5d2bf · inbound

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction cites this paper.

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction Learning to Prompt for Vision-Language Models

Reference 81

Resolution
metadata mismatch
doi, observed 2026-05-09T23:14:36.510352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T23:13:34.852488Z digest=sha256:e3f0509dbbcba6ae80ebf2ef15ee19dab40382a45e8d6a597354f12329094088

Observation 2aa58e67-2cf1-4014-8423-da1565bb5cfa · inbound

Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model cites this paper.

Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model Learning to Prompt for Vision-Language Models

Reference 16

Resolution
verified exact
doi, observed 2026-05-08T18:28:58.211029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:26:24.377120Z digest=sha256:e1513150bc57ffcf2a77cf7a607e4f061c23b60393f32807dfa141a2f139fb28

Observation fefcfcbb-8348-4a23-98db-fd3fecc40e7d · inbound

Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models cites this paper.

Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models Learning to Prompt for Vision-Language Models

Reference 6

Resolution
verified exact
doi, observed 2026-05-10T19:05:44.275069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T19:04:20.674374Z digest=sha256:a7d26c786eeb63e7dbcc8b8db38f676a64d1c24f9f0e7381340fb83cb2aca278

Observation 42ba4757-df8c-4608-bdb5-e337fe221c0c · inbound

FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection cites this paper.

FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection Learning to Prompt for Vision-Language Models

Reference 38

Resolution
verified exact
doi, observed 2026-05-09T01:24:39.013742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T01:26:11.811914Z digest=sha256:2a27ec15bb1d687a2864a335f2c1c8dfa235a22826ff6d1e9f31593a8a9264b0

Observation 142f6e9b-f692-49b2-a0b8-b0229d78405e · inbound

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs cites this paper.

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs Learning to Prompt for Vision-Language Models

Reference 14

Resolution
metadata mismatch
doi, observed 2026-05-08T21:29:15.071225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T13:17:48.295630Z digest=sha256:f5276fb1ba29a08a04bc37e7d263498a662bad030dac1db2f6964acace797ca3

Observation dd73e5e7-bfd3-4153-83c4-51e12cce0f41 · inbound

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning cites this paper.

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning Learning to Prompt for Vision-Language Models

Reference 33

Resolution
verified exact
doi, observed 2026-05-13T01:42:02.392171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:41:07.685217Z digest=sha256:429ec71d1ecc49a229b746165298dc4f9c2fa8773f996b79e01e7d69fc5172f1

Observation e15eada8-be81-4a30-9434-4e2fa8cb232e · inbound

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning cites this paper.

TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning Learning to Prompt for Vision-Language Models

Reference 33

Resolution
verified exact
doi, observed 2026-05-14T21:17:58.410163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T21:14:38.092918Z digest=sha256:99766fefa43852d5ad7ebaf03d76f9002c8864d7bc1c7ad34dcce477032f7e0f

Observation 6de3b287-6771-48c9-859f-749beb818b69 · inbound

Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation cites this paper.

Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation Learning to Prompt for Vision-Language Models

Reference 43

Resolution
verified exact
doi, observed 2026-05-14T19:57:52.550230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T19:55:33.741521Z digest=sha256:958eea87fe5e98e852bf8e814cb94a38bddcb83e3b1d6648a8deee93fbc48eae

Observation 65b83ab6-516d-4082-b17d-0a6e127439a7 · inbound

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media cites this paper.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media Learning to Prompt for Vision-Language Models

Reference 178

Resolution
verified exact
doi, observed 2026-05-20T14:08:20.777711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T14:05:14.737146Z digest=sha256:3400167ea4a0601d3e1308b692102fe92d6051b5ce52d39fff51bb22c160f588

Observation 276e70c1-fb5b-4276-883f-82c75cdbbf38 · inbound

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency cites this paper.

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency Learning to Prompt for Vision-Language Models

Reference 73

Resolution
verified exact
doi, observed 2026-05-20T11:38:14.383828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T11:33:39.121508Z digest=sha256:c746e1361d61287cb740a90c9110d86246c91f77a507e7ea6faf29d0a3907961

Observation 0aec0bb8-3d96-445c-b9aa-b45811663fdd · inbound

PERL: Parameter Efficient Reasoning in CLIP Latent Space cites this paper.

PERL: Parameter Efficient Reasoning in CLIP Latent Space Learning to Prompt for Vision-Language Models

Reference 38

Resolution
verified exact
doi, observed 2026-05-20T11:48:14.619516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T11:45:15.339048Z digest=sha256:3cdb114473859177733246bbd369efffdd12d9948385b6656536bdd62de7e8ad

Observation e4e39745-277a-47bb-ac9a-de4988b1e285 · inbound

Steering Vision-Language Models with Joint Sparse Autoencoders cites this paper.

Steering Vision-Language Models with Joint Sparse Autoencoders Learning to Prompt for Vision-Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:07.769905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T20:56:57.716246Z digest=sha256:e66a2f4d9f2db98e41e14453a73867a522c5a5bc12845bde13ccebc216d589e1