Pith. sign in

Paper Citation Record · LEDGER

ClipCap: CLIP Prefix for Image Captioning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2111.09734.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09734 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:18.994193Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.426723Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9e39cb2a-dfe9-4514-bac3-8f235c4f40e0 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language ClipCap: CLIP Prefix for Image Captioning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.718168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:5893e9513b907043de8da5fc7f7c5a1d70a84499f7426c68fc63810794e9540e

Observation 8cfeab8e-d3b3-4a90-9f76-16506dde437b · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning ClipCap: CLIP Prefix for Image Captioning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.476566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:9d7bff8590bc65f7b4c2205392183edb97447c486e958974485b71e77c1a59df

Observation 7f11e737-20da-46c7-9a1c-c8852ac93363 · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models ClipCap: CLIP Prefix for Image Captioning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:22:17.333000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:dfefbbbca045562d3c8a1b69a9a4a92f2b147c2708353aa9d512fa97aa3ec430

Observation 1ef5c014-1562-48d7-b252-38993a4e3396 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention ClipCap: CLIP Prefix for Image Captioning

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:07:42.528525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:d2f72990500f5f0f869a5ac6894adb3f3887af9de692d5ba0754075712ebe6b4

Observation f7f1d37a-c0ba-42f8-a149-1b8b39a8f713 · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model ClipCap: CLIP Prefix for Image Captioning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:41:04.887711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:708fc108b1fea59f363d842b20453b21a05e199fe3c98e637f4add6ca4cc1cc8

Observation c9e2d1eb-3be7-4be7-96fe-9b7e798c957e · inbound

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers cites this paper.

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers ClipCap: CLIP Prefix for Image Captioning

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:11:49.649827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:11:49.475825Z digest=sha256:a65e991d0daa1fec794d3d287019b8b717e0466b80d4c7c6295f9dc14db1be7a

Observation b33833fc-4b2a-4ccf-82ac-1e8cdf30cc3c · inbound

Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales cites this paper.

Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales ClipCap: CLIP Prefix for Image Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:18.994193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:18.994193Z digest=sha256:feac597ada9e75d4561f9a56285dd0000531413e3f5175a0f3c8a85c315bbc7d

Observation c22dbbde-1000-446b-a047-623d6876de8c · inbound

Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants cites this paper.

Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants ClipCap: CLIP Prefix for Image Captioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:19.824796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:40:19.824796Z digest=sha256:544615c90ec82e8f3b923e678adab7c3e00bd31c16bf9988e265bd22d77719a8

Observation ab95742f-af91-4837-a8be-418c7492f131 · inbound

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model cites this paper.

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:44.090056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:44.090056Z digest=sha256:df0e7213e971f63bfe89448a856030fc436016587fe8ee2a38179a18c3d43373

Observation 22bc1e94-042a-402c-aa38-1f117cddb53e · inbound

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models cites this paper.

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models ClipCap: CLIP Prefix for Image Captioning

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.351321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:36.351321Z digest=sha256:46e2dcc2717372d26b0793c91c8543e5b204824f5f3226fa7d4e27586f4ab73c

Observation 01ceb64c-abd4-4229-8c0d-a17cd390032c · inbound

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering cites this paper.

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:10.602373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:10.602373Z digest=sha256:f2fd43660257cd7b62be83d8efea541b12e9a163dacc3915ee746048d013d803

Observation fe79f438-6836-4404-9e98-d2e7b2503694 · inbound

Diffusion-based Cumulative Adversarial Purification for Vision Language Models cites this paper.

Diffusion-based Cumulative Adversarial Purification for Vision Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:59:18.927640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:59:18.927640Z digest=sha256:8aed347f717ddb42e8742d6aa948fc8338271828deb83a89f499ac68823d1ffc

Observation f8c28283-b524-4c6a-bf1d-4adbcaf03595 · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling ClipCap: CLIP Prefix for Image Captioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:15.589604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:15.589604Z digest=sha256:3d633b5aff1920b320d6bfb948d9c4143f1509f416483de45ec591dccef2302c

Observation cf1e06e5-7ae9-47e5-bee9-f8ce5f3120ed · inbound

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning cites this paper.

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:58:07.909306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:58:07.909306Z digest=sha256:278e34c3ed8e45ba570f20a4bb5f817aeefe56dbe68cf4f5da6681e941ecba4d

Observation 3bb5433a-a261-423b-a2dc-691adfbaa877 · inbound

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification cites this paper.

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification ClipCap: CLIP Prefix for Image Captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:56:01.612867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:56:01.612867Z digest=sha256:defcedfe3d94f5226a842252d4fe8e3aaaad8ba5757a552bdf08467c7f224e7c

Observation 59700003-ea25-4573-9220-5fa415e5f49c · inbound

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge cites this paper.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.828926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.828926Z digest=sha256:fe626f5ad010e00def0b3c3e9604b94356460ae6adda6466326df97ce76c8a5b

Observation 46177889-2581-457d-bd97-0e86f0a016ed · inbound

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition cites this paper.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.557921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.557921Z digest=sha256:6b4e8a7e39c59d7d68c4590a482f56b08bf797b7c923ae9d59e9fb3099efc4ac

Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.511161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.511161Z digest=sha256:89b05099011a786ca7635d9c3532428f0efc3a68dbbeec842c7ddcfaebce83c5

Observation 5772af9d-23eb-449e-94e4-c3813760953a · inbound

GLAD: Generalizable Tuning for Vision-Language Models cites this paper.

GLAD: Generalizable Tuning for Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.189934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.189934Z digest=sha256:cdb7bfd2e6bbc4ff5a0654d1efc0274f823f9b4cc30d6b8a8c57ddfdd1b0e884

Observation bc89fd79-83dc-4729-b9f0-7d938b076d9e · inbound

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals cites this paper.

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals ClipCap: CLIP Prefix for Image Captioning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:32:42.735751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:32:42.735751Z digest=sha256:c38db86741cf2089232bf07cad4311dbadbc0be74c6cf11c4a21d48c88caf064

Observation 31ab0b8a-1df3-43a3-a031-d3671c0d56f2 · inbound

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings cites this paper.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.416252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.416252Z digest=sha256:12a94da8d2e95d704036c7cc7dac0d0fa5bda84ec3c19238b83ddedd6957813f

Observation e640f75c-908b-4092-86a0-39c9f77051fa · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:16.018283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:16.018283Z digest=sha256:ed4bb610e82b0a71307b6d2c413953419ccba71bbb4f53e77a08831af207d11e

Observation 6ce12810-cef8-4ad1-84c7-e4a6e8381ad2 · inbound

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification cites this paper.

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification ClipCap: CLIP Prefix for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:10.551418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:04:10.551418Z digest=sha256:a7743535f486157e6506ae237b9a75a82968d73d8bcf598e4538674cfab13b95

Observation 37b79eda-e591-486e-aa15-4fabec394308 · inbound

From Image Captioning to Visual Storytelling cites this paper.

From Image Captioning to Visual Storytelling ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:39.338214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:39.338214Z digest=sha256:e823e0b488fbcf11ca7d635ad934d803248e6b121e01d15f092580ffce32b23e

Observation 572e5683-939d-454a-90ad-ac72b0abbdc5 · inbound

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval cites this paper.

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval ClipCap: CLIP Prefix for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:25:01.008591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:25:01.008591Z digest=sha256:c5f463b7f16ca268d35e229442b5d8410950ccf8ab1a5b1ca2374d17936a8230

Observation 5d63d557-c52c-4e6c-9495-328a74914cf5 · inbound

Sample-efficient Integration of New Modalities into Large Language Models cites this paper.

Sample-efficient Integration of New Modalities into Large Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.509894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.509894Z digest=sha256:781ce42cccf19b131db287f377cb7e14a07eb4d7b06ce146e8a6f044f64fd434

Observation cdb27136-dcf8-43a4-98ed-2f7bc45059c5 · inbound

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models cites this paper.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models ClipCap: CLIP Prefix for Image Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:40.077127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:40.077127Z digest=sha256:a99df999f496ead07e208a8b366d2e4331e08355be30c149998d76a724c2372c

Observation 2b145175-6648-4219-93e1-146d1a146dd3 · inbound

Unpacking Hateful Memes: Presupposed Context and False Claims cites this paper.

Unpacking Hateful Memes: Presupposed Context and False Claims ClipCap: CLIP Prefix for Image Captioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:26:22.120641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:26:22.120641Z digest=sha256:3df0b552b7c877feba52e5ea39e7bd3506afe272bf3e856f12c05e91281e7c85

Observation b975f0e4-ca19-4fb7-8c9c-0fd08cb67a97 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models ClipCap: CLIP Prefix for Image Captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:04.998120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:04.998120Z digest=sha256:a9040cae5919f3651a6f32065920b5301523edd07a9d1e92b70de90225fc8f11

Observation c1124a2e-2683-494b-8f0b-106e84704a6b · inbound

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain cites this paper.

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain ClipCap: CLIP Prefix for Image Captioning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:27:59.324632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T14:26:08.188950Z digest=sha256:ac26f7bfa130f115e7ab7af3e46b5713cf845b0383a261ff79e9c71ebf379260

Observation b0ce17dd-d1ad-4f4c-bd35-81170c9a0278 · inbound

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs cites this paper.

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs ClipCap: CLIP Prefix for Image Captioning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:52.217522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:52.217522Z digest=sha256:94d160edf899ec9c852502d9f36a8d555e44cdc61accc5f8b7e8787a3f987dff

Observation 9df69bec-bde3-4b33-8a69-849040417ee2 · inbound

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks cites this paper.

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks ClipCap: CLIP Prefix for Image Captioning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:52.655740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:26:42.128350Z digest=sha256:a1386ceb1efe36292f05bb432786261212b18f64233ceba27eca1cf35285857c

Observation 5f07988e-3ad6-4978-80fe-4f081982dca2 · inbound

Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection cites this paper.

Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection ClipCap: CLIP Prefix for Image Captioning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.240217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:52:34.752473Z digest=sha256:4fc82193eb3382b273afb0cd8af1fc4a16e6189ae964eb725fc98ce19610c7bf

Observation 71069924-22ff-4aa0-9434-e119450329e7 · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval ClipCap: CLIP Prefix for Image Captioning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.060959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:a2d903a2d9d96f8bf4bac7f5567fada282d60f218602b6ccce5e28fb36cd8019

Observation 7220993c-f54f-4276-a5d2-fa3928c388d9 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation ClipCap: CLIP Prefix for Image Captioning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.421777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:5813cfd1f288aa8fd943b6d7b7844a53ed823afcf30b6d0ce279e2eec969c904

Observation 8baee550-4afd-4f0e-93bb-73d6a391dc69 · inbound

Semantic Manipulation Localization cites this paper.

Semantic Manipulation Localization ClipCap: CLIP Prefix for Image Captioning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:21:01.497760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:32:30.345594Z digest=sha256:27fce9d642ad446946c3a0393e178c632889d9934b2230ebe847a95ee1d0d39f

Observation 0f5c2bc7-85a8-4cc2-9b2d-d5129afbf14b · inbound

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment cites this paper.

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment ClipCap: CLIP Prefix for Image Captioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:27.865530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:19:31.570783Z digest=sha256:31aeba3fa1bfd19c711edf6c2d987ef1e9d0b92369420c7a939b00b9f8781ddc

Observation 6448c78a-b35d-4ece-9770-8503072ef485 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:56.290381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:c8da5f9c40fd9ff2056618ce633dfd8335c96b78e266ac4db2d1329960111d09

Observation c8556bc0-b861-478f-a3f6-fd1aa2d5a44e · inbound

CB-SLICE: Concept-Based Interpretable Error Slice Discovery cites this paper.

CB-SLICE: Concept-Based Interpretable Error Slice Discovery ClipCap: CLIP Prefix for Image Captioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.862374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:24:05.284777Z digest=sha256:1651b374bd521fe9baa4983dc666c0ef75352c40e184744f8fd8f2997720aee6

Observation 7d56f8ed-0f49-49f4-9f85-a227732d0549 · inbound

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning cites this paper.

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:11.285729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:52:23.150188Z digest=sha256:c465fdca245b642e542f8ffb2be0f2ded5efafd7c294c65686efb211be7054d9

Observation f8dcd612-4f0c-4991-9f92-1439fb961543 · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.791066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:ebf534d679315b9a7a5d85d9e57e35c046e4d1ccfe17d129b889a82124a9e4fc

Observation 654c879a-756d-4bd0-8cac-797032805010 · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning ClipCap: CLIP Prefix for Image Captioning

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:48.402050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:7ccfa52cfeb5a141dc7d708a147a5710808e3dea06e964cce1a9ff1004ae4e5a

Observation a362ee84-0f95-42af-9045-4b243e1564fa · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:31.474526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:272c3b7f99562762a59f78d372724613a150ef47dbabc607e7a4f6c0b3d1e346

Observation 0d8381db-8362-4d59-a50f-42218480667e · inbound

Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors cites this paper.

Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors ClipCap: CLIP Prefix for Image Captioning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:45.673868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:21:31.374575Z digest=sha256:b5501854ac5a9c76f2a7bcb7d54b77d5e6d33a4434dc5b11629582f5762dd41d

Observation bafcd518-7a9b-4e9f-bb70-5fe86f8df660 · inbound

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations cites this paper.

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations ClipCap: CLIP Prefix for Image Captioning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.428197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:09:23.918802Z digest=sha256:488162493a202b23b6924a2eae7423ac8ea49c518912a8033d9ba31552c6b0ff

Observation a5b94a51-91e3-4fd7-a8e7-d2225fc2f004 · inbound

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection cites this paper.

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection ClipCap: CLIP Prefix for Image Captioning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T12:36:44.253836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:36:44.253836Z digest=sha256:9ef5045d90bcd86a3fc575d390dbd3c1af021fdbd1639795f7396344a0e2cc67

Observation 8a086778-493c-4d97-8136-8dd41944fa15 · inbound

REPREC: Representation Driven Parameter-Efficient Recommendation System cites this paper.

REPREC: Representation Driven Parameter-Efficient Recommendation System ClipCap: CLIP Prefix for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T04:11:19.003885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:11:19.003885Z digest=sha256:0ea222d294c74567e98f0a2600a2b8cf194cf0836298fbfb8d2891496dd88b2c

Observation 5627ffab-96fe-4073-beb5-8f53e77b3fa6 · inbound

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning cites this paper.

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning ClipCap: CLIP Prefix for Image Captioning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T15:58:39.520616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:58:39.520616Z digest=sha256:72911fabe3c17936ec822f3da42df7cc51ca1283cf3a73c70af8ea0173bd4aac