Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2111.09734.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:18.994193Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:09:44.426723Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9e39cb2a-dfe9-4514-bac3-8f235c4f40e0 · inbound
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language ClipCap: CLIP Prefix for Image Captioning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cfeab8e-d3b3-4a90-9f76-16506dde437b · inbound
Flamingo: a Visual Language Model for Few-Shot Learning ClipCap: CLIP Prefix for Image Captioning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f11e737-20da-46c7-9a1c-c8852ac93363 · inbound
LAION-5B: An open large-scale dataset for training next generation image-text models ClipCap: CLIP Prefix for Image Captioning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ef5c014-1562-48d7-b252-38993a4e3396 · inbound
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention ClipCap: CLIP Prefix for Image Captioning
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7f1d37a-c0ba-42f8-a149-1b8b39a8f713 · inbound
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model ClipCap: CLIP Prefix for Image Captioning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9e2d1eb-3be7-4be7-96fe-9b7e798c957e · inbound
EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers ClipCap: CLIP Prefix for Image Captioning
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b33833fc-4b2a-4ccf-82ac-1e8cdf30cc3c · inbound
Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales ClipCap: CLIP Prefix for Image Captioning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22dbbde-1000-446b-a047-623d6876de8c · inbound
Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants ClipCap: CLIP Prefix for Image Captioning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab95742f-af91-4837-a8be-418c7492f131 · inbound
Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model ClipCap: CLIP Prefix for Image Captioning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22bc1e94-042a-402c-aa38-1f117cddb53e · inbound
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models ClipCap: CLIP Prefix for Image Captioning
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01ceb64c-abd4-4229-8c0d-a17cd390032c · inbound
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering ClipCap: CLIP Prefix for Image Captioning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe79f438-6836-4404-9e98-d2e7b2503694 · inbound
Diffusion-based Cumulative Adversarial Purification for Vision Language Models ClipCap: CLIP Prefix for Image Captioning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c28283-b524-4c6a-bf1d-4adbcaf03595 · inbound
CoLMbo: Speaker Language Model for Descriptive Profiling ClipCap: CLIP Prefix for Image Captioning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf1e06e5-7ae9-47e5-bee9-f8ce5f3120ed · inbound
InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning ClipCap: CLIP Prefix for Image Captioning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb5433a-a261-423b-a2dc-691adfbaa877 · inbound
CLIP-HandID: Vision-Language Model for Hand-Based Person Identification ClipCap: CLIP Prefix for Image Captioning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59700003-ea25-4573-9220-5fa415e5f49c · inbound
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46177889-2581-457d-bd97-0e86f0a016ed · inbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · inbound
On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5772af9d-23eb-449e-94e4-c3813760953a · inbound
GLAD: Generalizable Tuning for Vision-Language Models ClipCap: CLIP Prefix for Image Captioning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc89fd79-83dc-4729-b9f0-7d938b076d9e · inbound
HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals ClipCap: CLIP Prefix for Image Captioning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ab0b8a-1df3-43a3-a031-d3671c0d56f2 · inbound
VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings ClipCap: CLIP Prefix for Image Captioning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e640f75c-908b-4092-86a0-39c9f77051fa · inbound
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models ClipCap: CLIP Prefix for Image Captioning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ce12810-cef8-4ad1-84c7-e4a6e8381ad2 · inbound
RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification ClipCap: CLIP Prefix for Image Captioning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b79eda-e591-486e-aa15-4fabec394308 · inbound
From Image Captioning to Visual Storytelling ClipCap: CLIP Prefix for Image Captioning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572e5683-939d-454a-90ad-ac72b0abbdc5 · inbound
Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval ClipCap: CLIP Prefix for Image Captioning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d63d557-c52c-4e6c-9495-328a74914cf5 · inbound
Sample-efficient Integration of New Modalities into Large Language Models ClipCap: CLIP Prefix for Image Captioning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb27136-dcf8-43a4-98ed-2f7bc45059c5 · inbound
Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models ClipCap: CLIP Prefix for Image Captioning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b145175-6648-4219-93e1-146d1a146dd3 · inbound
Unpacking Hateful Memes: Presupposed Context and False Claims ClipCap: CLIP Prefix for Image Captioning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b975f0e4-ca19-4fb7-8c9c-0fd08cb67a97 · inbound
MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models ClipCap: CLIP Prefix for Image Captioning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1124a2e-2683-494b-8f0b-106e84704a6b · inbound
Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain ClipCap: CLIP Prefix for Image Captioning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0ce17dd-d1ad-4f4c-bd35-81170c9a0278 · inbound
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs ClipCap: CLIP Prefix for Image Captioning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df69bec-bde3-4b33-8a69-849040417ee2 · inbound
Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks ClipCap: CLIP Prefix for Image Captioning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f07988e-3ad6-4978-80fe-4f081982dca2 · inbound
Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection ClipCap: CLIP Prefix for Image Captioning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71069924-22ff-4aa0-9434-e119450329e7 · inbound
WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval ClipCap: CLIP Prefix for Image Captioning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7220993c-f54f-4276-a5d2-fa3928c388d9 · inbound
UIPress: Bringing Optical Token Compression to UI-to-Code Generation ClipCap: CLIP Prefix for Image Captioning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8baee550-4afd-4f0e-93bb-73d6a391dc69 · inbound
Semantic Manipulation Localization ClipCap: CLIP Prefix for Image Captioning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f5c2bc7-85a8-4cc2-9b2d-d5129afbf14b · inbound
Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment ClipCap: CLIP Prefix for Image Captioning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6448c78a-b35d-4ece-9770-8503072ef485 · inbound
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models ClipCap: CLIP Prefix for Image Captioning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8556bc0-b861-478f-a3f6-fd1aa2d5a44e · inbound
CB-SLICE: Concept-Based Interpretable Error Slice Discovery ClipCap: CLIP Prefix for Image Captioning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d56f8ed-0f49-49f4-9f85-a227732d0549 · inbound
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning ClipCap: CLIP Prefix for Image Captioning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8dcd612-4f0c-4991-9f92-1439fb961543 · inbound
Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models ClipCap: CLIP Prefix for Image Captioning
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 654c879a-756d-4bd0-8cac-797032805010 · inbound
Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning ClipCap: CLIP Prefix for Image Captioning
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a362ee84-0f95-42af-9045-4b243e1564fa · inbound
Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality ClipCap: CLIP Prefix for Image Captioning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d8381db-8362-4d59-a50f-42218480667e · inbound
Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors ClipCap: CLIP Prefix for Image Captioning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bafcd518-7a9b-4e9f-bb70-5fe86f8df660 · inbound
READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations ClipCap: CLIP Prefix for Image Captioning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5b94a51-91e3-4fd7-a8e7-d2225fc2f004 · inbound
Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection ClipCap: CLIP Prefix for Image Captioning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a086778-493c-4d97-8136-8dd41944fa15 · inbound
REPREC: Representation Driven Parameter-Efficient Recommendation System ClipCap: CLIP Prefix for Image Captioning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5627ffab-96fe-4073-beb5-8f53e77b3fa6 · inbound
Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning ClipCap: CLIP Prefix for Image Captioning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.