Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2111.07783.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:47:50.404759Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:07:53.330969Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7589c8ec-0140-4b7f-81b2-ff07c748df29 · inbound
Florence: A New Foundation Model for Computer Vision FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f7e8f8d-54ba-487d-829a-68adacc53d8b · inbound
Flamingo: a Visual Language Model for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2adb3506-7cbf-4907-86c4-9ca1be6fd63b · inbound
CoCa: Contrastive Captioners are Image-Text Foundation Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8122669-16c3-4e8f-8f51-891e7666a1e8 · inbound
DetailCLIP: Injecting Image Details into CLIP's Feature Space FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 394409b5-09ee-46ff-b7a7-49f06cd927e2 · inbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 25df6251-3b0c-4dbc-8297-b07522936587 · inbound
VideoChat: Chat-Centric Video Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0674e91f-b242-4d8b-aa76-cbcaa4141965 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7e3e877-1b7e-4f8f-9dab-1f673822d2e1 · inbound
LPT: Less-overfitting Prompt Tuning for Vision-Language Model FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 636e17f1-91bf-48be-b23d-60c25914a9b4 · inbound
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b568bdd-c956-4ff0-8550-f4dff723ec8e · inbound
Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18e71dd-2079-4c9d-97c5-8feaabc0f774 · inbound
Multimodal Medical Image Binding via Shared Text Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeff4d42-cf50-4119-b915-40f1fc14ae37 · inbound
Global and Local Entailment Learning for Natural World Imagery FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · inbound
On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc37e39e-0277-4cf0-8044-0384f944813d · inbound
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f619dc8-8c22-41ad-9e55-a49108b73bb6 · inbound
MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b7ccce-4d7c-41f6-b11a-f4bcb7231546 · inbound
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 226
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e716f882-24dc-4ea4-b2f4-4fc56b7832f5 · inbound
Adapting Vision-Language Models Without Labels: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e870abf-54be-4205-ab3e-7161bff43bc1 · inbound
AttriPrompt: Dynamic Prompt Composition Learning for CLIP FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca5c6c14-8e03-4eeb-ad10-4c9d3e4ce9cb · inbound
On the Provable Importance of Gradients for Language-Assisted Image Clustering FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dfddfc6e-3ae2-4001-9e1f-44a32fa27263 · inbound
Attention Grounded Enhancement for Visual Document Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf755d44-5954-44b6-a8a1-e62993e965a3 · inbound
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f74b3d-5bc3-4bf8-926b-8b40bd2cf6d7 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1fa4fc50-24b3-4829-ad7b-5031b6650632 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d292c2d8-336a-43b3-80d5-44e76beb8772 · inbound
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ffa69be8-970c-43ce-b060-40900efc81b8 · inbound
MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e47a07f3-5fe6-4557-860a-1b6f3f1911fe · inbound
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51116604-23c3-45d9-9f87-95a5e3786c6f · inbound
Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a8a828c-4b96-4f39-9789-f04baf8c659c · inbound
Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 81dc42ca-f89f-4d22-9fc7-e63ce45a8ce6 · inbound
Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7764adf1-8b35-47b7-850b-823adbea6315 · inbound
Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2555d6d5-57f1-4b73-bb1a-4715cd7b8f65 · inbound
Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 434480b4-1d01-4c68-a1c3-c492c53ada42 · inbound
Neutral-Reference Prompting for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e4d049c0-dbf6-4f4b-966f-21c2cd4be757 · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33708721-13b5-468a-8a68-4aed6827ca90 · inbound
Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 30cbd54a-dd56-4b46-8408-ea7436258751 · inbound
LARE: Low-Attention Region Encoding for Text-Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e0994651-052d-49b9-90dc-93bc2c9e2906 · inbound
Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a4cf79a-e396-4fed-ab52-d5483a6b56f3 · inbound
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7172ae34-e34a-4300-b63a-66b73dd6694c · inbound
SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 89035cf7-a9c1-43be-81c4-a70c68d4366d · inbound
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8392932-71c2-4c4c-acea-841098fe600c · inbound
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.