Pith. sign in

Paper Citation Record · LEDGER

FILIP: Fine-grained Interactive Language-Image Pre-Training

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2111.07783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.07783 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:47:50.404759Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:07:53.330969Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7589c8ec-0140-4b7f-81b2-ff07c748df29 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.582884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:71bb1a7e99814e6b3d9a25cf144014272a736203168625c295afd60aa0bc12a0

Observation 5f7e8f8d-54ba-487d-829a-68adacc53d8b · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.453088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:cf9b5feaf8dbd845b10b19a23a2b4ea047de76e908cc6cc04b93d557f5e7fc49

Observation 2adb3506-7cbf-4907-86c4-9ca1be6fd63b · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:53:08.484747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:ab5e4ae15946d25a3d618617cd35933b4eea9270d04dfb887d06edb823883666

Observation e8122669-16c3-4e8f-8f51-891e7666a1e8 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.391885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:85320fca0acdb84021b73f32e2a95e5627d13fe9e49c2c849904d85e2e4a0058

Observation 394409b5-09ee-46ff-b7a7-49f06cd927e2 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.293435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:a37743ea53c3ad6c74f175a4991d642f659b24fdf622ab2419ac49108b759dcc

Observation 25df6251-3b0c-4dbc-8297-b07522936587 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.687974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:c005e361a04df258ae9961c2ab6c016f54196a557f2e1cdff595c55290f870bb

Observation 0674e91f-b242-4d8b-aa76-cbcaa4141965 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.575279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:3af745e8b4bf06eeae4187f2a71fa74fcb9a0103c6a73874c7554a09b0a9b6fa

Observation b7e3e877-1b7e-4f8f-9dab-1f673822d2e1 · inbound

LPT: Less-overfitting Prompt Tuning for Vision-Language Model cites this paper.

LPT: Less-overfitting Prompt Tuning for Vision-Language Model FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:47.187508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:03:25.106193Z digest=sha256:9446c6bac6054d5f03d81fddebd95024f219f2dc05c5b6496573e474f52ca8a5

Observation 636e17f1-91bf-48be-b23d-60c25914a9b4 · inbound

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping cites this paper.

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:41:36.652595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T13:38:17.640841Z digest=sha256:35ae6185c5ea6d4767fb07c4d965fdc153cd7bf2a728f49cb697a4be42e57002

Observation 8b568bdd-c956-4ff0-8550-f4dff723ec8e · inbound

Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency cites this paper.

Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:47:50.404759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:47:50.404759Z digest=sha256:0b6a3ac65b4e6f0e46e6a2ee68fe1d079b50fd3bb843973749049dfd273f4703

Observation e18e71dd-2079-4c9d-97c5-8feaabc0f774 · inbound

Multimodal Medical Image Binding via Shared Text Embeddings cites this paper.

Multimodal Medical Image Binding via Shared Text Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:12.914424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:12.914424Z digest=sha256:3d9270744bec2a9d29e4d0638e5b028ae685de04cdb19dc6b69fa962a24eb3d5

Observation aeff4d42-cf50-4119-b915-40f1fc14ae37 · inbound

Global and Local Entailment Learning for Natural World Imagery cites this paper.

Global and Local Entailment Learning for Natural World Imagery FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:33:35.382549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:33:35.382549Z digest=sha256:9d2648b0e041b776dcdd29ce39656034d5ada4b641c0ae8d4b36fae8b53d1f21

Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.667082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.667082Z digest=sha256:e33527345f0f2682838014ab6b5aa3b0bc8f2154e558b80800610157b06ea2ea

Observation bc37e39e-0277-4cf0-8044-0384f944813d · inbound

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text cites this paper.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.156463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.156463Z digest=sha256:0d98b61b96b7625596375b8e121ac04dcd8941942f82c38c0aaa5f400a3bb5aa

Observation 1f619dc8-8c22-41ad-9e55-a49108b73bb6 · inbound

MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning cites this paper.

MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:42.549669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:42.549669Z digest=sha256:525f1ddae54722778c6836d7bece12509214eef120c523cf2066daa75b45e80f

Observation 69b7ccce-4d7c-41f6-b11a-f4bcb7231546 · inbound

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey cites this paper.

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-06T11:20:07.035625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:20:07.035625Z digest=sha256:be5d125850137a7362420b836f482bfb2f4db2c2efc4165433795f860b0e05bb

Observation e716f882-24dc-4ea4-b2f4-4fc56b7832f5 · inbound

Adapting Vision-Language Models Without Labels: A Comprehensive Survey cites this paper.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.859329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.859329Z digest=sha256:0f8d2df3ec031375a59c832bed87915b1d95c004438cbc648bbe36df0d4b221e

Observation 7e870abf-54be-4205-ab3e-7161bff43bc1 · inbound

AttriPrompt: Dynamic Prompt Composition Learning for CLIP cites this paper.

AttriPrompt: Dynamic Prompt Composition Learning for CLIP FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:31.087108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:31.087108Z digest=sha256:5ef38c6a38886d35b051d8745e847e680ee7d33d6f632263d08b53bd215898cd

Observation ca5c6c14-8e03-4eeb-ad10-4c9d3e4ce9cb · inbound

On the Provable Importance of Gradients for Language-Assisted Image Clustering cites this paper.

On the Provable Importance of Gradients for Language-Assisted Image Clustering FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:35:29.635355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T07:30:36.437935Z digest=sha256:cd685929f96de73dfd924a9f5f2c33f2cb9a9ad2e79aa61079472c3569bbf1f1

Observation dfddfc6e-3ae2-4001-9e1f-44a32fa27263 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.314742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:7c739bae29f83bca3bc65675810b2a0ae62173d41666d6da38a759e7761fa886

Observation bf755d44-5954-44b6-a8a1-e62993e965a3 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.073555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.073555Z digest=sha256:72c2c3172a88c1e206b3d117be1b39a0c2a070edc03c6bffea664c729e6ee3ee

Observation a7f74b3d-5bc3-4bf8-926b-8b40bd2cf6d7 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.577358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:7ae2abf14af663953299d5d26d69d85a14edc90034dea9fe760da4bff0bacd9b

Observation 1fa4fc50-24b3-4829-ad7b-5031b6650632 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:5abea0b269de5c4703629935a175ed13260ee58066fa6835be0abbd5a2556ab0

Observation d292c2d8-336a-43b3-80d5-44e76beb8772 · inbound

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference cites this paper.

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:26:01.808588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:38:00.522094Z digest=sha256:5e1d06667123892f3867976785196325b087cef46d695b25823db1d415ae0af6

Observation ffa69be8-970c-43ce-b060-40900efc81b8 · inbound

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images cites this paper.

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.375166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:33:02.946839Z digest=sha256:62b5b2fc42d6a0efd4b2ad6df7a96e0b195022a7bf74123a3ba408d182c48ccb

Observation e47a07f3-5fe6-4557-860a-1b6f3f1911fe · inbound

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval cites this paper.

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:50:19.936226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:50:00.850857Z digest=sha256:0108706a696773275dc42334fad7b42a1990a9878fea2e29c8065947dda2db21

Observation 51116604-23c3-45d9-9f87-95a5e3786c6f · inbound

Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning cites this paper.

Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:08.369925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T17:41:05.464502Z digest=sha256:800c633f40545d54c8a1ac46130343bb95612a82a5644f98cd6617fa2a8cd8b0

Observation 8a8a828c-4b96-4f39-9789-f04baf8c659c · inbound

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search cites this paper.

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.806567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T13:39:20.777029Z digest=sha256:f7e42cf657d4f44af28ec0a01ecdc523c1f42a4e8caed04eb675d686139507ce

Observation 81dc42ca-f89f-4d22-9fc7-e63ce45a8ce6 · inbound

Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference cites this paper.

Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:27.287942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T02:35:05.535416Z digest=sha256:09917ffafbcb350a65d10f32bf9298f023cc155977301dffc3ebaad66026831b

Observation 7764adf1-8b35-47b7-850b-823adbea6315 · inbound

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection cites this paper.

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.251633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:25:13.709254Z digest=sha256:076ddfe42c2965f387c53f2ed647479de259acb5fc72927557bfdb32f1652437

Observation 2555d6d5-57f1-4b73-bb1a-4715cd7b8f65 · inbound

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models cites this paper.

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:22.557394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:56:31.648428Z digest=sha256:3cec5b3127c21a4aad5fbf329b37a98d1886704ca3b9c12cc1cd627676ad1097

Observation 434480b4-1d01-4c68-a1c3-c492c53ada42 · inbound

Neutral-Reference Prompting for Vision-Language Models cites this paper.

Neutral-Reference Prompting for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.570655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T19:15:54.152498Z digest=sha256:91c93a788fe04c9bda97777fbf3b38e5d538cbe3bb743ae222fa9cd966c0cc70

Observation e4d049c0-dbf6-4f4b-966f-21c2cd4be757 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.219272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:208b445e237e7730bdfb83ae7d8bc21126f09ef67463f9eb2371876684a9e034

Observation 33708721-13b5-468a-8a68-4aed6827ca90 · inbound

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models cites this paper.

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.418341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:40:26.803098Z digest=sha256:fcb3fbe6cadb56ff6153827e6be027302e8b476a072b834c24e539bf71d1ac8c

Observation 30cbd54a-dd56-4b46-8408-ea7436258751 · inbound

LARE: Low-Attention Region Encoding for Text-Image Retrieval cites this paper.

LARE: Low-Attention Region Encoding for Text-Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:09:15.213914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T21:25:12.373068Z digest=sha256:f39439ded16b2a256fae55823b25572e31bf3768f3e80b9ffd7e2a1841b61895

Observation e0994651-052d-49b9-90dc-93bc2c9e2906 · inbound

Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings cites this paper.

Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:39:49.797760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T06:16:51.174659Z digest=sha256:44b6d95d1241a1b9a5ee78e4555bf5c49d11fdb725cd40cf9654deefc4747f64

Observation 8a4cf79a-e396-4fed-ab52-d5483a6b56f3 · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.418042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:acf64794c3c94f2a6709384bad5457c26956c0cf05028cbdc7b46e184098219b

Observation 7172ae34-e34a-4300-b63a-66b73dd6694c · inbound

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs cites this paper.

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:07:53.355970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T02:59:38.193627Z digest=sha256:e90cfd9b194b9d85d29c9538fd19b999bace799bf36b9fa3a02829d0452fe1f3

Observation 89035cf7-a9c1-43be-81c4-a70c68d4366d · inbound

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks cites this paper.

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T15:47:23.353134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-10T15:38:58.361411Z digest=sha256:6bc816b209e2b2a5ac304bc7dff1f17166407dc7bc8387e7bed2e6823d27a523

Observation f8392932-71c2-4c4c-acea-841098fe600c · inbound

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision cites this paper.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.969697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.969697Z digest=sha256:13148476af2de19bdfd1b52807c0c5473a0586509dba68d516179058177a779f