Pith. sign in

Paper Citation Record · LEDGER

FILIP: Fine-grained Interactive Language-Image Pre-Training

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 74 inbound Pith citation observations for arXiv:2111.07783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.07783 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 74 of 74 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:52.478823Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:07:53.330969Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7589c8ec-0140-4b7f-81b2-ff07c748df29 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.582884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:c93dca1c005858a59d617c13e261a917f0ec56a325ad2d182743fcec810d10db

Observation 5f7e8f8d-54ba-487d-829a-68adacc53d8b · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 139

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.453088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:d3747899b4d4455b92b5dda1c0f9ce5779ac5022772108dd85a1199754314413

Observation 2adb3506-7cbf-4907-86c4-9ca1be6fd63b · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:53:08.484747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:d8b14ebe3a777aeacf72b9267b256ea7dadedd5117c933bceafc75c34a3d43d1

Observation e8122669-16c3-4e8f-8f51-891e7666a1e8 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.391885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:8c841ed765f0db6b7cd15afd595c2f68c1f347c916aa551de4e281c3d4284f78

Observation 394409b5-09ee-46ff-b7a7-49f06cd927e2 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.293435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:789af09713f8f4e48ab82a9608d013442a2f65c2d84f7d5f04d6a6e820d639f2

Observation 25df6251-3b0c-4dbc-8297-b07522936587 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.687974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:cdb32723d667d9b9a21edfd56020ddb0ae9dac42ef78032de35c658c85b4fadc

Observation 0674e91f-b242-4d8b-aa76-cbcaa4141965 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.575279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:052c8b61c32841565a5699ced90ba25cb0118be168388aaea2180eb837c6d6aa

Observation b7e3e877-1b7e-4f8f-9dab-1f673822d2e1 · inbound

LPT: Less-overfitting Prompt Tuning for Vision-Language Model cites this paper.

LPT: Less-overfitting Prompt Tuning for Vision-Language Model FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:05:47.187508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T19:03:25.106193Z digest=sha256:2c65e353d4b9648802dc0577b40b8c79de3f153d32a7a7be4662442c7820294c

Observation 8d34c447-ea31-418c-88bb-2644885dfc7a · inbound

AstroM$^3$: A self-supervised multimodal model for astronomy cites this paper.

AstroM$^3$: A self-supervised multimodal model for astronomy FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:23:18.884168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:23:18.884168Z digest=sha256:9a9e536d5a417fec2d30a29ee65d2d4300795fcd7a60e5cce0b0e747bfc0f21a

Observation 3981110d-85f6-4da4-a44c-0f94aad247b3 · inbound

Dissecting Representation Misalignment in Contrastive Learning via Influence Function cites this paper.

Dissecting Representation Misalignment in Contrastive Learning via Influence Function FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:22:17.180042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:22:17.180042Z digest=sha256:4aeddf723991012b4d20ef9568a2fef0abd993a34be3dc5701eb5da6d5d394e5

Observation 08a0c0fb-f71e-41cd-9394-4ac6a690eb78 · inbound

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training cites this paper.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.029722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.029722Z digest=sha256:447d5eb054cc8fb467682b65a14574ec555a2e1819f93d2e6282eaf42d8fdb2f

Observation 95ba7f40-9837-4a6a-96f4-48f3644cb71e · inbound

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers cites this paper.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.127302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.127302Z digest=sha256:5bfb7fc37fc66b13d523541344ee6bf5c5d5cf66edbe67ceefd06d0eaa143304

Observation 283ff450-16cd-43ac-abed-4b9272363e32 · inbound

Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training cites this paper.

Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:50:17.587771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:50:17.587771Z digest=sha256:cb3107dcfb0be8c5dde8dc8142b43856643223fa247c4711ea8ad0854ccaaa16

Observation 3964313e-b3b3-41aa-a326-138ad3d039dd · inbound

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference cites this paper.

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T13:53:33.566440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:53:33.566440Z digest=sha256:2f03e1757f1aead45e36d5557e43b2971f1fa9d603281ac695a765c5c39a62f0

Observation b19e2ac0-7835-4940-8197-219afc716d48 · inbound

Style-Pro: Style-Guided Prompt Learning for Generalizable Vision-Language Models cites this paper.

Style-Pro: Style-Guided Prompt Learning for Generalizable Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T13:43:17.237195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:43:17.237195Z digest=sha256:5051a119ef8472eea20b253d9f898f1e3a48c3ed5f96a88ea9304fa11002951b

Observation 60393313-46e1-47ce-ab04-12388d1ba672 · inbound

Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation cites this paper.

Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:38.781663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:38.781663Z digest=sha256:bdda33bf1f07182968b4d4565cb1cb6fcda77a42f108f6b920b87f47acf037ed

Observation dbd676ce-f04c-43ee-944b-1e300a7a59b0 · inbound

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions cites this paper.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.027307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.027307Z digest=sha256:863dbdf5f1dbe4c7af738918e22a66e7b0b4b32af191f78b3ef2d3797194e949

Observation aa041bf4-5f56-4cc0-94b9-6668a172ba94 · inbound

CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections cites this paper.

CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T10:20:26.899763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:20:26.899763Z digest=sha256:bb89cb7ea764a9b704cc93c351f9b121a87686d908edd9b079203025372fa06b

Observation 2de2d24c-3e8f-4434-811f-b3d03c8464be · inbound

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives cites this paper.

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:42.952416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:42.952416Z digest=sha256:b4f6affcce2af4659b1ffc018389c00130f2af382ffea62584aceef713d572b1

Observation 07678994-8233-4f1b-afca-1b458ba8ff31 · inbound

FLAIR: VLM with Fine-grained Language-informed Image Representations cites this paper.

FLAIR: VLM with Fine-grained Language-informed Image Representations FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:06.895716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:06.895716Z digest=sha256:3a7b98148a20f006fe6f3f395c4601c45e0da6d9fa12e35bdff545ee9078b93f

Observation 76ffd2fe-4075-4158-819f-bfc6fad1f3d7 · inbound

Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling cites this paper.

Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T19:15:01.549245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:15:01.549245Z digest=sha256:66ff92c67926d195e80ad6134642d275c599720001080b5a4f3b8ac91de69a64

Observation 9f595821-b815-4d6a-a722-f67e614d1f41 · inbound

AmCLR: Unified Augmented Learning for Cross-Modal Representations cites this paper.

AmCLR: Unified Augmented Learning for Cross-Modal Representations FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T18:24:37.422289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:24:37.422289Z digest=sha256:82a5eb1d54a4739cd02dfc567c51850b7902e12aa897faec1d5eb326ac6e3f71

Observation 7dc86bcb-35af-4f97-9c0a-d329449e0029 · inbound

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey cites this paper.

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:54.315913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:54.315913Z digest=sha256:31b249927cd8425db42d04d70ec8f7d576229be967ef189f807630c1438cccc8

Observation 2c36153d-63d1-48ad-a815-2d6ffac7da70 · inbound

Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves cites this paper.

Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:55:41.025736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:55:41.025736Z digest=sha256:d6d2fcb212143097902b6b56a4fdba62e4505d15a983f6f2ec7f57be73e23ea6

Observation 4a5a1d1d-f725-477f-88ac-428285e010fb · inbound

ProtCLIP: Function-Informed Protein Multi-Modal Learning cites this paper.

ProtCLIP: Function-Informed Protein Multi-Modal Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:45:12.270436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:45:12.270436Z digest=sha256:4706b95af5ab12808a3383bb784cc0cf2d588e3bd180dbe68ca9efe3ac82c640

Observation c3581977-98a9-40e7-8023-8914219f26dc · inbound

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models cites this paper.

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T13:31:23.611194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:31:23.611194Z digest=sha256:742a8b13437b3acfee79d114ac2eaf72170abc2160d29bb808034772aa7a2a66

Observation 44e175a0-ec5e-4e91-9d78-8d43d48b7cc9 · inbound

UniCoRN: Unified Commented Retrieval Network with LMMs cites this paper.

UniCoRN: Unified Commented Retrieval Network with LMMs FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.111778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.111778Z digest=sha256:d58b931d1a10cf34e3cbdeab9254bcb85b5bb8e78bc452a83d29a3cc95c73384

Observation d329b046-032a-416e-a75b-64acc28813de · inbound

Decoupled Global-Local Alignment for Improving Compositional Understanding cites this paper.

Decoupled Global-Local Alignment for Improving Compositional Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:52.478823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:59:52.478823Z digest=sha256:186828b44c8a70b6bbab8a08c54fe537c2ccd57d84eb48839a2678e77ba3ab66

Observation ff3d4352-6886-499b-98fe-9d44b5b46777 · inbound

HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights cites this paper.

HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T23:24:46.762569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:24:46.762569Z digest=sha256:029026c15e244946549e92a201a5ff34832d412cbcc81b79e8f7076e1de1f2b2

Observation 9404eccd-3c1c-4828-a55e-e8df4c203e64 · inbound

MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models cites this paper.

MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:22:01.598766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:22:01.598766Z digest=sha256:ef191954cbe4b0a7f51166cfd96d802f0135594c245fd8247a52c2859420be7e

Observation 957b594d-6f13-47fe-96d9-93b44a7a7eb9 · inbound

GeoMM: On Geodesic Perspective for Multi-modal Learning cites this paper.

GeoMM: On Geodesic Perspective for Multi-modal Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.868708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.868708Z digest=sha256:a096166d813602667226ae5f8a0ea222ab972e87e458d8e91af7d826f44dde69

Observation 774ff214-62cb-479b-8021-7e583c09f806 · inbound

CRISP: Clustering Multi-Vector Representations for Denoising and Pruning cites this paper.

CRISP: Clustering Multi-Vector Representations for Denoising and Pruning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:08.963598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:56:08.963598Z digest=sha256:b86ca3a220f2526782ee2a47aec895b41a5c34d17f8425fac9869ff15b830b11

Observation 7ccb72d7-813b-4bde-aaa2-3c35c2daaa72 · inbound

DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation cites this paper.

DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:55:24.657625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:55:24.657625Z digest=sha256:73903cf153a3ae3e07c8af5ba5d7027ec1af17f2c0f3deb3e2e930965b215226

Observation 636e17f1-91bf-48be-b23d-60c25914a9b4 · inbound

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping cites this paper.

Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:41:36.652595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T13:38:17.640841Z digest=sha256:397a4635faab50709b9f0b0426aaa516330c5a0b7750f8b644bef64f08018d46

Observation 6335e2a4-9f40-4714-829a-2ea7718e1101 · inbound

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification cites this paper.

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:12.558456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:12.558456Z digest=sha256:06ee81afc32a448536779317a0f8896b728ac6450b78e15d1d4b5b2bcdfe8f04

Observation 9d521d6d-a866-4b45-b9b7-f271b658c00b · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:15.993448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:15.993448Z digest=sha256:79a472f7d55932e1ea32a04f04715640c00172a5a4cc5dbcbcf7d0afc329d74e

Observation 964fa5b6-aa92-4547-acd9-5067dec726f1 · inbound

Rethinking Causal Mask Attention for Vision-Language Inference cites this paper.

Rethinking Causal Mask Attention for Vision-Language Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:33.125369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:33.125369Z digest=sha256:253e6aa08574840d440e681c0281c4782eec00bc4f62a8a7f7d0dfe61cf84913

Observation f290b0e3-db47-4343-b8e2-f59785e43ed0 · inbound

DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models cites this paper.

DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:07.729511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:07.729511Z digest=sha256:1d14db1a774298cb176f5f18a3a6c9fc2f7e880546669349ab661b339d625a18

Observation 2c35fb94-e3e6-4786-b4bf-a649b9ea3e60 · inbound

ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval cites this paper.

ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:19.694499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:19.694499Z digest=sha256:96eb1de3fc506ecf640b6b5e310dc9e56fd9a5e3673aab65f576cda0dc4d1eb7

Observation 7474f0fe-0d7a-4030-81fd-dcd40343605c · inbound

Target Semantics Clustering via Text Representations for Robust Universal Domain Adaptation cites this paper.

Target Semantics Clustering via Text Representations for Robust Universal Domain Adaptation FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:30.756023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:06:30.756023Z digest=sha256:846c8c8da565d3573c74dfa16247bdb85d23ba432b778ad9df1d00eef35a9b7a

Observation 91644c81-a5f5-4f68-a16d-b45b1a3b89c3 · inbound

FREE: Fast and Robust Vision Language Models with Early Exits cites this paper.

FREE: Fast and Robust Vision Language Models with Early Exits FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:46.451945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:53:46.451945Z digest=sha256:e69071c11886844f9f8fc05273a324bb721a7466db5a5951687a4a6f735dfbda

Observation 8b568bdd-c956-4ff0-8550-f4dff723ec8e · inbound

Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency cites this paper.

Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:47:50.404759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:47:50.404759Z digest=sha256:bdcf2bb891d41414566a2e2e1f2a46b55f43662bafa476c9cfdd86e2503c9d5c

Observation 7b42301b-1592-49b0-8b5b-871882dbbb3f · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:19:04.788682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:19:04.788682Z digest=sha256:3808fe7efa9157c28da3fc76d41b4eb4f4b2dba2deb1fbfd735cf452176a60ba

Observation e18e71dd-2079-4c9d-97c5-8feaabc0f774 · inbound

Multimodal Medical Image Binding via Shared Text Embeddings cites this paper.

Multimodal Medical Image Binding via Shared Text Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:12.914424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:12.914424Z digest=sha256:d63b4d44ff5c24e384848848c64eb7ef069c774c72065fb99a1da7f841db01a1

Observation aeff4d42-cf50-4119-b915-40f1fc14ae37 · inbound

Global and Local Entailment Learning for Natural World Imagery cites this paper.

Global and Local Entailment Learning for Natural World Imagery FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:33:35.382549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:33:35.382549Z digest=sha256:0c4c8af6c2206636be9d64b7425ed786a44b2aef2d9ddf806a31b1849857c820

Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.667082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.667082Z digest=sha256:c5307b7f6230a3a507424fac078cc20a65dba21d1e40ab32dca41453791b659d

Observation bc37e39e-0277-4cf0-8044-0384f944813d · inbound

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text cites this paper.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.156463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.156463Z digest=sha256:07785e0c8f18118899854811bc110d6dda2d9c01959daeb75911989900af59eb

Observation 1f619dc8-8c22-41ad-9e55-a49108b73bb6 · inbound

MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning cites this paper.

MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:26:42.549669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:26:42.549669Z digest=sha256:7895d820db1d7df9b801588ebbead63a80d403bc1ebfd6a61f1b9a097a697db0

Observation 69b7ccce-4d7c-41f6-b11a-f4bcb7231546 · inbound

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey cites this paper.

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-06T11:20:07.035625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:20:07.035625Z digest=sha256:b46c2bc02ee4590e24e2583585482f4c2883b9a9ad24d47d4d54f1caa453dc47

Observation e716f882-24dc-4ea4-b2f4-4fc56b7832f5 · inbound

Adapting Vision-Language Models Without Labels: A Comprehensive Survey cites this paper.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:05.859329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:05.859329Z digest=sha256:2db25bf7eef29e5bf96c00c087c32d70a7e2f06d528579225f57efbc4b9789f0

Observation 7e870abf-54be-4205-ab3e-7161bff43bc1 · inbound

AttriPrompt: Dynamic Prompt Composition Learning for CLIP cites this paper.

AttriPrompt: Dynamic Prompt Composition Learning for CLIP FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:31.087108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:31.087108Z digest=sha256:918096fa58d45be2885539e374dd17c7a698858f0a35654fbfd49722c6aa8338

Observation ca5c6c14-8e03-4eeb-ad10-4c9d3e4ce9cb · inbound

On the Provable Importance of Gradients for Language-Assisted Image Clustering cites this paper.

On the Provable Importance of Gradients for Language-Assisted Image Clustering FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:35:29.635355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T07:30:36.437935Z digest=sha256:71980f47601b8d86d76c5e981988cda5a03adb01e3ad13e2df3821ada230f980

Observation dfddfc6e-3ae2-4001-9e1f-44a32fa27263 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.314742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:cfba33827d134fac7df0e4ccceab9ea37088aba1cc167f9de995eb58532755b5

Observation bf755d44-5954-44b6-a8a1-e62993e965a3 · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.073555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.073555Z digest=sha256:40050c78d5a6abad4cfbf003d8754965f8424825ebbdaf814f1cc71e5b853199

Observation a7f74b3d-5bc3-4bf8-926b-8b40bd2cf6d7 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.577358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:3f8544e4d1dd36b1f3132cdac5e371ba29d88debc9d999b76bbae5952eccffce

Observation 1fa4fc50-24b3-4829-ad7b-5031b6650632 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:14e7c779ca6cc7b2903a8847037ee6f3bfa58d8106f93a4f5b76175f3021a0e9

Observation d292c2d8-336a-43b3-80d5-44e76beb8772 · inbound

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference cites this paper.

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:26:01.808588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:38:00.522094Z digest=sha256:649c3c5919145144943d2952a362325e7f18aa6117dbf0983d63c67dd84b4aa2

Observation ffa69be8-970c-43ce-b060-40900efc81b8 · inbound

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images cites this paper.

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.375166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:33:02.946839Z digest=sha256:d299aa0ba206d7c0978de73ea22aafd15f24a84a4ebdabb9435b363d757fd834

Observation e47a07f3-5fe6-4557-860a-1b6f3f1911fe · inbound

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval cites this paper.

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:50:19.936226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T11:50:00.850857Z digest=sha256:18c441eb24d520df2c2c3215f1b0f718d285e4014f5802cccf6db452f1b0c64e

Observation 51116604-23c3-45d9-9f87-95a5e3786c6f · inbound

Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning cites this paper.

Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:08.369925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T17:41:05.464502Z digest=sha256:da42805eda3e098622184a0ccead318e0b77ad636a67961d5bcae8b91b99ea1d

Observation 8a8a828c-4b96-4f39-9789-f04baf8c659c · inbound

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search cites this paper.

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:07.806567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T13:39:20.777029Z digest=sha256:e8182cbf042170dc0ca781a371c5db2a32f02be55c8d5ddaeafd51a968516421

Observation 81dc42ca-f89f-4d22-9fc7-e63ce45a8ce6 · inbound

Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference cites this paper.

Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:27.287942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:35:05.535416Z digest=sha256:56edeb59b6cfd22ebf4f9a66efbfb0f28d6403bbc63b4dd68a782370221577e2

Observation 7764adf1-8b35-47b7-850b-823adbea6315 · inbound

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection cites this paper.

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.251633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:25:13.709254Z digest=sha256:0fbcdfcd5678091c87f950463815e663dccdc158669816e60be794af102a0af9

Observation 2555d6d5-57f1-4b73-bb1a-4715cd7b8f65 · inbound

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models cites this paper.

Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:22.557394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:56:31.648428Z digest=sha256:2682719e6ac5f64fc01617075216dde6ef08754e6ab13c4fa2cc119a3dc05ef5

Observation 434480b4-1d01-4c68-a1c3-c492c53ada42 · inbound

Neutral-Reference Prompting for Vision-Language Models cites this paper.

Neutral-Reference Prompting for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.570655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T19:15:54.152498Z digest=sha256:e02cfbc7dd3b8caae47cec163d821f6a761db0366a826d45774a4a5010f037eb

Observation e4d049c0-dbf6-4f4b-966f-21c2cd4be757 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.219272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:59084f3d05640a00bd43f746c414ef1de58714a416b9d0291d17cb290269f845

Observation 33708721-13b5-468a-8a68-4aed6827ca90 · inbound

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models cites this paper.

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.418341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:40:26.803098Z digest=sha256:3fdd99cc6eefad271eb6352eae48c1ee183887f9b1126de1c7a0f145c8ef2bf8

Observation 30cbd54a-dd56-4b46-8408-ea7436258751 · inbound

LARE: Low-Attention Region Encoding for Text-Image Retrieval cites this paper.

LARE: Low-Attention Region Encoding for Text-Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:09:15.213914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T21:25:12.373068Z digest=sha256:3c85dab00a97777fc678664a3e6f6da8a81d7fff4d86b85a7bfe6bbe9abc6565

Observation e0994651-052d-49b9-90dc-93bc2c9e2906 · inbound

Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings cites this paper.

Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:39:49.797760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T06:16:51.174659Z digest=sha256:5b10e8337bb35c4f9771d910970a5c483fc5a70beb52755455983d69c75eff11

Observation 8a4cf79a-e396-4fed-ab52-d5483a6b56f3 · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.418042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:d573876813119a806521a9822ba7718abf6c9a3f670204527a306ac6c25bfbb7

Observation 7172ae34-e34a-4300-b63a-66b73dd6694c · inbound

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs cites this paper.

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:07:53.355970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-11T02:59:38.193627Z digest=sha256:402985f73ce623238d8fa6605765ce721d4e7923729669be1c5b617a1a7f6081

Observation 89035cf7-a9c1-43be-81c4-a70c68d4366d · inbound

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks cites this paper.

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T15:47:23.353134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-10T15:38:58.361411Z digest=sha256:7e5203ac45db3cb7b62981a81de135994be7d5d07ba780eaa1504e50f7f8b880

Observation f8392932-71c2-4c4c-acea-841098fe600c · inbound

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision cites this paper.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.969697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.969697Z digest=sha256:029a197cc443b779bbeccc96a78128fe40dd45c1dc4747acdcd2c921646350aa

Observation 8461ec2a-4129-4a40-a87d-e7149cd3afe5 · inbound

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval cites this paper.

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:25.491006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:25.491006Z digest=sha256:4448b0a4fed289c185607160eba537039347fb97319ca533560140b7e25d7a2e