Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 74 inbound Pith citation observations for arXiv:2111.07783.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:52.478823Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:07:53.330969Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7589c8ec-0140-4b7f-81b2-ff07c748df29 · inbound
Florence: A New Foundation Model for Computer Vision FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5f7e8f8d-54ba-487d-829a-68adacc53d8b · inbound
Flamingo: a Visual Language Model for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2adb3506-7cbf-4907-86c4-9ca1be6fd63b · inbound
CoCa: Contrastive Captioners are Image-Text Foundation Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e8122669-16c3-4e8f-8f51-891e7666a1e8 · inbound
DetailCLIP: Injecting Image Details into CLIP's Feature Space FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 394409b5-09ee-46ff-b7a7-49f06cd927e2 · inbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 25df6251-3b0c-4dbc-8297-b07522936587 · inbound
VideoChat: Chat-Centric Video Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0674e91f-b242-4d8b-aa76-cbcaa4141965 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b7e3e877-1b7e-4f8f-9dab-1f673822d2e1 · inbound
LPT: Less-overfitting Prompt Tuning for Vision-Language Model FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8d34c447-ea31-418c-88bb-2644885dfc7a · inbound
AstroM$^3$: A self-supervised multimodal model for astronomy FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3981110d-85f6-4da4-a44c-0f94aad247b3 · inbound
Dissecting Representation Misalignment in Contrastive Learning via Influence Function FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08a0c0fb-f71e-41cd-9394-4ac6a690eb78 · inbound
FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ba7f40-9837-4a6a-96f4-48f3644cb71e · inbound
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283ff450-16cd-43ac-abed-4b9272363e32 · inbound
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3964313e-b3b3-41aa-a326-138ad3d039dd · inbound
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19e2ac0-7835-4940-8197-219afc716d48 · inbound
Style-Pro: Style-Guided Prompt Learning for Generalizable Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60393313-46e1-47ce-ab04-12388d1ba672 · inbound
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbd676ce-f04c-43ee-944b-1e300a7a59b0 · inbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa041bf4-5f56-4cc0-94b9-6668a172ba94 · inbound
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de2d24c-3e8f-4434-811f-b3d03c8464be · inbound
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07678994-8233-4f1b-afca-1b458ba8ff31 · inbound
FLAIR: VLM with Fine-grained Language-informed Image Representations FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76ffd2fe-4075-4158-819f-bfc6fad1f3d7 · inbound
Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f595821-b815-4d6a-a722-f67e614d1f41 · inbound
AmCLR: Unified Augmented Learning for Cross-Modal Representations FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc86bcb-35af-4f97-9c0a-d329449e0029 · inbound
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c36153d-63d1-48ad-a815-2d6ffac7da70 · inbound
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a5a1d1d-f725-477f-88ac-428285e010fb · inbound
ProtCLIP: Function-Informed Protein Multi-Modal Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3581977-98a9-40e7-8023-8914219f26dc · inbound
LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e175a0-ec5e-4e91-9d78-8d43d48b7cc9 · inbound
UniCoRN: Unified Commented Retrieval Network with LMMs FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d329b046-032a-416e-a75b-64acc28813de · inbound
Decoupled Global-Local Alignment for Improving Compositional Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3d4352-6886-499b-98fe-9d44b5b46777 · inbound
HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9404eccd-3c1c-4828-a55e-e8df4c203e64 · inbound
MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 957b594d-6f13-47fe-96d9-93b44a7a7eb9 · inbound
GeoMM: On Geodesic Perspective for Multi-modal Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774ff214-62cb-479b-8021-7e583c09f806 · inbound
CRISP: Clustering Multi-Vector Representations for Denoising and Pruning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ccb72d7-813b-4bde-aaa2-3c35c2daaa72 · inbound
DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636e17f1-91bf-48be-b23d-60c25914a9b4 · inbound
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6335e2a4-9f40-4714-829a-2ea7718e1101 · inbound
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d521d6d-a866-4b45-b9b7-f271b658c00b · inbound
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 964fa5b6-aa92-4547-acd9-5067dec726f1 · inbound
Rethinking Causal Mask Attention for Vision-Language Inference FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f290b0e3-db47-4343-b8e2-f59785e43ed0 · inbound
DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c35fb94-e3e6-4786-b4bf-a649b9ea3e60 · inbound
ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7474f0fe-0d7a-4030-81fd-dcd40343605c · inbound
Target Semantics Clustering via Text Representations for Robust Universal Domain Adaptation FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91644c81-a5f5-4f68-a16d-b45b1a3b89c3 · inbound
FREE: Fast and Robust Vision Language Models with Early Exits FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b568bdd-c956-4ff0-8550-f4dff723ec8e · inbound
Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b42301b-1592-49b0-8b5b-871882dbbb3f · inbound
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18e71dd-2079-4c9d-97c5-8feaabc0f774 · inbound
Multimodal Medical Image Binding via Shared Text Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeff4d42-cf50-4119-b915-40f1fc14ae37 · inbound
Global and Local Entailment Learning for Natural World Imagery FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · inbound
On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc37e39e-0277-4cf0-8044-0384f944813d · inbound
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f619dc8-8c22-41ad-9e55-a49108b73bb6 · inbound
MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b7ccce-4d7c-41f6-b11a-f4bcb7231546 · inbound
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 226
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e716f882-24dc-4ea4-b2f4-4fc56b7832f5 · inbound
Adapting Vision-Language Models Without Labels: A Comprehensive Survey FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e870abf-54be-4205-ab3e-7161bff43bc1 · inbound
AttriPrompt: Dynamic Prompt Composition Learning for CLIP FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca5c6c14-8e03-4eeb-ad10-4c9d3e4ce9cb · inbound
On the Provable Importance of Gradients for Language-Assisted Image Clustering FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dfddfc6e-3ae2-4001-9e1f-44a32fa27263 · inbound
Attention Grounded Enhancement for Visual Document Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bf755d44-5954-44b6-a8a1-e62993e965a3 · inbound
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f74b3d-5bc3-4bf8-926b-8b40bd2cf6d7 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1fa4fc50-24b3-4829-ad7b-5031b6650632 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d292c2d8-336a-43b3-80d5-44e76beb8772 · inbound
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ffa69be8-970c-43ce-b060-40900efc81b8 · inbound
MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e47a07f3-5fe6-4557-860a-1b6f3f1911fe · inbound
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 51116604-23c3-45d9-9f87-95a5e3786c6f · inbound
Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8a8a828c-4b96-4f39-9789-f04baf8c659c · inbound
Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 81dc42ca-f89f-4d22-9fc7-e63ce45a8ce6 · inbound
Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7764adf1-8b35-47b7-850b-823adbea6315 · inbound
Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2555d6d5-57f1-4b73-bb1a-4715cd7b8f65 · inbound
Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 434480b4-1d01-4c68-a1c3-c492c53ada42 · inbound
Neutral-Reference Prompting for Vision-Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e4d049c0-dbf6-4f4b-966f-21c2cd4be757 · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 33708721-13b5-468a-8a68-4aed6827ca90 · inbound
Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 30cbd54a-dd56-4b46-8408-ea7436258751 · inbound
LARE: Low-Attention Region Encoding for Text-Image Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e0994651-052d-49b9-90dc-93bc2c9e2906 · inbound
Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8a4cf79a-e396-4fed-ab52-d5483a6b56f3 · inbound
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7172ae34-e34a-4300-b63a-66b73dd6694c · inbound
SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 89035cf7-a9c1-43be-81c4-a70c68d4366d · inbound
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f8392932-71c2-4c4c-acea-841098fe600c · inbound
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8461ec2a-4129-4a40-a87d-e7149cd3afe5 · inbound
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.