Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:37:15.438060Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 14 inbound Pith citation observations for arXiv:2412.08802.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:37:15.438060Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:03:30.104899Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:30:07.771573Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 28433655-bb18-4a4b-aa78-c991a8a4d8e0 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc283364-e5ae-4c17-bc7a-c584773a54c8 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unsupervised Cross-lingual Representation Learning at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9059c358-c83e-478c-bff5-ecca89bc7e3a · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab468643-9f9f-42f2-89a4-e5eebadf5ef1 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Data Filtering Networks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ad892e-b79f-46f0-8a18-37d0c1ad7c71 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images ColPali: Efficient Document Retrieval with Vision Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 564b053c-546a-4b15-9cc3-2d0f596153a4 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Mistral 7B
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b75c2a7e-c72e-463b-9853-8458486d0283 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Jina CLIP: Your CLIP Model Is Also Your Text Retriever
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc086943-97d5-490f-a507-39f456c1d25d · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Matryoshka Representation Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a8b09ee-2e85-42a9-a5a2-5f987db6ec56 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images doi: 10.18653/v1/ 2024.acl-long.775
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbaad966-44ca-4a23-9b6d-da522237e392 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7692dd87-d1f6-4799-aab7-50ca735c6692 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e207a0-f5f8-452c-9bf6-736a41fec216 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Decoupled Weight Decay Regularization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3fd9621-eb0b-4d3c-8131-a5d9e4e250ba · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images MTEB: Massive Text Embedding Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b06bff-b775-4ada-9f8b-47ac0bee19db · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e37158a8-1165-4c72-bc6f-5f5478232ecf · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images doi: 10.18653/v1/2021.naacl-main.466
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91aa8421-c6da-4aef-86a8-1fee89eaaceb · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Learning Transferable Visual Models From Natural Language Supervision
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4349020-5751-42d9-aa17-5b7c0ea28b12 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a20f4d8-2722-47f9-8116-f4bec7a48711 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images jina-embeddings-v3: Multilingual Embeddings With Task LoRA
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5c3c0ae-35c5-4c62-b197-865b9ad31262 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db83f1c0-5257-4c02-a39e-54e4259800af · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 695f1236-a10c-4582-afd9-23853ea0b66d · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Representation Learning with Contrastive Predictive Coding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250fbe52-f52d-4903-a573-66d629c886bb · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c2a72f-95e1-438d-8fe3-e472c690db29 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Sigmoid Loss for Language Image Pre-Training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826e06e9-bb60-488c-ba87-b1a017e86fe2 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images URL http://dx.doi.org/10.1145/3503161.3548422
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4dbeeb-e1ea-4b1b-9b90-5245d9e4a1da · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4ca267e2-befd-4fb2-9727-d0706de30ae6 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c28bcd36-472b-4b20-a5e8-2444605feb43 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f6cd3441-e01b-4b0e-bd7b-0543ac1020d3 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dc9b4900-f8a7-4d21-bdb1-088c472d540b · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 30a51f61-333d-41db-9760-280872923a04 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b945e95-9262-4134-913c-b37b1c67f7d3 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9ecdbf1c-4da4-40b3-97ac-6830418a7be4 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 55737288-8ab8-4a9f-8cb3-acd6cc5ca755 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5302035c-553f-443e-9988-ecc1524f40dd · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0db1e746-023a-4287-a64e-85bc72203c8b · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5fff53ef-862d-44ad-9e7e-027fe326aef9 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b165dc71-12e0-4300-b34b-4302c578a031 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5fa377ac-33ef-43d7-a28a-8054d68514d6 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 42d8bc64-d8f2-4028-83bb-bc4c311fc9ba · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4312103-6f6e-4152-b7f4-b00a9030982a · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f1d63ec3-a23e-4dc6-93d1-17c014a17e7f · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images URL https://aclanthology.org/Q14-1006
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d205eae-7313-4eda-a2de-409606e5105d · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a192aaa9-990f-49f2-86ff-2b99f27be51b · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3ea58816-7331-405f-bffe-61faed31980b · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images NLLB-CLIP -- train performant multilingual image retrieval model on a budget
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e923487-1318-4e60-9efc-1921fb78ca44 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images doi: 10.18653/v1/K19-1049
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0adfd5c6-9845-409a-b5fe-2d5d40a64969 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Rethinking benchmarks for cross-modal image-text retrieval
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f658bfd-b45c-4a0a-a375-191d3ce971ab · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images LoRA: Low-Rank Adaptation of Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7872148d-6719-45d8-a895-9308a75c46a2 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46aae501-ac6a-4601-a235-b9969bd1f334 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images DataComp: In search of the next generation of multimodal datasets
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8172c23d-fd52-4830-8cfd-3904fed8f058 · outbound
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a37951d5-4c65-4996-8dbf-85523f262b56 · inbound
MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c233abe-685d-4479-aed1-9faae35b6db7 · inbound
RGB-Pointmap Pretraining for Unified 3D Scene Understanding jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f5186e79-b65c-4286-91ea-410d98f875a5 · inbound
HIVE: Query, Hypothesize, Verify An LLM Framework for Multimodal Reasoning-Intensive Retrieval jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6c5de1f-0bdd-4e19-893a-9e758e94f51b · inbound
jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8b86ce07-2594-4b42-96d6-f4e48c6c4bd3 · inbound
jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fb41aea0-70b2-483a-a861-eaea954a8b99 · inbound
jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f1cf3dec-3d74-43a2-8536-ce16658b0faf · inbound
MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9f548490-1d6f-4406-8dfd-aa5baf953dfb · inbound
One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 20b3544d-2409-4544-8762-e46faac2b04a · inbound
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 67de105b-4d74-4a96-86cb-b0b461385c3a · inbound
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 994aec71-e745-4473-8b9b-2de9c6dc08df · inbound
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5530c45c-12d9-46d5-bb05-c6171f617c16 · inbound
KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ae2735-f4c0-4927-ba64-3ea619c019e7 · inbound
Illuminating Visual Identity in Universal Multimodal Embeddings jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d832f383-4bd3-4833-acf3-633ec92e1182 · inbound
MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.