Pith. sign in

Paper Citation Record · LEDGER

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 14 inbound Pith citation observations for arXiv:2412.08802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08802 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:37:15.438060Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:03:30.104899Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.771573Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved36
  • parse uncertain0
  • malformed identifier10
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28433655-bb18-4a4b-aa78-c991a8a4d8e0 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.225863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.225863Z digest=sha256:68469c4d8cca765d3136ab5895311af63c94d3a00ddbe191753d190e4ab39824

Observation fc283364-e5ae-4c17-bc7a-c584773a54c8 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unsupervised Cross-lingual Representation Learning at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.239351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.239351Z digest=sha256:0ecf40db780c492cc47ca896f5c518f6c341e2f30d008319911f804dfe5930bd

Observation 9059c358-c83e-478c-bff5-ecca89bc7e3a · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.244438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.244438Z digest=sha256:0217530fc4deeb438133806c8725860ac18702224c1c64fed122617631fa4e49

Observation ab468643-9f9f-42f2-89a4-e5eebadf5ef1 · outbound

This paper cites Data Filtering Networks.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Data Filtering Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.249390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.249390Z digest=sha256:e6b3ca0b4eb81d224ecc2c918d86fe212acea5b9e8e2754c6fc6b65038aa9637

Observation 27ad892e-b79f-46f0-8a18-37d0c1ad7c71 · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images ColPali: Efficient Document Retrieval with Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.253813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.253813Z digest=sha256:5108d4498919ef95113012dc5cfd99156358ec48a2cd0eb1a1b858a0944c0e19

Observation 564b053c-546a-4b15-9cc3-2d0f596153a4 · outbound

This paper cites Mistral 7B.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.271878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.271878Z digest=sha256:f71b3fe796245cb92fdd717fa9675d749f495d52a5db90b9d86ea25cd2ae3c23

Observation b75c2a7e-c72e-463b-9853-8458486d0283 · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.276535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.276535Z digest=sha256:e60ce4e85eadd40c5a2c13d879ca33b182dc8e07fc5a14d63b34410665ea91ae

Observation fc086943-97d5-490f-a507-39f456c1d25d · outbound

This paper cites Matryoshka Representation Learning.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Matryoshka Representation Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.281003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.281003Z digest=sha256:372ac81a6c59ca925eed1000587380b6a801346f2103c4c952561c312fd6afa4

Observation 1a8b09ee-2e85-42a9-a5a2-5f987db6ec56 · outbound

This paper cites doi: 10.18653/v1/ 2024.acl-long.775.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images doi: 10.18653/v1/ 2024.acl-long.775

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-11T17:37:15.285204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.285204Z digest=sha256:dc1c9d4bd881367d287946a86e6d2c83944316583ad12c918c941eb7c2cba3f6

Observation fbaad966-44ca-4a23-9b6d-da522237e392 · outbound

This paper cites Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.289873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.289873Z digest=sha256:f45244e9de6d679f5e04cca458a079edf85f1313ebe9197e8eb9f501ea43098d

Observation 7692dd87-d1f6-4799-aab7-50ca735c6692 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.294739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.294739Z digest=sha256:7f0d33de1c0923721bcc4d951e6fe2541a311be5c7c233d7223946847fe2da32

Observation 76e207a0-f5f8-452c-9bf6-736a41fec216 · outbound

This paper cites Decoupled Weight Decay Regularization.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Decoupled Weight Decay Regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.299504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.299504Z digest=sha256:0da680bec826970715bc9a0a8e49dce61bfab0af4bbb4d977365379228dc7b4d

Observation f3fd9621-eb0b-4d3c-8131-a5d9e4e250ba · outbound

This paper cites MTEB: Massive Text Embedding Benchmark.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images MTEB: Massive Text Embedding Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.304053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.304053Z digest=sha256:a98dab7ae8b096340d9459baf502f48627457b89ee6055dc571e80b514aee558

Observation 61b06bff-b775-4ada-9f8b-47ac0bee19db · outbound

This paper cites Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:37:16.334288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.308438Z digest=sha256:510a6da4be825515abb28ac1e2b056fef82a3b73e7612558fb55ba238c531c4b

Observation e37158a8-1165-4c72-bc6f-5f5478232ecf · outbound

This paper cites doi: 10.18653/v1/2021.naacl-main.466.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images doi: 10.18653/v1/2021.naacl-main.466

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.313095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.313095Z digest=sha256:898298ccd316367ac5eb608a689a4718291144bfacd16d050cb70eca97bec6de

Observation 91aa8421-c6da-4aef-86a8-1fee89eaaceb · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Learning Transferable Visual Models From Natural Language Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.317251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.317251Z digest=sha256:90d291e16615bb2451a4688d4c5d8492aa03e1b6153a0de57f95189f57ea2f6d

Observation a4349020-5751-42d9-aa17-5b7c0ea28b12 · outbound

This paper cites WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.327747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.327747Z digest=sha256:293267ecad106949c52286e07b86d7366546a77bb5c060b8ebbdd08c0ce09503

Observation 2a20f4d8-2722-47f9-8116-f4bec7a48711 · outbound

This paper cites jina-embeddings-v3: Multilingual Embeddings With Task LoRA.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images jina-embeddings-v3: Multilingual Embeddings With Task LoRA

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.332090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.332090Z digest=sha256:5a27ed6298e36293ca9358967bd3aa53425529ed8b253fa4e8dba6ee2bb21698

Observation f5c3c0ae-35c5-4c62-b197-865b9ad31262 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.336464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.336464Z digest=sha256:7884aa9293360be3e0e88902bb29f2712a6954f6a312669e0c785bd95bc716e8

Observation db83f1c0-5257-4c02-a39e-54e4259800af · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.340949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.340949Z digest=sha256:03516c0aa758f3333c1fb06680bcc0b42615a7dc7077459263014aaa5c466124

Observation 695f1236-a10c-4582-afd9-23853ea0b66d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Representation Learning with Contrastive Predictive Coding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.345325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.345325Z digest=sha256:cd9d80f8b14a74a89b7c5e8ee89dd3fda74e421c6b07350200c7dc9afdbecbd1

Observation 250fbe52-f52d-4903-a573-66d629c886bb · outbound

This paper cites Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.353809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.353809Z digest=sha256:8871ea06e907962f1c06fc598c14343ca00e1c17977cef44262d156d594cfaad

Observation 96c2a72f-95e1-438d-8fe3-e472c690db29 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Sigmoid Loss for Language Image Pre-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.362559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.362559Z digest=sha256:a4971f8a378bdb13f0d6602fafed56ff7cf27a8f3630c9330b54aaf41d2ba91a

Observation 826e06e9-bb60-488c-ba87-b1a017e86fe2 · outbound

This paper cites URL http://dx.doi.org/10.1145/3503161.3548422.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images URL http://dx.doi.org/10.1145/3503161.3548422

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.366612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.366612Z digest=sha256:84f30f735ed722ad4bbeeb0e1c31e26d84dc00effc71521547f919589c989d4b

Observation 4f4dbeeb-e1ea-4b1b-9b90-5245d9e4a1da · outbound

This paper cites an unresolved cited work.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:37:16.323790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.370977Z digest=sha256:d9c1ce9679e3b2a68760fdd9a146a2b272d737017af13786ab0e3748c1e30e45

Observation 4ca267e2-befd-4fb2-9727-d0706de30ae6 · outbound

This paper cites an unresolved cited work.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work

Reference 36

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.312780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.375306Z digest=sha256:51fa0d70cefc446a853e6826dc4441dd73c7a8c7d2151e97809f0bec75e22160

Observation c28bcd36-472b-4b20-a5e8-2444605feb43 · outbound

This paper cites an unresolved cited work.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:37:16.302071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.380031Z digest=sha256:0c278cec264adac3731aab56fe130ef33c5d8dc9a5417057fc39d3be41d64d50

Observation f6cd3441-e01b-4b0e-bd7b-0543ac1020d3 · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.291854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.384262Z digest=sha256:cfc43fa5bb62f4171a61794819638777088997df8226145d813a71370902c467

Observation dc9b4900-f8a7-4d21-bdb1-088c472d540b · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 39

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.280362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.389257Z digest=sha256:95de2559684bf3c6054f91ea881fbbceeaf3721acfe54922f8d13e0879d78251

Observation 30a51f61-333d-41db-9760-280872923a04 · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:37:16.267200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.393505Z digest=sha256:be62afb87cadedb06f2bce57bd9f5e03041cb3245cf47ab4ec616bba29cccdf0

Observation 2b945e95-9262-4134-913c-b37b1c67f7d3 · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 41

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.252690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.398245Z digest=sha256:5ba80123a5be904beb801190136cc67276e254cef4a316627cadbb60fdcb6a8d

Observation 9ecdbf1c-4da4-40b3-97ac-6830418a7be4 · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.239046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.402986Z digest=sha256:8bd9490966aa889718674221ea65903013e44be90270b9a4d4bb6b3d4d89ee23

Observation 55737288-8ab8-4a9f-8cb3-acd6cc5ca755 · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 43

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.225670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.407522Z digest=sha256:42ea53c44f8e283f98c1eb5121fc0634c4b38ec63e40112ac359f5e304e03af6

Observation 5302035c-553f-443e-9988-ecc1524f40dd · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:37:16.211344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.411829Z digest=sha256:77270ea38e0aa86e9aeed5498e4de88324e60c0f7c59b75c6aa29191b761b6aa

Observation 0db1e746-023a-4287-a64e-85bc72203c8b · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.197339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.416026Z digest=sha256:95c4c7f8ce2b238d67ab9a563dbdefe64db7ec69ce555d2ba198db5303cd584b

Observation 5fff53ef-862d-44ad-9e7e-027fe326aef9 · outbound

This paper cites These tasks were excluded either due to bugs in the evaluation code or excessive computation times.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images These tasks were excluded either due to bugs in the evaluation code or excessive computation times

Reference 46

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:37:16.183707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.420774Z digest=sha256:c9d538911f7942d88f31fd4ef34e7dc3361e20bc4c86dad5778962bcaa29fea2

Observation b165dc71-12e0-4300-b34b-4302c578a031 · outbound

This paper cites an unresolved cited work.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:37:16.170726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.425603Z digest=sha256:be26c1374b6bd68c37cd823730a335035ea1ddb5f7282f857cd2eb698da5687a

Observation 5fa377ac-33ef-43d7-a28a-8054d68514d6 · outbound

This paper cites an unresolved cited work.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:37:16.158064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.429597Z digest=sha256:200f67ef62e1b2e8f98275692bd80f8b800519ce4b97e4df82a39fa1dfd94a5b

Observation 42d8bc64-d8f2-4028-83bb-bc4c311fc9ba · outbound

This paper cites an unresolved cited work.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:37:16.144389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.433835Z digest=sha256:36334f504d7f4e0481eee4a4685c71be5c6467864341ad99f247b5686b78b247

Observation c4312103-6f6e-4152-b7f4-b00a9030982a · outbound

This paper cites an unresolved cited work.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:37:16.128988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.438060Z digest=sha256:9e8447f5a0d3640d9c8e113228f7f3e86c7c6ce54dd561eee2e8ee6280fc311e

Observation f1d63ec3-a23e-4dc6-93d1-17c014a17e7f · outbound

This paper cites URL https://aclanthology.org/Q14-1006.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images URL https://aclanthology.org/Q14-1006

Reference 2014

Resolution
malformed identifier
no resolver link, observed 2026-08-11T17:37:15.358098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.358098Z digest=sha256:75e8154d04340638821c342d97697c458b22083e0c1553ca2724b7d23c79ff4d

Observation 1d205eae-7313-4eda-a2de-409606e5105d · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.234736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.234736Z digest=sha256:07bf7bb05baab5449004268c6a340688c24a8765da74224b908358d251729bd8

Observation a192aaa9-990f-49f2-86ff-2b99f27be51b · outbound

This paper cites Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning

Reference 2016

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T17:37:15.805982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T17:37:15.321555Z digest=sha256:be665610a396d74127407047feaa20d8c55d6b2b1f5829ef70fed351e958c416

Observation 3ea58816-7331-405f-bffe-61faed31980b · outbound

This paper cites NLLB-CLIP -- train performant multilingual image retrieval model on a budget.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images NLLB-CLIP -- train performant multilingual image retrieval model on a budget

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.349401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.349401Z digest=sha256:9fb93f50b760dc170a89e0aad3511b71a918a98973781270d05d7d25fef1dc03

Observation 5e923487-1318-4e60-9efc-1921fb78ca44 · outbound

This paper cites doi: 10.18653/v1/K19-1049.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images doi: 10.18653/v1/K19-1049

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.263107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.263107Z digest=sha256:17f0adefcbdd8b3e4809394047fa6803a78f5fcdbe710d9825bfe5891e13390a

Observation 0adfd5c6-9845-409a-b5fe-2d5d40a64969 · outbound

This paper cites Rethinking benchmarks for cross-modal image-text retrieval.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Rethinking benchmarks for cross-modal image-text retrieval

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.230394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.230394Z digest=sha256:964a3e61a098a96c5c7139520f6d7616e9598792173bda2fb28da7e47bab7692

Observation 6f658bfd-b45c-4a0a-a375-191d3ce971ab · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.267310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.267310Z digest=sha256:69bbdd07910db6c7f98e694a8b39b6e890b9d8351249a959598ae608c8442b68

Observation 7872148d-6719-45d8-a895-9308a75c46a2 · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.220990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.220990Z digest=sha256:510274ac5a198d4246cfa8d7ba9c6ef30590f660dca81dbf64b9cfede6341ec9

Observation 46aae501-ac6a-4601-a235-b9969bd1f334 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images DataComp: In search of the next generation of multimodal datasets

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.258312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.258312Z digest=sha256:b1303d828db05fbf1f3debe89ba219a53e037db7fe63bf2d7c484851e1124c0a

Observation 8172c23d-fd52-4830-8cfd-3904fed8f058 · outbound

This paper cites Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.215596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.215596Z digest=sha256:9685406e09ec0857301fef387746f14cf4d1e04432dbce279bace8f64d6e1082

Pith citing papers

Observation a37951d5-4c65-4996-8dbf-85523f262b56 · inbound

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation cites this paper.

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T23:20:01.029650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:20:01.029650Z digest=sha256:96004f2e759ed4a963cdb5655aa332fd7091c4523180579da400b2dfa8016513

Observation 4c233abe-685d-4479-aed1-9faae35b6db7 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.939130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:d4f7175518626329f9d3a7900cb671c933aa9ff70e62dfd8fe4624ac503a015f

Observation f5186e79-b65c-4286-91ea-410d98f875a5 · inbound

HIVE: Query, Hypothesize, Verify An LLM Framework for Multimodal Reasoning-Intensive Retrieval cites this paper.

HIVE: Query, Hypothesize, Verify An LLM Framework for Multimodal Reasoning-Intensive Retrieval jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:56:02.585194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:22:21.682534Z digest=sha256:d7fb70781053db12d574f385f47cc3c11feffca31c3dba45df6dcb39e079f1d1

Observation e6c5de1f-0bdd-4e19-893a-9e758e94f51b · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.838737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:ae2baf3135592f4ab13ce2bc24cb598bf8772f550e06169b7f2f188624f9df7c

Observation 8b86ce07-2594-4b42-96d6-f4e48c6c4bd3 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.779441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:20722dfe0c43221688dc9e51c420cc958f964de0ca36050ed1c99105577cb16d

Observation fb41aea0-70b2-483a-a861-eaea954a8b99 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.703034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:0d84c62b1d4db44ebf5c8ac49fac5c0b04b78f9d5be51f2a84190a68517b2d45

Observation f1cf3dec-3d74-43a2-8536-ce16658b0faf · inbound

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset cites this paper.

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:23:58.441043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T05:21:18.369534Z digest=sha256:6501921a4f1f527ba9045107d697f43dfc478d2047b155263dc1e95845bbbcec

Observation 9f548490-1d6f-4406-8dfd-aa5baf953dfb · inbound

One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling cites this paper.

One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.771000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T19:57:15.988771Z digest=sha256:28190fbf34a57e099d12cc2b6b25d671c9a37814b5459de9f10c6bf92fe05755

Observation 20b3544d-2409-4544-8762-e46faac2b04a · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:30:07.773095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T21:15:12.242955Z digest=sha256:9dac023419c86c9319ae9a1890b57e25861e6cb7a1c72ab68d518053ac50687f

Observation 67de105b-4d74-4a96-86cb-b0b461385c3a · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.253079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T05:07:28.537679Z digest=sha256:e6b11d88d523beb88ce22ef079d996ffaabe2f160493d9063e855a9cb4ae4694

Observation 994aec71-e745-4473-8b9b-2de9c6dc08df · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.719907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:41dab095f6d2d58092c1baf0b77d1ed0205004aeec489ebcae162df7bc0b94f3

Observation 5530c45c-12d9-46d5-bb05-c6171f617c16 · inbound

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval cites this paper.

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T00:17:09.450205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:17:09.450205Z digest=sha256:45318b225180e0d077f4aa1483a4426e3580e8cbaa56123f1b24145f49603e0a

Observation 73ae2735-f4c0-4927-ba64-3ea619c019e7 · inbound

Illuminating Visual Identity in Universal Multimodal Embeddings cites this paper.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.254863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.254863Z digest=sha256:1ae9946e17b9f5427d498d054b1b9018c758b742c7854ae706b404e4bb734776

Observation d832f383-4bd3-4833-acf3-633ec92e1182 · inbound

MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding cites this paper.

MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:30.104899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:03:30.104899Z digest=sha256:272304a85096d12f37ad0e3e7f52a281be49158012a7610698c2ee8eee25242c