Pith. sign in

Paper Citation Record · LEDGER

Jina CLIP: Your CLIP Model Is Also Your Text Retriever

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2405.20204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20204 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:45:05.954607Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03a6eaf3-1f33-47d8-b5fd-cbadd5602711 · inbound

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training cites this paper.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.890177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.890177Z digest=sha256:7fdd3f59d5d2d88d34e52917d0aa5926f00002610e23374445c2fe0aa63211c1

Observation 3997a9ce-1f6b-45bc-a7c4-e62e457061ee · inbound

Human Action CLIPs: Detecting AI-generated Human Motion cites this paper.

Human Action CLIPs: Detecting AI-generated Human Motion Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.650879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.650879Z digest=sha256:9779895ec7bab5d9f64d5cb45a824f967ef153687f075a9861ad30e33c525800

Observation b75c2a7e-c72e-463b-9853-8458486d0283 · inbound

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images cites this paper.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.276535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.276535Z digest=sha256:e37bdb6db8b19fa781176cb6e547b07f6310cc35d952ba345d03133190bd0620

Observation 82cd6f0a-d026-4d91-ae4d-78054994ccac · inbound

Progressive Multimodal Reasoning via Active Retrieval cites this paper.

Progressive Multimodal Reasoning via Active Retrieval Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:55:10.290249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:55:10.290249Z digest=sha256:df09b7ac4c157569b263634acd510cd8eb57319084faac81a7ea1ce598b7316c

Observation 27933465-9277-4982-a242-568cca4b7967 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:54.587472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:54.587472Z digest=sha256:4e288393f8c217f9af38484360561961dfb7b3a40dac630be36aac62a347b7ce

Observation 8f096dad-c98c-4e6c-b3e8-2e41080854c1 · inbound

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning cites this paper.

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:18:24.201734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:18:24.201734Z digest=sha256:3d823c744e0899197a6731c89477f6db93cb531da5af58a55b94afd852f0dc19

Observation a17f97fd-89cf-4b74-9030-2c605e030027 · inbound

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval cites this paper.

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:02.482363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:02.482363Z digest=sha256:262753f56b6e195e313325715bd8f523ab9a2f3029d2893159b71b3651a08a2a

Observation 3f2f21b7-9173-4403-9463-a615ed30990a · inbound

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval cites this paper.

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T18:45:05.954607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:45:05.954607Z digest=sha256:0655ca65684390995588adbb0d60c8a356509ec6864aeb9787b009fd8393aa86

Observation 33061be7-7d06-4e84-8641-a6c5f7f813a7 · inbound

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning cites this paper.

CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:16.629099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:51:16.629099Z digest=sha256:25158f4e0bf5408786d8b32bb336a88a65dd055c40859f316aea934da4d2afad

Observation 2b12457c-5b01-447d-ac71-1375b9ea22f4 · inbound

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning cites this paper.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.681915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.681915Z digest=sha256:a8a5dead216d74641ea54180aa1fc674838eb58784ba9861f95840125a390cc5

Observation 34f2df9d-6b4b-40db-baf5-a60c3f7a1609 · inbound

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL cites this paper.

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:00:58.367741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:51:16.675856Z digest=sha256:8d9229c9f327a59b6d1663205d2d2aa439f7152a673533f2cca8901e1106852f

Observation 7272083d-76f2-4baa-a4ee-b92f93047d00 · inbound

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment cites this paper.

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:58.336606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:28:59.838565Z digest=sha256:68cd229b4f81c837d15c029c9a93573de855a269372fbdea3e5219a6f79734f5

Observation d4ea2ad9-f244-488b-a5b1-93850ff9909e · inbound

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models cites this paper.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.685664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:ccf7344803eeabc76e02a458d4a1148865144536f85bdb710207268a42120f10

Observation 99934a97-6fed-4880-9dca-26f56211ef13 · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.071233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:829dabe68dce044f831f80f45852e29dd8029508f5bb37c5a0ea766a6887ca8b

Observation 0704295e-41db-4e75-98ba-6c0ec752b223 · inbound

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations cites this paper.

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-08T19:09:02.722999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T19:04:40.817413Z digest=sha256:cffa68a284e869dec718411a3954d934a389787206ac7944959b3dbca53b48d7

Observation 5c59e35f-0378-40f6-96f4-751a7f31a020 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.867871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:01b583ca782d04aa3a9e77828ac53098b6fac7549ee3b21fab60d1a6aade1723

Observation 041fe031-0521-4a9b-9763-085fada59064 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.816288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:2629cba0fe47c8e3942371e623df935e5ffd989be15415b67733c9ce05196295

Observation 55ed93fa-6ac9-4c67-b0a2-a885ab380a75 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.700545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:0319affb09621e9567d88b4b1e9993bcfefa56862006d65444fab0a5309eb53a

Observation 3b2932d2-1ebd-4140-9973-cb608e17bd23 · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.462858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:0650da62f1fc1aaf9bb5085eb1ee8efff9ace835f53c313d9fc204c6babb1496

Observation af075db8-671f-4551-9335-c014c5432b89 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.782048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:351f2815130204d4ddc0b54ff6c8df4e98cd22a8999cfb46b6b6914c380ec083