Pith. sign in

Paper Citation Record · LEDGER

Jina CLIP: Your CLIP Model Is Also Your Text Retriever

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2405.20204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20204 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:54.587472Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 27933465-9277-4982-a242-568cca4b7967 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:54.587472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:54.587472Z digest=sha256:7488698f1fda525b0b91ed313fcf7cd27572aa24dceddcd97384ed6f6ed13970

Observation 8f096dad-c98c-4e6c-b3e8-2e41080854c1 · inbound

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning cites this paper.

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:18:24.201734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:18:24.201734Z digest=sha256:1594f4470c1961578ee32709a23b78fd33388c4a182c7f587388ac49f275565d

Observation a17f97fd-89cf-4b74-9030-2c605e030027 · inbound

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval cites this paper.

MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:02.482363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:02.482363Z digest=sha256:af94b32d89ec0a10aafabbbda2ea59a29e57554ff2d4ba4336c3c1540823fa51

Observation 2b12457c-5b01-447d-ac71-1375b9ea22f4 · inbound

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning cites this paper.

AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:15.681915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:09:15.681915Z digest=sha256:cf6f8bfd4285179fefec63f7e1d9c26459ef8d6673554b323af202be81ce8c95

Observation 34f2df9d-6b4b-40db-baf5-a60c3f7a1609 · inbound

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL cites this paper.

MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:00:58.367741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:51:16.675856Z digest=sha256:68a24f18c4c2ff40f0342bcbd2863c44cf2873c25b5510c711a0a63d6fc9d430

Observation 7272083d-76f2-4baa-a4ee-b92f93047d00 · inbound

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment cites this paper.

BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:58.336606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:28:59.838565Z digest=sha256:14ce084a8746cb9eed6ef3313016afda683172324331de7948f94debaa0d504d

Observation d4ea2ad9-f244-488b-a5b1-93850ff9909e · inbound

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models cites this paper.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.685664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:182bc9af48b92afc2d98c1707929de36bded0399fef66a863cff7f4a1b5d0e94

Observation 99934a97-6fed-4880-9dca-26f56211ef13 · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.071233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:bf2c1949ffa6c54b219ee6f73334101d219d114e9eb6ae45e5c7c0ffa0e0097f

Observation 0704295e-41db-4e75-98ba-6c0ec752b223 · inbound

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations cites this paper.

TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-08T19:09:02.722999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T19:04:40.817413Z digest=sha256:d7312d23acf1743a1754524cf0ecbc84f071da4a4323847cd73ff76b93f333db

Observation 5c59e35f-0378-40f6-96f4-751a7f31a020 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.867871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:5da8c29c745ef763079cf8a8ddb43f691dba6d74e6aeaca9072ec0d139fcf8f4

Observation 041fe031-0521-4a9b-9763-085fada59064 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.816288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:1f4ce9b0534bfdadf1df4c7be73a62d3a3f3f4fe73b13dced2b73a324e04e3e2

Observation 55ed93fa-6ac9-4c67-b0a2-a885ab380a75 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.700545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:c2c2acb75a3f7fcb9450c0e2d2602f5839daf29f80ce261d8084328cc8bbe7aa

Observation 3b2932d2-1ebd-4140-9973-cb608e17bd23 · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:28:31.462858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:9dda8f41d90537b7930cab7d974fda68ef17fc94430aa89cc7aeae873a629113

Observation af075db8-671f-4551-9335-c014c5432b89 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.782048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:1ea7c2030a6aad41a8c87e5c0251e018eecfbc0b3cb830ccf1208d5ca2d49233