Pith. sign in

Paper Citation Record · LEDGER

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2203.02053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.02053 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:19.855755Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

98
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ad9eeb77-3851-4f76-8702-3702171fdb4a · inbound

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings cites this paper.

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:14:16.094675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T09:11:37.220487Z digest=sha256:13f97155a33eb72cce8acaf96ee936c83743ccb6669cba0bc3e0b0bef45a4fcd

Observation 9a8f0407-498c-4677-afa6-5800b0e5c46b · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.855755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.855755Z digest=sha256:2c09d8dcb2a90617475ab232a9e3fbd604f6aaba7182b507d925866f297ba6a3

Observation 4ea027f2-e3fe-4a8d-8cd9-a931b167e2b1 · inbound

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps cites this paper.

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:37:23.778509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:37:23.778509Z digest=sha256:47423718836eaf7e6f6a87a5f3e2515cd7ff4586beea9e0f0d22384d2989a724

Observation fcc47bb4-a808-48f3-84ad-483bb1702b32 · inbound

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models cites this paper.

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:50.461183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:50.461183Z digest=sha256:63cb513767bab44faa3ef2ec169766a4b5c5d878a1e76bf2c8ef7920e12bbb2a

Observation eb9a4f0d-d496-4c51-98a7-aed0c5d3577c · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:21.947358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:07:21.203513Z digest=sha256:c9f44fc16997ca4d17d8b08ab3b00ef627352109847e8e827243cc74dd8d2786

Observation 31fade60-4c4a-48ce-9548-16a5c26c91dd · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T20:05:20.702092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:05:20.702092Z digest=sha256:010d8f427fdc6b4935b43eb3535f2e5d6922b6fe807e1b6dec1573f846615f71

Observation 65e85d52-f2ec-425b-92e6-29a8222780e3 · inbound

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs cites this paper.

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:06.934858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T13:17:48.295630Z digest=sha256:d60326bc6da4a381984ce571bd895c0d9547c881db887b5f571fcd4ff095cd8f

Observation 90dd2a2a-699c-4669-a914-e563b100455d · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.955356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:51847f6aeb11fcf742089ca7de19093697f522a859dbf42b446b551ddd013148

Observation 672f9bdc-1ebe-4665-a0c0-58a83e959e86 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.740739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:934c6d69aada444396896183a903dd7eab67fb91dc51c96c2c45725baed2ad7b

Observation 4ed090c5-521d-49d9-8a7c-e078e60cf360 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.736252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:bdc373e5bc3d7bb8a3ff13daf76d337deba7368abe2a451c4068d3376db75561

Observation 65b23aaa-f7d2-4ab2-a270-9cc6ba0fc6c0 · inbound

DeconDTN-Toolkit: A Library for Evaluation and Enhancement of Robustness to Provenance Shift cites this paper.

DeconDTN-Toolkit: A Library for Evaluation and Enhancement of Robustness to Provenance Shift Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:05.858809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T02:12:54.796350Z digest=sha256:21a92f18246f1d0db57c85f391440c76b98138a9ab75392b1eca8227b61b89cb

Observation 77e2541c-9271-478f-9107-975d68981210 · inbound

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning cites this paper.

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.665073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:32:00.361991Z digest=sha256:8f1948e5ae5c0694ebe7dfcb5907a12f311e3ec4eb5e83c753c50f817743032e

Observation c5ef43e4-90d6-4834-a917-68e5c3113c5b · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.050488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:7b215a5abe846f9fdfd2f7e5d90b7fb76e60b887c5f026da4a5816cb21dddf11

Observation 40b2cdd8-71d1-4a78-b8e7-790c42dd38ce · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T14:10:57.796470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:c18a3ff6d3e5f6e9e563e85b5df733c55ebf3f03077861824eba97f290423efd

Observation d65a7949-bf34-45f5-aecf-fbe5eee2f8a7 · inbound

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms cites this paper.

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:34:13.357363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T03:27:31.177489Z digest=sha256:464114f49cb9678524373ce560b3bad95284b78179f4bc9c51f5c5c6d057a2ef

Observation 58328478-9e2b-4265-9cf4-83b2d7999dc3 · inbound

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model cites this paper.

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T04:33:53.287239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:33:53.287239Z digest=sha256:65ca60131aa840d6a23625d924275221fcc8045e5a1ac16a0d2f68da5fe72cab

Observation fdc3b59e-ffdc-4230-b0ab-23375b4f97d9 · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:48.890280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:48.890280Z digest=sha256:2364927991e0e54c1202e0a9f37282820ef26c084ac285c0d59f34297fb90b3c