Pith. sign in

Paper Citation Record · LEDGER

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2203.02053.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.02053 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:19.855755Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

98
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ad9eeb77-3851-4f76-8702-3702171fdb4a · inbound

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings cites this paper.

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:14:16.094675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-24T09:11:37.220487Z digest=sha256:4e42200100c527643188f2e814d0e069a977a12cebac453896c0bdce3396a4d9

Observation 9a8f0407-498c-4677-afa6-5800b0e5c46b · inbound

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning cites this paper.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.855755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.855755Z digest=sha256:2c09d8dcb2a90617475ab232a9e3fbd604f6aaba7182b507d925866f297ba6a3

Observation 4ea027f2-e3fe-4a8d-8cd9-a931b167e2b1 · inbound

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps cites this paper.

QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:37:23.778509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:37:23.778509Z digest=sha256:f9d36982ac4b4250d46aa548937a3e6b65d72ee0dea2ea4032415492b9b8579c

Observation fcc47bb4-a808-48f3-84ad-483bb1702b32 · inbound

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models cites this paper.

Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:50.461183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:50.461183Z digest=sha256:63cb513767bab44faa3ef2ec169766a4b5c5d878a1e76bf2c8ef7920e12bbb2a

Observation eb9a4f0d-d496-4c51-98a7-aed0c5d3577c · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:21.947358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T12:07:21.203513Z digest=sha256:81843909d32585723f364e9df71e6a60c87829129bb189c5061ac6012e405993

Observation 31fade60-4c4a-48ce-9548-16a5c26c91dd · inbound

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts cites this paper.

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T20:05:20.702092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:05:20.702092Z digest=sha256:010d8f427fdc6b4935b43eb3535f2e5d6922b6fe807e1b6dec1573f846615f71

Observation 65e85d52-f2ec-425b-92e6-29a8222780e3 · inbound

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs cites this paper.

GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:06.934858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T13:17:48.295630Z digest=sha256:41b8d6739b498c67aef9b209392744837e6a36fc44588943d0d4ad4d3110bf83

Observation 90dd2a2a-699c-4669-a914-e563b100455d · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:28.955356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:64cf039beb0a0822c733c009805419d34e6778edba863dcce52875b15182f972

Observation 672f9bdc-1ebe-4665-a0c0-58a83e959e86 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.740739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:1c7c4fc271cc5c60cc7fde8d4e1a7acda4b5d9e35bccac0c68a87a0dc37ac2fe

Observation 4ed090c5-521d-49d9-8a7c-e078e60cf360 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.736252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:616688355cdc0b83ad8ff668d4e6ec4d712561de1a1f8ec654873d6f4a136464

Observation 65b23aaa-f7d2-4ab2-a270-9cc6ba0fc6c0 · inbound

DeconDTN-Toolkit: A Library for Evaluation and Enhancement of Robustness to Provenance Shift cites this paper.

DeconDTN-Toolkit: A Library for Evaluation and Enhancement of Robustness to Provenance Shift Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:05.858809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T02:12:54.796350Z digest=sha256:8ac1739b4ef18eff5b34f9ffe5ca880de1b5cee6ae4c8882447da555a83f18de

Observation 77e2541c-9271-478f-9107-975d68981210 · inbound

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning cites this paper.

SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.665073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:32:00.361991Z digest=sha256:c0611da8f3445ee1b5d589cc73c962f543921ded8a73858d31dcfe453dd90c22

Observation c5ef43e4-90d6-4834-a917-68e5c3113c5b · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.050488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:09c18833d489605c1a3c98e2bc9212cf001b4e111c0a2474500a34bc5597924f

Observation 40b2cdd8-71d1-4a78-b8e7-790c42dd38ce · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T14:10:57.796470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:beae9f629aa4235813ebf96e98aa7bf7220d62c462ceaafa6576927ec3e2883f

Observation d65a7949-bf34-45f5-aecf-fbe5eee2f8a7 · inbound

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms cites this paper.

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:34:13.357363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T03:27:31.177489Z digest=sha256:fa24ccb7805a0c711f7cd41f6404d7b94dc398ba4cede2270ca7a8f12dcba8ac

Observation 58328478-9e2b-4265-9cf4-83b2d7999dc3 · inbound

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model cites this paper.

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T04:33:53.287239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:33:53.287239Z digest=sha256:65ca60131aa840d6a23625d924275221fcc8045e5a1ac16a0d2f68da5fe72cab

Observation fdc3b59e-ffdc-4230-b0ab-23375b4f97d9 · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:48.890280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:48.890280Z digest=sha256:0ff1d8d31b2bbf9e0f1da9a742ea13d49613f2433858c96a2135c788940d5a71