Pith. sign in

Paper Citation Record · LEDGER

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

As of 20 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 4 inbound Pith citation observations for arXiv:2505.03703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03703 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:49:24.756575Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:39:05.154618Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T09:42:30.741220Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b05f7d0c-d197-4e7b-9b58-546959f31903 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.699051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.699051Z digest=sha256:f538141a5a88157b1f28cc076597855394ab9dc77ce7f7c74738fe851df2e814

Observation 25b28392-3b5f-4f04-a5e8-6a963e629f53 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Sigmoid Loss for Language Image Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.704727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.704727Z digest=sha256:9c2d599bbd557c6823c889612f5561b7062d644e01f6c8d7cb86266eaffda34f

Observation fa06d3a8-55a0-4e05-8d22-527827118964 · outbound

This paper cites Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint 2411.04997, 2024.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint 2411.04997, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.709451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.709451Z digest=sha256:04de013bfc74443f3ac1c71e783b77395e24d7b50d6973c46ee312b1d404b82e

Observation 081861d1-f654-4530-b405-766702e8acd9 · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning ColPali: Efficient Document Retrieval with Vision Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.714057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.714057Z digest=sha256:87ea6a8cdbc3a02f157e3f3915ad3e797ed149f8d39b8fd9c696aee972966c29

Observation 8b75a658-b4a7-4591-b72c-cc076fbd4dc4 · outbound

This paper cites Mind the gap: understanding the modality gap in multi-modal contrastive representation learning.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Mind the gap: understanding the modality gap in multi-modal contrastive representation learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:25.008244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:49:24.719757Z digest=sha256:e1d62309b224fd1560d48bf7acb22e7514a98cdab1720d4c0cb10835c3eebeb2

Observation c221131b-7b09-4d05-af05-a5623206d7c7 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retriev- ers.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Uniir: Training and benchmarking universal multimodal information retriev- ers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.990060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:49:24.724652Z digest=sha256:256fbcc9d6f9946aa1f00b193d6ed64669ddedd29db756aa56bb5ed7c11aa5ef

Observation fe76457b-ceb6-483c-be0c-78bbde9e4b48 · outbound

This paper cites an unresolved cited work.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:49:24.974389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:49:24.729637Z digest=sha256:5fe97bae21bda2544b3eb1b6696795c5373feb96b87dbdda43c393b15946f208

Observation ae0b1950-e9fc-481f-b128-63189e614000 · outbound

This paper cites A Tutorial on Spectral Clustering.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning A Tutorial on Spectral Clustering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.734226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.734226Z digest=sha256:321a06747f6683df5db66d689fcb150f26aded5b2bc99a08ba72c8bc49404987

Observation f735136c-af65-493d-908f-9ea012613aab · outbound

This paper cites Regularized discrete optimal transport.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Regularized discrete optimal transport

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.958064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:49:24.738839Z digest=sha256:2b811a1ab04f85aa681cb376b2568aed4a549559d13c5d47c07b11222672e374

Observation 7f3fb7f5-8139-4cb5-ae49-b7827fe6d1c7 · outbound

This paper cites Optimal transport with laplacian regularization.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Optimal transport with laplacian regularization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.942605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:49:24.743237Z digest=sha256:28310afc79193a36bdaaaa3fa1b9eaaf47b72d4a74da8937ed0d12424840af3d

Observation 1e1980cd-8e85-465b-8798-6ecdb1719910 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.926768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:49:24.747664Z digest=sha256:4cb8e587ac767bee3903b03992c94cf0c247375e681ada64353d8dab03cc5625

Observation 7fccb29f-b114-440f-bce8-0473a2caddef · outbound

This paper cites Lawrence Zitnick.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Lawrence Zitnick

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:49:24.752120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:49:24.752120Z digest=sha256:2d099113b7d291168fb2def8383e43a42b776db8462c21192eaf7f5dd00f491e

Observation 2b385166-435a-4361-9a3d-781166696547 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:49:24.901369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:49:24.756575Z digest=sha256:7b637ddf375ff830b0c4fbbd6f507b2d1ed628a0b71ff237913e2c75ad2b310f

Pith citing papers

Observation 54cde237-0bc4-4ebc-acaf-e3f5582b2e54 · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:05.154618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:05.154618Z digest=sha256:720e1c448a5483eaf8e750a34351b77433788ecbc760a52114db8cc33e0bb150

Observation e6954add-8b1d-4272-9f05-16a6c374c73a · inbound

Asynchronous Federated Learning with non-convex client objective functions and heterogeneous dataset cites this paper.

Asynchronous Federated Learning with non-convex client objective functions and heterogeneous dataset Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T05:28:38.325633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:28:38.325633Z digest=sha256:1e59a50219cb549b9ea8c746caf0a9249b8cae75f997532c3ae6ac77a50ce2b3

Observation 372ad32d-ad96-4cd9-9f47-e6bdeb8bf8f6 · inbound

Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization cites this paper.

Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:42:30.745067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T09:41:38.974413Z digest=sha256:f5320c7c9b5a98c5858afff891fcded4ce7579871a796c0806d448f8182f1a7a

Observation aaa20988-4bfe-48ae-b506-7093ef8b35f1 · inbound

On the modality gap and the contrastive loss in multi-modal representation learning cites this paper.

On the modality gap and the contrastive loss in multi-modal representation learning Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:53b09ba46b059e74228e087484056eafc099f387cb125f4d1b8d81df4738f83f