Pith. sign in

Paper Citation Record · LEDGER

Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2404.07983.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.07983 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:52.101762Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:56.194117Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0db09d49-9aa0-466f-8fe7-43051d776125 · inbound

Aligning Multimodal Representations through an Information Bottleneck cites this paper.

Aligning Multimodal Representations through an Information Bottleneck Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:52.101762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:43:52.101762Z digest=sha256:d7ec3e8c5c2edee5464cb16be0fe7078157877f6fc1d803f1008646ae2b42675

Observation 75e0a77a-2fa4-4472-ae3c-4ba1f8c1d869 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.310154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.310154Z digest=sha256:b85cd74af5df39aa01b62302335513e80537db1eb57cf7872e2c8bae1e19bb2a

Observation 3e104e80-3169-4f91-9d14-f6f87cd195d6 · inbound

Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation cites this paper.

Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:28:42.092582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:28:42.092582Z digest=sha256:24c62d8ba6300304cc56143cce00deb6af020cd823be19523690645940f62e6d

Observation 664a590a-6d3d-4696-bbed-00748126c35f · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:29.198405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:29.198405Z digest=sha256:c3c0bd6f87ced76061a45ead3f855c9eae26b6f1978cabbf9fc52bcb5c308d88

Observation 8d894cb5-c695-425d-b294-db6419b67047 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.255620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:5b54ffae9705219e97d4b6e7be2e28e5f5f50b1d40a71a0925ea89c529d19461

Observation 8dac79d7-f1df-4eab-ab21-bca0d0161410 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:30.264059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T01:24:56.236650Z digest=sha256:5b6dfe76fe482fae777fbd9f9370d8af31a3e553b8bb5b5e5f26c028a0668c8d

Observation fbc8683d-1392-4ed2-b25a-0fbc605e5064 · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:35.800091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:14:12.092478Z digest=sha256:303868e7a41c038d9993489b181baec7f1e340b70cdfc4d434d70f522419fae6

Observation bb8ba82d-434e-4d52-96f4-801f7d839b1e · inbound

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection cites this paper.

VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:48.477044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:33:55.317107Z digest=sha256:c87961f1772593ac8147eb20a114bbb785c812f3935edb9a071ea20cb56d79d6

Observation 204c49ec-2a83-4a32-bec3-00acdb7e2416 · inbound

Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning cites this paper.

Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:02.543555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:24:22.361181Z digest=sha256:fc4be9d9e7c7122b8253dd79bcc1b6697c5b599badd4a5f2baa93fc79db0196e

Observation 0604384a-580b-4eea-be5d-350088ca2ec5 · inbound

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion cites this paper.

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.083496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:20:05.081980Z digest=sha256:5dc1ceb09856697b15801cb97fbf6556a55e9232a638eb53b1c5e8b50c7a8fb2

Observation ebe164c0-0317-4617-a594-1be920bbd7b6 · inbound

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models cites this paper.

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:46.081302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T06:56:00.782128Z digest=sha256:e8a219b856838f80b0ffb4c3be5427b11909a3287168f1a464249dd2f6c573fd

Observation 05a2ecf6-785e-4c8d-93df-da99bba02c73 · inbound

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model cites this paper.

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.196361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T01:31:57.305121Z digest=sha256:af30c50b1d378347cade33578752d09b7aef6ee1c7a515dcada84bdcf49d73ff

Observation 638d44d3-5364-4180-b6bf-c4a6319ab41d · inbound

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization cites this paper.

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:43:47.409552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:43:47.409552Z digest=sha256:07566022bcced1510e1d8defca548ce69bd4025ef9fec1307f64ae63323ef35d

Observation 00f0db15-6aa4-411f-a033-4e59e2061d5e · inbound

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs cites this paper.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.037233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.037233Z digest=sha256:156d6366dc0983e3e9dbe9a4837150bfd0bea5a153ba9a63b3f799acf2d962ac