Pith. sign in

Paper Citation Record · LEDGER

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

As of 22 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2607.28640.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28640 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:55:27.313147Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved55
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bf742f8-da7f-42ec-84c5-bb0a48f8177d · outbound

This paper cites Vision-language models struggle to align entities across modalities.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Vision-language models struggle to align entities across modalities

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.224783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:22.224783Z digest=sha256:fc291035a7ddd46adf02a3d577ca4ed4a6bf8b25af788e6d5f56548b4ee7f065

Observation 3a0966a1-e105-4bf6-838f-4d5fa7fe0b14 · outbound

This paper cites Introducing Claude Haiku 4.5.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Introducing Claude Haiku 4.5

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.388224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:22.388224Z digest=sha256:fa866fbfcd03591d431499f77a6bf989db891b71a7d7477478d63c759204f9c7

Observation 00fe7b66-8fb3-4a60-b459-90d41d47cdc4 · outbound

This paper cites Introducing Claude Opus 4.6.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Introducing Claude Opus 4.6

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.568678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:22.568678Z digest=sha256:955abc95d954b2bc4c11e5ad1f86bb6bbff9338ba78e7860f4b1d5be2edb4499

Observation dc994809-00b1-4113-a0d8-727255b33736 · outbound

This paper cites Introducing Claude Sonnet 4.6.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Introducing Claude Sonnet 4.6

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.702350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:22.702350Z digest=sha256:3b8e19256146af24e7ac236b252f2251397c89f5d557ed1bca588c12b6452fc0

Observation 42d50484-fd84-4d72-bb1f-45c4d67beb0d · outbound

This paper cites Qwen3-VL Technical Report.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Qwen3-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.846358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:22.846358Z digest=sha256:12578ade66010d4fd68ae5048a76960e6f90f9f97f01f6ec9de57c4b817d0d59

Observation 1a322ee7-bd9c-43d0-8fb1-c815e3775b67 · outbound

This paper cites Qwen2.5-VL Technical Report.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.982366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:22.982366Z digest=sha256:1ecfc64dadd76bd040d8f053bf753f6ab15144b09657726cad5dfa7d9d82aec0

Observation 6704d8d7-3840-44fe-a0a4-fcc9e4296d58 · outbound

This paper cites OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:23.079089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:23.079089Z digest=sha256:049a0f6ecf01e4c2e62746f2ee46f918f3d48cee375aca91b94aff3f8808cd6b

Observation b9728877-6b05-4fac-836d-8c8630b34d32 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:23.218254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:23.218254Z digest=sha256:197ccaf5e604a7c428326459938adba73e98dca3f278b7c1f14d7aca111195b6

Observation 3b5de65c-cc3c-4a3d-a736-808e8970bbc6 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Reproducible scaling laws for contrastive language-image learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:23.365501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:23.365501Z digest=sha256:a4cf15d2f30ee14656274f02a9814b71546e437a1a3522e3bbc7481a7da8ceea

Observation 8ab08e30-816b-402c-9b67-834b125bf141 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:23.535097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:23.535097Z digest=sha256:26191d699969ecc9ff3745731fb99c09a4b437ad5da7664a5f63f8bd4acb3b2a

Observation 20e46a5d-9c0b-4505-9b39-08dc69a291c2 · outbound

This paper cites Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:23.673276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:23.673276Z digest=sha256:e10ec8712f160c1950f8dc0da4cbd8073e694eb3c621bac4cb57165c386301a5

Observation ee249f0b-2bac-43a7-aee1-5d80eed4bd09 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:23.784825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:23.784825Z digest=sha256:2d3c2610dbd77b1c4d38da0b2b2b50f24e81e08209032656d1c03ca205ed3bb9

Observation dc9a9875-1f58-44a3-8bfb-79d9dd641b67 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:23.885090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:23.885090Z digest=sha256:db424ec924475eaedf25be5b388432c9dee47dc8ca46ff11864ce6955d064f5d

Observation f539f254-2506-4a97-b330-42f890d0f58a · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092–27112, 2023.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Datacomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Systems, 36:27092–27112, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.035488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.035488Z digest=sha256:a1595a03710ed943acad6c190e52c8413e7d4fe67285efac987f48342854c137

Observation 566e1361-173d-43c9-8ab6-142c4e7aaa54 · outbound

This paper cites Gemini API Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Gemini API Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.111180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.111180Z digest=sha256:366ef7c188da4d3fea1b9451f96a363dacf5cb58b06d9e82a775ed55526b6415

Observation 4f45055a-0d49-435c-a375-3294185d348b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.238463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.238463Z digest=sha256:0292a8f4b33cde3cc8d2402b888ec385ef8e11e87825e44c472cb934925bcc4e

Observation da25ee71-5af0-4d47-affb-c6f345106c52 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.375267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.375267Z digest=sha256:5d23da4f0340243a190058a376b72a0c4a9ad2d8e2801f50cd8cc4b0d18db8c1

Observation adbf1c56-c270-4099-918e-c7efecfbb718 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.586305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.586305Z digest=sha256:88ac67f1c1294bfeba7634bababbaae9f4f2e4d79559979d3402f2a418ac7c3e

Observation 1de3104f-f5af-406c-816f-21b71a80e72d · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Scaling up visual and vision-language representation learning with noisy text supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.718228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.718228Z digest=sha256:0d4d7e669aea516421c2953f6da48e2f163149933eff936d81755a6bcf1d4076

Observation 9f529eec-e8cb-4847-a198-10bc23fba8ef · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.874739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.874739Z digest=sha256:d27223a12b1fb1889f84a33e91d7c8109817596cd535b72c4cb16cd0f65ce897

Observation 8975576b-32f3-4353-9b1e-dc1c66504366 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.983000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:24.983000Z digest=sha256:a3ba651acb67b50b1367466d46634060c68485ec32ed3ccdbc7a28adbdfc10a2

Observation 2142a2af-4d36-44b3-9c7d-7975c404ab32 · outbound

This paper cites Improved baselines with visual instruction tuning.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Improved baselines with visual instruction tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.149705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.149705Z digest=sha256:c9257177feca9c4c62cb8049cf2405820e41b145154edf8ef6942fb145a9858d

Observation 012b3093-ea5f-4bdc-85bd-f04864a493b1 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.276883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.276883Z digest=sha256:9f1fddeb334034a571e05911f28cdd1667f7f872bbb09df2814f102e53fee03d

Observation 86dc6205-2485-4ff5-8862-d7a54f39269d · outbound

This paper cites Scene text recognition using higher order language priors.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Scene text recognition using higher order language priors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.361355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.361355Z digest=sha256:db22d5edc3b8753c02481a77be85ce978fc7adf9dfaa3c8e4558eacfe7ea604c

Observation d7ea08f0-a96b-48ba-a685-82f1f60e7849 · outbound

This paper cites GPT-4o System Card.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.450674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.450674Z digest=sha256:3d0f6219575b90fe923ae28a0bc313ab26693136179d649291b8400853249b08

Observation 4fc60180-ba70-4308-8b01-b080645a3c73 · outbound

This paper cites Introducing GPT-4.1 in the API.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Introducing GPT-4.1 in the API

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.639465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.639465Z digest=sha256:8843355f4f0f6b14e8c8ba8aa3a2df4b5e6f04befc404d880827dd42266ce06c

Observation 4f89df5b-97bc-4ea7-a589-c765dbdefef6 · outbound

This paper cites Introducing GPT-5.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Introducing GPT-5

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.735867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.735867Z digest=sha256:0fc23d270f3c480f85ff92712371bfed26b34a24dab05866ef967db895461af4

Observation 436fbab9-e63f-4f85-8eb8-147172e5f104 · outbound

This paper cites Introducing GPT-5.1 for developers.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Introducing GPT-5.1 for developers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.804458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.804458Z digest=sha256:fcaa47be7ed55d69cfe599780db25b8d3f341702271a36e1b8f74eacb2c3707f

Observation 3ec1c4cd-3148-4c4e-a76a-c064e3633caa · outbound

This paper cites Introducing GPT-5.2.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Introducing GPT-5.2

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.899217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.899217Z digest=sha256:98b4e04525702efdfb9e6f5d3e18bff4c943ccedc52d62a337ed7d16151c7387

Observation 0cfd5b92-4b19-4b07-9f29-0776e74f1302 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Learning transferable visual models from natural language supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.999038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.999038Z digest=sha256:e63076f407ed66db8be38fb11708737d58e65ac5b821be6e41d3f3a2d038679f

Observation 00f0db15-6aa4-411f-a033-4e59e2061d5e · outbound

This paper cites Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.037233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.037233Z digest=sha256:2713a4ae7c8a4583615d92cfaf98af388fb993a9580c961b0c87cafa885d42ac

Observation 274f2c6e-a451-4eb8-a31d-569a5b5476c5 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.080641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.080641Z digest=sha256:9516ecc30e790b8e42770e1794f21a74fc78be3bf031384259887b741aa94078

Observation 4c2ba4d4-a67a-41a8-abd2-5f8ec0eebcfc · outbound

This paper cites Understanding the modality gap in clip.ICLR, Stockholm, Sweden, pages 8–10, 2023.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Understanding the modality gap in clip.ICLR, Stockholm, Sweden, pages 8–10, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.122685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.122685Z digest=sha256:421db59aca9894771b581fcde1c524cc8c36ffabd90846aa52ef09db2d260a39

Observation e2486b27-fd6e-437c-aea4-920b8f8e2bec · outbound

This paper cites Towards vqa models that can read.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Towards vqa models that can read

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.166125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.166125Z digest=sha256:b8289c5b0d5cb773e139efc23aa7391f20a4dada18e87610d8844460077c6ea9

Observation efc69535-bf0e-4bd0-8bd8-f5406d23872f · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.204914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.204914Z digest=sha256:58740fdfe50d2f60d5a530c8b4321b11b506336c1db222305fc2f5562e4d4966

Observation 351a9023-0a61-4de3-8da0-72858a7b8fe0 · outbound

This paper cites SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.251613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.251613Z digest=sha256:3cd6731db324b92a9882f1698d064184e2ec34acc8bae64e1ab5f02b60077a9f

Observation 217acad4-03c2-4acb-bdf2-ffef09a28bf1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.298996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.298996Z digest=sha256:f78c401e12ee5ea8a2e665336c59867828f35d1e76562dd559ac4fca46a31710

Observation 795124b8-4d41-470b-a4a5-3e20985a1bf5 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.353484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.353484Z digest=sha256:6a4564f03ab5c9928b12503b535eb0ab29e6b1807fec959c445cb90410189802

Observation c707c0d3-f389-4d8f-8197-e927dbff2e72 · outbound

This paper cites Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.402402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.402402Z digest=sha256:4f5c3b326efa4a6bb7d92d967f373b3163d97dc9944bf0a0ff4aeb1b71616457

Observation 1eabe3f4-4d68-4aa1-a539-42d62e00c877 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.464199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.464199Z digest=sha256:4bab066201fbda7aa1a121a1e75ae21f16777a750d48b79dd7240b21e0c76b92

Observation ebbec044-eab8-4e3c-8dd3-ceeac2ed0673 · outbound

This paper cites XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.517339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.517339Z digest=sha256:826205fb388c163b35255773a4f80deaae0d2001fdff88b0c26de750237fc301

Observation 8ca04334-b4c5-4f82-901d-24312791a546 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.568021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.568021Z digest=sha256:907a94820592a974b341e0bf33de86de2cd26eff0c2bd7953952b18353e14bff

Observation 9b1609f3-bfac-4722-875a-be9f74c531c6 · outbound

This paper cites Explaining and mitigating the modality gap in contrastive multimodal learning.arXiv preprint arXiv:2412.07909, 2024.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Explaining and mitigating the modality gap in contrastive multimodal learning.arXiv preprint arXiv:2412.07909, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.635411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.635411Z digest=sha256:a04c1a3c8d45723bd56d0a60985a14ade63a558d329dd7bcc5774b89538fc4cd

Observation 1e4f6912-b653-4bc9-8eac-cd853d9e8242 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.669492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.669492Z digest=sha256:75040db4a8c798f7497b4a84d2f9674533d72d1d4e25c4e64093c674648574ed

Observation 62d4f79f-429e-4acf-9239-635abca262b9 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.741382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.741382Z digest=sha256:333d3a31ac61fff19f24c848ae029871188cdffb0f460918f4ba29f31e5d56bb

Observation 45c5474f-a472-4e98-a4de-1637082407b5 · outbound

This paper cites Sigmoid loss for language image pre-training.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Sigmoid loss for language image pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.794791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.794791Z digest=sha256:8f6059274250c02eccad9a86f184354d34036d7312808c464ead38ddc449a71b

Observation 74446a3f-8d0e-4e36-ba2a-d031d5df0af7 · outbound

This paper cites Lost in Translation: When GPT-4V(ision) Can't See Eye to Eye with Text. A Vision-Language-Consistency Analysis of VLLMs and Beyond.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Lost in Translation: When GPT-4V(ision) Can't See Eye to Eye with Text. A Vision-Language-Consistency Analysis of VLLMs and Beyond

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.849480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.849480Z digest=sha256:0e625ef998b7eb39f26f817123d4a0ba995973ccb8becbfde182199ac9aaf07f

Observation 0a8925eb-0edb-4761-9cb5-c604d27d11c5 · outbound

This paper cites Cross-Modal Consistency in Multimodal Large Language Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Cross-Modal Consistency in Multimodal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.897977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.897977Z digest=sha256:b3607caca1dc67564f5a0450a700036653495bcd7251a73d58f5da0bc62c6d1d

Observation a8150157-dfc4-48d9-b641-ca7572dff2f1 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.936954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.936954Z digest=sha256:dbf661781bc6ba9e65b24294a106470690bd399755ea58be2b46568a6bf80d49

Observation ce086900-7c1e-4b2e-b0ed-6ecb8c0fa0cc · outbound

This paper cites - **RETURN FALSE IF**: The word is a verb but the image shows a noun (e.g., ’die’ as in death vs.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs - **RETURN FALSE IF**: The word is a verb but the image shows a noun (e.g., ’die’ as in death vs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.982981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:26.982981Z digest=sha256:a20dd33c7fac5f483e1ffab51deb195c8b8b0a74dc0b228a2acbb9f9dbe7959e

Observation 36b1e162-d94f-4555-9544-6780d6379617 · outbound

This paper cites human head.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs human head

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:27.028042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:27.028042Z digest=sha256:9c15521f4c993138e68bbca6a27f7aa462e84783246118d7a76126dc576482bb

Observation defab8ea-7503-4279-af6a-33b0d096f24f · outbound

This paper cites hot desert.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs hot desert

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:27.061895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:27.061895Z digest=sha256:3c2d0c964b4201794d9fa8790d1f39d27e9c6a62ce750d184a2aaacacb975df1

Observation f74b7389-febf-4d15-9008-3199f1dcac5a · outbound

This paper cites an unresolved cited work.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:27.106744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:27.106744Z digest=sha256:134f1bef3aebd326486d17d9ea1b9e73bc0bcdf76120a7643fdc2e7d8412f64e

Observation 1462c91f-f99c-4c5d-aaf1-14bea1506142 · outbound

This paper cites The hot <image> (sand) scorched his feet.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs The hot <image> (sand) scorched his feet

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:27.183439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:27.183439Z digest=sha256:aaffbc7834d68a97a757f34930df5e0be9f5747d066a5a6878eed6546b84d769

Observation 2c58c819-1f7a-4a70-bb93-684284bb516c · outbound

This paper cites XModBench [41] further extends cross-modal consistency evaluation to tri-modal settings involving text, vision, and audio.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs XModBench [41] further extends cross-modal consistency evaluation to tri-modal settings involving text, vision, and audio

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:27.313147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:27.313147Z digest=sha256:1ecf15f0fbede0731d5eb7906165a40027a3fce0941ff4c7738584265eab1352

Observation a6958987-e9fb-4823-ad12-a23b65922f55 · outbound

This paper cites an unresolved cited work.

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs Unresolved cited work

Reference 2025

Resolution
parse uncertain
no resolver link, observed 2026-08-03T00:55:25.958444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:25.958444Z digest=sha256:35f05071af2c78ee25fac7bef37a88309efcf32be4d5c62b991d656ccc39cbe1

Pith citing papers

No inbound Pith citation observations are available.