Pith. sign in

Paper Citation Record · LEDGER

MLLMs are Deeply Affected by Modality Bias

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2505.18657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18657 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T01:00:28.861351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.223756Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 409490f8-ec1c-4261-8af4-a4482362a558 · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability MLLMs are Deeply Affected by Modality Bias

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:28.861351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:28.861351Z digest=sha256:ff88cf1c372aecfbcc866e816b10b70748d6ee2f27fc7a075636c6e9f1d3d260

Observation 44f4ebe2-070c-4bb0-85d9-e8365943b253 · inbound

Omnidirectional Spatial Modeling from Correlated Panoramas cites this paper.

Omnidirectional Spatial Modeling from Correlated Panoramas MLLMs are Deeply Affected by Modality Bias

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:37.791379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:37.791379Z digest=sha256:9ec01f5ec92711722ea76f4b86d6d2349bdd89ba34f14630eea928d94be935ae

Observation 163ddd2a-a86a-413e-84b0-168791481c7c · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs MLLMs are Deeply Affected by Modality Bias

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:31.631473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:31.631473Z digest=sha256:1d2c9f48f95b562bcb6f317deff2ffbe32a96c27c4c6aa97d9168e54130d70d5

Observation fbd1d54a-dd4f-4ddb-b55d-2cb5b89c675f · inbound

Token-Efficient Multimodal Reasoning via Image Prompt Packaging cites this paper.

Token-Efficient Multimodal Reasoning via Image Prompt Packaging MLLMs are Deeply Affected by Modality Bias

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.601782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:52:25.713970Z digest=sha256:a48f005cf5031b72c15aa0d33f9b0f6be556067bf22b57929adca935946cf113

Observation f8d109ff-404d-49b4-b09b-18a6839cf17d · inbound

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning cites this paper.

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning MLLMs are Deeply Affected by Modality Bias

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:38:02.902544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T17:34:10.089555Z digest=sha256:c6cc3d00caeb6170571dfac764715264213c3034a13f26f522b818e5a4d072d8

Observation b07b222e-4d13-46a3-a610-d57a13e7a1b0 · inbound

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models cites this paper.

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models MLLMs are Deeply Affected by Modality Bias

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.732553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T07:02:02.752466Z digest=sha256:75b4cb219914a45fcb91d5caa26460c5a06f7b7bdee13ccf21180327cabf2a4d

Observation 4a808afb-0e03-4f25-96d0-9de968dac8e9 · inbound

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment cites this paper.

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment MLLMs are Deeply Affected by Modality Bias

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:07.562485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T21:56:38.004300Z digest=sha256:ffe34ca5d5cadf561ed55e4dde77ea7888cc53c5e390eafb24059cf2c1eb80c3

Observation 977cd533-9d06-4171-ad7e-30545c81c955 · inbound

Do Composed Image Retrieval Benchmarks Require Multimodal Composition? cites this paper.

Do Composed Image Retrieval Benchmarks Require Multimodal Composition? MLLMs are Deeply Affected by Modality Bias

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:23:44.429486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T21:20:49.609788Z digest=sha256:40b4637d5a1662cf98ccef09d3192b31f5829f4e8d589532b823cae1ba3cc990

Observation cc28aa76-c4c2-43ba-8e65-61703fccfcf6 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models MLLMs are Deeply Affected by Modality Bias

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.210167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:b612bb10956caa2fbcbb7b6ac657835df700fc358d307d3dceef4a5287944f19

Observation 69715a54-940a-4a56-bd7d-4de3a12c796a · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models MLLMs are Deeply Affected by Modality Bias

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.347720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:b8c2a7d898c8b246cf99812507eed87e7ed298d5b7427fbb91bb831fd7bb911b

Observation 762b4ad1-b48e-40dd-95e3-99e53e76a9a2 · inbound

Measuring Cross-Modal Synergy: A Benchmark for VLM Explainability cites this paper.

Measuring Cross-Modal Synergy: A Benchmark for VLM Explainability MLLMs are Deeply Affected by Modality Bias

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.750971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:05:34.039507Z digest=sha256:30dbdc03efd8c7ffff3c64977ac8d425cbe9e2020f2cdd638625b5119d5d7a00

Observation 7a9952f2-aa1f-416c-a298-703ef7fd3898 · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems MLLMs are Deeply Affected by Modality Bias

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.882322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:b52c20933dec6d8126f520bdada3eab6b7563b2d53efb109b82f18c653a567b7

Observation 4c0022df-3301-495b-a119-c24451d8c3f6 · inbound

Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models via Pareto-Optimal Gradient Integration cites this paper.

Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models via Pareto-Optimal Gradient Integration MLLMs are Deeply Affected by Modality Bias

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.380374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:31:02.911020Z digest=sha256:d1ba1191a99f766ed5158bbde68bd994456e199f1f309f2c10a8d6cf7483a269

Observation d21d4326-f5ba-4049-a2d6-e75804a0f5a1 · inbound

Vesta: A Generalist Embodied Reasoning Model cites this paper.

Vesta: A Generalist Embodied Reasoning Model MLLMs are Deeply Affected by Modality Bias

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:29:35.710669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T16:55:12.518255Z digest=sha256:969f8f943c971e4776c77d6f78194774aabe63e8b0380ac3b21c3f5d33754b0f

Observation 9f8ed3fd-eddc-4c01-be3d-3b7158916600 · inbound

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models cites this paper.

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models MLLMs are Deeply Affected by Modality Bias

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.226221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:28:54.731659Z digest=sha256:d4b1dff68cf2970da81966d1c88db831886fb3fc2387ff4f07e21b329fb80e8a

Observation 2428e129-1e3e-41a7-a55b-7e035829ebef · inbound

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs cites this paper.

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs MLLMs are Deeply Affected by Modality Bias

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:39.041606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:39.041606Z digest=sha256:a62ded14432c282248e6b8d3346f771ead0ea6b084fc87baddc2b61f85e2bdc9