Pith. sign in

Paper Citation Record · LEDGER

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

As of 10 August 2026, this Paper Citation Record lists 1 of 1 outbound references and 12 inbound Pith citation observations for arXiv:2508.10552.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10552 v1

Coverage vector

measured 1 of 1 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:25:06.520011Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:47:44.047025Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:59:50.274843Z

Reference resolution

1 of 1 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8907c983-dd9a-4091-a5e0-15aeec607250 · outbound

This paper cites When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models.

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:25:06.520011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:25:06.520011Z digest=sha256:e6535d311216811762db02e85277b612d530f1273f20140d993d8ed8624f07c1

Pith citing papers

Observation 8907c983-dd9a-4091-a5e0-15aeec607250 · inbound

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models cites this paper.

When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:25:06.520011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:25:06.520011Z digest=sha256:e6535d311216811762db02e85277b612d530f1273f20140d993d8ed8624f07c1

Observation 6afbbff3-4ac5-4c18-a358-39f24d832a0d · inbound

Token-Efficient Multimodal Reasoning via Image Prompt Packaging cites this paper.

Token-Efficient Multimodal Reasoning via Image Prompt Packaging When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:53:19.621054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T21:52:25.713970Z digest=sha256:216157ec9c530c2250c8ef4826abc6862d548f5d8c216be9413136da4fabd887

Observation e203d505-c702-443f-9dc0-1889ae26d7f0 · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:46.771486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:a843b7b7a7251c1a95dc88a6cafac0a319eb5ae89cf23a36842518645e0959db

Observation feadaba6-dedb-41a3-8d0e-df473df2b503 · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.190847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:ef55f0bb7009d0c37804a5a8f2fe47404cc0322e2e288a7dc1baa0403d57a016

Observation 4822d922-35d1-4da0-bc30-a3e6167eb3a4 · inbound

Information Router for Mitigating Modality Dominance in Vision-Language Models cites this paper.

Information Router for Mitigating Modality Dominance in Vision-Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.163886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:0e14c93a6837ceb3a87821490afae808b8a7c32e043610eeb54e9da1357b8893

Observation 1e21e9b8-db5c-4a07-b01d-92667dfef6aa · inbound

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment cites this paper.

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:07.568028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T21:56:38.004300Z digest=sha256:b50d765b27647afce545c71e126413a25d65a550b79b485d79c6206224a0f173

Observation 0557267d-94f2-42bf-8900-6dedf3e3b4ab · inbound

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation cites this paper.

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.940714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:04:50.483329Z digest=sha256:5b797e67811c4c494cfd99274559fbb46644d3efb5a7bf832bfa5c197305f10d

Observation d88d3522-0088-462e-860d-1a2fb0ac740a · inbound

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs cites this paper.

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:39:24.641008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T19:36:40.901674Z digest=sha256:1ee9e4d35d2935e5f676b2e7c625d0dfdf55d7b7c317bb419d926e4c301056ef

Observation 108ed84f-519d-4932-99c9-927fa4411a9b · inbound

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models cites this paper.

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.276916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T07:28:54.731659Z digest=sha256:49f8de13c0a2bcbb07555c190ad2a9b5bfb313fd66300b8c8b8efa74a902840e

Observation 3743324e-f58d-479e-9d1f-e5a99017c5e8 · inbound

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection cites this paper.

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T07:04:04.812401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:04:04.812401Z digest=sha256:e8ff0152daf85d326847cc248ed87177bd903a729620ae096c25d8fc8f1e2972

Observation 1292bf7c-cbee-4843-8715-43181321eafc · inbound

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs cites this paper.

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:38.657616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:38.657616Z digest=sha256:594e54263d89204a00049c89ad4ab8739cdd09dfd3943991144d3b2157d2b3f9

Observation fe3929af-e1b9-4de8-a0b4-6ce2f51ee134 · inbound

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination cites this paper.

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:47:44.047025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:47:44.047025Z digest=sha256:8a67ec5768fe94b3a86c81b1854b18da9a9ed367388efefa3831d900d17b1b40