Pith. sign in

Paper Citation Record · LEDGER

Information Router for Mitigating Modality Dominance in Vision-Language Models

As of 5 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2604.16264.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16264 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:41:54.503158Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact7
  • verified fuzzy17
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3275d42a-0d48-4709-806f-3f04d16565f8 · outbound

This paper cites Internvl: Scaling up vi- sion foundation models and aligning for generic visual- linguistic tasks.

Information Router for Mitigating Modality Dominance in Vision-Language Models Internvl: Scaling up vi- sion foundation models and aligning for generic visual- linguistic tasks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.389505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:c54ba4c0d13e8ffcc8b788bdf7f793ef1ba4094793d7fd2790078696f19f1b39

Observation e66960ee-d957-41c2-87ee-c3fb35f28fb5 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

Information Router for Mitigating Modality Dominance in Vision-Language Models Vilt: Vision-and-language transformer without convolution or region supervision

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.406078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:cd878a0694b6b74d1a5d273eb0f249f2e8c3532c0a07386cb9c459f347c2a293

Observation ce9cfe12-bf1e-4f51-99a7-918e036feabb · outbound

This paper cites Can vlms actually see and read? a sur- vey on modality collapse in vision-language models.

Information Router for Mitigating Modality Dominance in Vision-Language Models Can vlms actually see and read? a sur- vey on modality collapse in vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.397570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:d33c05847f04b6e0b74ca2db5f4572f1d25a0f18f05007dce7825a8900c14b50

Observation 51251733-3a7c-4e34-8f7c-83370dc76356 · outbound

This paper cites Multimodal large language mod- els: A survey.

Information Router for Mitigating Modality Dominance in Vision-Language Models Multimodal large language mod- els: A survey

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.395573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:b995bccc582c28862d0ab2e447ce98e4be08c059844d04ea5622898099cc43e6

Observation 6cddc04a-31f1-403d-a843-fc49f5e3d9cb · outbound

This paper cites Synthesize diagnose and optimize: To- wards fine-grained vision-language understanding.

Information Router for Mitigating Modality Dominance in Vision-Language Models Synthesize diagnose and optimize: To- wards fine-grained vision-language understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.391567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:c9dcea7986b95f2eec64b5576f75471a6b39dabbe5c97fb18aea5ed606b210ba

Observation ba85d94f-5558-4b8a-b32b-2ed2d00d18cc · outbound

This paper cites Mllm as video narrator: Mitigating modality imbalance in video moment retrieval.

Information Router for Mitigating Modality Dominance in Vision-Language Models Mllm as video narrator: Mitigating modality imbalance in video moment retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.399764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:0b760b9fb9a697eff884144dcd5c18e2f9ad334a4d5163510dc250e76f14cc5f

Observation fd5e1827-8cc6-4b4d-8519-905cae0d5b9f · outbound

This paper cites Learn to explain: Multi- modal reasoning via thought chains for science question answering.

Information Router for Mitigating Modality Dominance in Vision-Language Models Learn to explain: Multi- modal reasoning via thought chains for science question answering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.412232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:89143368a22bd2caba234cbc736b37fda6647c50e1d3ff9aac3d68b0cf6e723e

Observation 0f6df2ba-8624-4f9d-a601-73480860c380 · outbound

This paper cites Improved baselines with visual instruction tun- ing.

Information Router for Mitigating Modality Dominance in Vision-Language Models Improved baselines with visual instruction tun- ing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.416780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:f544deec8a947cd7b945bd870077fcc82b6e618ebc09d1fa49c62332ae40c9fd

Observation 4acc3667-87ca-44d5-b43c-717949272b72 · outbound

This paper cites Qwen2.5-VL Technical Report.

Information Router for Mitigating Modality Dominance in Vision-Language Models Qwen2.5-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T08:43:01.156032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:bc433ffea0b9b3976258b82c78dbac6be4d4db6b153b6b02df17afaf3624d9f7

Observation 01b74e83-1f91-4543-a190-85634bdc049d · outbound

This paper cites Lora: Low-rank adaptation of large lan- guage models.

Information Router for Mitigating Modality Dominance in Vision-Language Models Lora: Low-rank adaptation of large lan- guage models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.393549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:812043edaf07a448ea19704a2cb64a403a78724f119787bd4c2d5b2529b86348

Observation 66b1e7de-c85d-48b7-ac36-99e226c02fec · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Information Router for Mitigating Modality Dominance in Vision-Language Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.422746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:aad27b5e914a512717257f6e996022afc21f3cd7486fd45eb7d4a5612d246d1f

Observation 95d0e755-37dc-4c7a-91eb-251431ab7159 · outbound

This paper cites Mmbench- video: A long-form multi-shot benchmark for holistic video understanding.

Information Router for Mitigating Modality Dominance in Vision-Language Models Mmbench- video: A long-form multi-shot benchmark for holistic video understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.410307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:520926df0a71ec6d0b84c6ed93bb8c0983e41063c727c69ca93442f32ee426bd

Observation 4822d922-35d1-4da0-bc30-a3e6167eb3a4 · outbound

This paper cites When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models.

Information Router for Mitigating Modality Dominance in Vision-Language Models When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.163886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:f0517606a7bcfddfeff994ff29779d6a802be6cca89ff8a431db7e0a52cbbc85

Observation 45015d1f-d69f-4a0c-9ad4-2162c9e1b89a · outbound

This paper cites Mitigating modality collapse in multimodal vaes via impartial optimization.

Information Router for Mitigating Modality Dominance in Vision-Language Models Mitigating modality collapse in multimodal vaes via impartial optimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.404105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:ec5ab4b3f1101dc3995467d3f057081fbf4c11b3fecaca885a7571a402ba8685

Observation 510e0e70-c6fc-4581-adcb-664ff92539aa · outbound

This paper cites Assessing modality bias in video question answering benchmarks with multimodal large language models.

Information Router for Mitigating Modality Dominance in Vision-Language Models Assessing modality bias in video question answering benchmarks with multimodal large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.418704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:bb351ae5d3c591d7b8e55b07a2846b5d33f1730151762bc5c4c37810c9d71835

Observation 6c794de5-5000-4821-8da2-b93d4501d4b0 · outbound

This paper cites HEX: Hierarchical Emergence Exploitation in Self-Supervised Algorithms.

Information Router for Mitigating Modality Dominance in Vision-Language Models HEX: Hierarchical Emergence Exploitation in Self-Supervised Algorithms

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.146562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:1367b59aaf380a220ee1264e4ffbd172dc96a4c9fb1445e155539615baa70f47

Observation 766ef1a4-2ddc-4d11-8b5b-8f5bb5365651 · outbound

This paper cites Countering multi- modal representation collapse through rank-targeted fu- sion.

Information Router for Mitigating Modality Dominance in Vision-Language Models Countering multi- modal representation collapse through rank-targeted fu- sion

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.149746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:69450e6f9066933c4269f4b13f1eb2f53a8cfe4743355427c6c9040de3f93397

Observation 17ac0f16-b9c4-4ddb-b4b4-f8eff948a24f · outbound

This paper cites arXiv preprint arXiv:2505.12576 , year=.

Information Router for Mitigating Modality Dominance in Vision-Language Models arXiv preprint arXiv:2505.12576 , year=

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:43:01.159640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:6e3ae2caa49d8e0d4d78c121dd786b840e57f5563c1804e7865a3cd64a530494

Observation a11ee470-c5d9-4beb-87c5-224c99a5d9e5 · outbound

This paper cites Multimodal Deep Learning.

Information Router for Mitigating Modality Dominance in Vision-Language Models Multimodal Deep Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.143127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:458bb89f385b208f03af8afbb561f075ea0fff3aa049e3787be90a6439b1fc7a

Observation 70d4b906-47d1-46fd-9a2c-5f875efbc746 · outbound

This paper cites Hierarchical and Multimodal Data for Daily Activity Understanding.

Information Router for Mitigating Modality Dominance in Vision-Language Models Hierarchical and Multimodal Data for Daily Activity Understanding

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-08-03T02:10:27.131456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:c25e07818ad52e97ea57522d38f0444dcb8d95cdf305eb52b095b4b9557f825e

Observation 662ff9ba-6ef4-4b1b-bf49-f2a76b02bd15 · outbound

This paper cites Multi-level and Multi-modal Action Anticipation.

Information Router for Mitigating Modality Dominance in Vision-Language Models Multi-level and Multi-modal Action Anticipation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.167703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:cf13f931779371c7595a6d3822e488ae18dcda3908cb5d69dcb6268e7d673d35

Observation 2b4cb0bd-21c6-4c3c-8761-e4a9bb78e876 · outbound

This paper cites Multimodal machine learning: A survey and taxonomy.

Information Router for Mitigating Modality Dominance in Vision-Language Models Multimodal machine learning: A survey and taxonomy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.401959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:5c5142a48e2832d0c9191c22dcadc0d21a2ee7292905eb53df7f6f27a714449f

Observation d16da830-3500-4943-85fe-ab570d773984 · outbound

This paper cites Intra-and inter-modal curriculum for multimodal learning.

Information Router for Mitigating Modality Dominance in Vision-Language Models Intra-and inter-modal curriculum for multimodal learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.414805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:06b9712abbd6889c8507af50adf9e5127703899cd36f221593cd48d87499631b

Observation 685c97da-759b-43ac-b77b-fdc487781475 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

Information Router for Mitigating Modality Dominance in Vision-Language Models Overcoming catastrophic forgetting in neural networks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.407952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:61c92ccdf0b50f79178d7d9055c55508f97c3d9787fe68b27b83ffed535c9b3c

Observation b7ccc8ea-c1c7-4d11-8a50-48a3d1ddc1a8 · outbound

This paper cites Decoupled Weight Decay Regularization.

Information Router for Mitigating Modality Dominance in Vision-Language Models Decoupled Weight Decay Regularization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-10T08:43:01.171028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:e0cd4dc2355d700ed3037c369abfad6f4f9bbe85741421e0e36f5fb982fabb20

Observation ea1f52eb-5eb4-47cf-8172-89a0b964fb5e · outbound

This paper cites The effective rank: A measure of effective dimensionality.

Information Router for Mitigating Modality Dominance in Vision-Language Models The effective rank: A measure of effective dimensionality

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:54:01.420960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:41:54.503158Z digest=sha256:8836140ecbe3a3e75a374fd11f7f752977729e2999cd1a33d8816bf36c366820

Pith citing papers

No inbound Pith citation observations are available.