Pith. sign in

Paper Citation Record · LEDGER

Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2404.18624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.18624 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:43:20.913891Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:06:16.588829Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a743814e-47b3-472e-ad4b-8c60f1159b28 · inbound

Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence cites this paper.

Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:43:20.913891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:43:20.913891Z digest=sha256:a15f3443b45b4a6f217160d2a2db162f44e0f1100209aeaf98e10c0efe624aff

Observation c65432de-aa7e-42d4-b967-18ee6a206c22 · inbound

On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI cites this paper.

On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:23.223689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:23.223689Z digest=sha256:223e3efc813924d6df2b80d8dcc3e34a99491675b0b851fa71bab912177b4fa9

Observation c7eeb6ec-974e-42a0-905d-091b4df582d1 · inbound

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't cites this paper.

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:00:49.710005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:24:48.212344Z digest=sha256:e1c4371577039f774f4b5a0565316873f796221d3ee8eb7a006ca53196b1d711

Observation 92e8d1eb-a8ea-4443-b7dd-505820fb9694 · inbound

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement cites this paper.

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:21.408263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T05:44:05.093926Z digest=sha256:5670670848a4691773b7008f3cf1777a7a700fd9e19b253463ad7b9011aed9d5

Observation 88efbdee-8564-4b3e-ba26-903f1452f388 · inbound

Medical Context Distorts Decisions in Clinical Vision Language Models cites this paper.

Medical Context Distorts Decisions in Clinical Vision Language Models Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:03:24.795322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T14:59:11.496757Z digest=sha256:171d2fa2842db4dc486efdff776d8b19fd844879e98984b295f73ec3644b87c0

Observation 362d4bae-2551-4d28-8683-906cb4e3e56e · inbound

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning cites this paper.

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.590145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T15:45:26.891621Z digest=sha256:22dda501c4d9cced43d6bdf93aad5a5a3e6924475338eba8fc88118f326d99b7