Pith. sign in

Paper Citation Record · LEDGER

Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2404.18624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.18624 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:22:23.223689Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:06:16.588829Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c65432de-aa7e-42d4-b967-18ee6a206c22 · inbound

On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI cites this paper.

On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:23.223689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:23.223689Z digest=sha256:0b99cc34e5d14cdd49ea061679d528cc7735cae1fcd4c8c71e113242fd7fe48d

Observation c7eeb6ec-974e-42a0-905d-091b4df582d1 · inbound

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't cites this paper.

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:00:49.710005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:24:48.212344Z digest=sha256:e2efaa4f13f4f8e920ca5dbc0d8fa685378c957d7a4cf0c5eba78ba01070565a

Observation 92e8d1eb-a8ea-4443-b7dd-505820fb9694 · inbound

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement cites this paper.

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:21.408263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:44:05.093926Z digest=sha256:d8410d8ef63795ce3230733f1df45acc5b6cbae6cd1e5402406b94898aa84106

Observation 88efbdee-8564-4b3e-ba26-903f1452f388 · inbound

Medical Context Distorts Decisions in Clinical Vision Language Models cites this paper.

Medical Context Distorts Decisions in Clinical Vision Language Models Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:03:24.795322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T14:59:11.496757Z digest=sha256:be08824586f0b90fde5455eeaa44881ff4c21f9ec1fdc3acc79d8c6716569fb3

Observation 362d4bae-2551-4d28-8683-906cb4e3e56e · inbound

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning cites this paper.

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.590145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T15:45:26.891621Z digest=sha256:b4cd6a2f55e357ad42f9708dc2977b47e83d71cd869802e17fd853685fe18404