Pith. sign in

Paper Citation Record · LEDGER

Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2503.02199.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.02199 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:57:50.355178Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0fa55a5-b713-49fc-9f3b-ff77372806c6 · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.355178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.355178Z digest=sha256:3636eee4414209f40cf450d7646fe9eede061941e862e9d99928b49d2b22286d

Observation 921bb0d0-f044-467b-b493-b92e6524b6c4 · inbound

How Do Vision-Language Models Process Conflicting Information Across Modalities? cites this paper.

How Do Vision-Language Models Process Conflicting Information Across Modalities? Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:03.900892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:47:03.900892Z digest=sha256:d335cc7016e606ebb6baade5aabbc20264f92d0c52f3c8a61f0b74d6a2b200bc

Observation 4f818ba9-efa0-41f6-a64f-7df6b755f483 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.114952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.114952Z digest=sha256:657b4f3294be8045373c98ca93645e21c8591cb3a52c0096001f18f1764f9a73

Observation 657d33d3-96b3-4675-bcd3-7d6e1fa4e4f6 · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:26.259597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:26.259597Z digest=sha256:c2131816afd20d73650683af29ae893240a4111246d63835eddccfb660b0e96d

Observation 38ae6aba-df00-41ad-8310-1b471635598f · inbound

Challenges in Understanding Modality Conflict in Vision-Language Models cites this paper.

Challenges in Understanding Modality Conflict in Vision-Language Models Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T11:26:17.339378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:26:17.339378Z digest=sha256:8102d87f7b0edc55f59419efa1b67bbc64d493a1abbda295ccfa06f7226ba312

Observation 073cb235-3660-4485-a335-80e01dbba94e · inbound

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models cites this paper.

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:15:50.361352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:11:01.039309Z digest=sha256:1c6548d0187bb36503b0401efd94cf0853801243bec0854b6636d6746051d363

Observation 3db06706-e5dd-47a6-bd64-883674fdd09b · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.881249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:6cbc9721ef9b7b15547a9eb5425719736d100a20ee35acc00920ebba3af24e09

Observation af8cdcdf-708e-4b3c-80ae-72edf47fa4cf · inbound

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices cites this paper.

Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T13:02:42.673767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:02:42.673767Z digest=sha256:f418ad1c5af4840f965c379d5b2b9b84cca97c765ebd381153382c60387a744b