Pith. sign in

Paper Citation Record · LEDGER

Vision language models are blind: Failing to translate detailed visual features into words

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2407.06581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06581 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:13.161806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 85364f4f-4887-461d-a6f9-59f4f65d3781 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Vision language models are blind: Failing to translate detailed visual features into words

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:13.161806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:13.161806Z digest=sha256:85e12ca2dd7d7bd9e990e9b80f7e2b553136b08ccb4091c75f3b52ae90f23688

Observation 2d172e97-d16d-4b23-96e8-07f47b5f7817 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:57:13.399915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:be86f43ea25a77f7c175dd02532d33e23d2725fbd7614685ea930e745bb8d5aa

Observation 90ab0e07-92d8-469e-a697-85fb8720ac3f · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.234599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ad3a3e9199f77115636ef69116a5a5e58a686e82b1908affa9b7048b1b0189e3

Observation fea50c66-695f-4e0d-8199-6795462b794a · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Vision language models are blind: Failing to translate detailed visual features into words

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.991717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:31.991717Z digest=sha256:d5aa4c0f5652272a6472ae5b8425c001d5a6d30df47b43660919e366128bc10b

Observation 25b7d359-8125-4274-9e8c-390a39552d4b · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.218019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:6bf26d07a93546042e3cf140b44c82f19ab3e70f8b5e936b76dab9c1ed1b092e

Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.204341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.204341Z digest=sha256:c892cd4ad56aa855a97c1b10db6c072df3a0ac2254b0ef80a0c2df5682b46ae6

Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · inbound

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation cites this paper.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.423858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.423858Z digest=sha256:c2ab7aa3b194dc47467afbace64e26469b8303f43140b11ee5a3fde2b173924d

Observation 8bd60bd1-6dd6-41cf-8065-10bf0bc2c12c · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:07.991005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:1a3380a64d045d327828ee8437ab94f1b1c38c42477973ad51bc37f09bffbf98

Observation 54beb4b6-eab4-4c6a-a62f-4932d1ec6ffa · inbound

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production cites this paper.

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:25:34.429111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:25:34.429111Z digest=sha256:946ff0cb0f14b96180663c45d7b6ca9c802323333a8d1ef6801fb04b9e0954dd

Observation f940b53a-8cb9-47ba-ba9b-7bc8d974cbef · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:10.881826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:10.881826Z digest=sha256:00c6771ac3699f28999a5dd41b3f8415554990e43bcd63a34ae6835bc4b6e283

Observation 50c24374-7d38-43d7-8e26-b3ad50485149 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.704948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:169b4c399387b194feb2948b0f03535ef8609e2b7686cacb911d1df5ecfe4acd

Observation fb0aee7d-01c0-4023-a49a-babad6dda0c5 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:37c7c7ea7b363d58aa80f2b4c3d5f90f6c4ba61b00c40e7e680c316dc1dd9548

Observation 5f717ea5-af40-4dee-a8c1-6ba866d69223 · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory Vision language models are blind: Failing to translate detailed visual features into words

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:11:02.305599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:ea1712e4c0f8127077a2c2ed1cfd3ae997de3c5610b7b3ca13fd8a375ff093c7

Observation 507e832f-3a9e-4d07-b76f-eac62eab3c70 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.274527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:04415089775cdead4db31a1aad43afad9ac786d448dbc946e78accd53eadd6e1

Observation 53b4db02-75ec-40cf-9ebc-67c356f64060 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.871090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:2d82ccb20326bd939125531389484d1762afe4248a5fdddcd3e8baab77818a5c

Observation 74437cb0-9d1c-4e63-bed6-ff4a7af29f26 · inbound

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? cites this paper.

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Vision language models are blind: Failing to translate detailed visual features into words

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:36:18.598684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:35:39.884350Z digest=sha256:1ff0fa4c343ce5bde4fc5f56badd5aded0199f34d3166d481fa1808bacc8143c

Observation fdb723b4-7ba4-4341-8e5a-1cc979e0f73d · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:59:45.707454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:3abb3c85e25ca135d283c2e8aa8dab4b0e346c1b9cca695c3923bee38ed5d57d

Observation a0e4a19a-dfff-4ede-84aa-bc715d538aab · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:54:57.786081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:b8f36f38b25c649eed23bd5ba593bd5bc3a4c5c124b429f778fa4098f06678e0

Observation e5f8806d-6710-47e2-a050-161487d2d9e5 · inbound

Binding Visual Features Point by Point cites this paper.

Binding Visual Features Point by Point Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:03.285291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:56:44.793896Z digest=sha256:5fc7d10d33551fec1c24b2499acdc6eb176124293bd6921ad3754ccaaf5bdf8f

Observation d137ed69-6c52-4aad-b193-8cee47b04f23 · inbound

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes cites this paper.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Vision language models are blind: Failing to translate detailed visual features into words

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.666857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:2a2445761724711a0f835c9f618dd0f610dd859bda3c5dd54b8e7b07c6ae4a3a

Observation e26043c4-e69a-4fba-bf18-736ff1b56518 · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.051325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:90e209f1ee1eeb87cd9b75b57733feb6b5b425d47ded9d542ff1f5e6a823875c

Observation 8068ad65-ea6a-480c-9149-ba38d04dddc3 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Vision language models are blind: Failing to translate detailed visual features into words

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.978721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:cba80f181d1109e5ba8c70c417a01c085b3befefaab7c29fd9478db86d57cdb2

Observation a38d71dd-185c-4860-8360-5c4924cca342 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.439678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:e7a523d9ad213a9efda7b58f569b7d3fe73c5825e45f4a848a59ccf0c2da965d

Observation 58793a2d-0661-4d17-a12c-6c5c5216f1e3 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Vision language models are blind: Failing to translate detailed visual features into words

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.076380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:09dc695ba4c80e67fed35d7b10396e67d80aa6431b120bffc1996c1447b537a9

Observation 287748e6-e590-4333-866d-e815b816a8c8 · inbound

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR cites this paper.

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR Vision language models are blind: Failing to translate detailed visual features into words

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:31.073799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:32:39.976049Z digest=sha256:6c1d8c2d2327c0d6ffa2bfb3edf0fca184aee7a74a0b3eb19a76aa99400d09fd

Observation 1626bef0-46b0-43d5-9dcb-86301cac682a · inbound

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms cites this paper.

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms Vision language models are blind: Failing to translate detailed visual features into words

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:10:02.337868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T22:58:41.991573Z digest=sha256:0f5fb014b840ea8d9e3dd3c250e49ba1ed5a7b2b29e1cfa3d92251dd4c5ea492

Observation fe5085fe-7473-43f1-869e-7b787f611cb6 · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.238996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:24:11.228672Z digest=sha256:b05dfcb5284ea146fb0d1a2e22a7985f01e3ba7c28f1a419423c1aa833f45a8e

Observation 6dd3fd32-7674-4264-9f5e-f80b07eefd3f · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T11:58:52.700648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:58:52.700648Z digest=sha256:3eba495acf1cfa474bf6d9f6b7b060be66948d8f113b0ac6156054037411b9e6

Observation 90a92488-f13b-4b92-a37b-504dac08ccda · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:42:46.146973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:42:46.146973Z digest=sha256:444e73e43f8d51f74649adbd58e6b23b107022132bd8be1318c4106d83b3d6d1

Observation b9c3781e-bf2f-49f8-8d26-6fbf92dacf31 · inbound

Information-Regularized Attention for Visual-Centric Reasoning cites this paper.

Information-Regularized Attention for Visual-Centric Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.256397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T15:12:43.475802Z digest=sha256:f179c68be65c1d79a910a608a47db53175b84b69b0cbf1ede47a653ebbd60def

Observation 64fac733-b34b-4efc-8a60-fa541118694f · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers Vision language models are blind: Failing to translate detailed visual features into words

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:06.631736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:06.631736Z digest=sha256:4e22cbfd0feb16aeb0cd6de1870e78acacc32cbbf7ee0b6838c48553bf2b53fa