Pith. sign in

Paper Citation Record · LEDGER

Vision language models are blind: Failing to translate detailed visual features into words

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2407.06581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.06581 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:51:35.338457Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f7b1ae63-f76c-4a55-a981-3775661d04d5 · inbound

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games cites this paper.

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games Vision language models are blind: Failing to translate detailed visual features into words

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T16:21:08.316762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:21:08.316762Z digest=sha256:08d0bb0f955594b5d17fc0ef8164debc5dfade69c5e4ed2848be87273637a178

Observation a475d0bc-3f9e-4cb2-ae1e-2501bbea77d6 · inbound

De-biased Multimodal Electrocardiogram Analysis cites this paper.

De-biased Multimodal Electrocardiogram Analysis Vision language models are blind: Failing to translate detailed visual features into words

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:09.967925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:58:09.967925Z digest=sha256:420e2e6dda78728c295038bc19fe9fa09e37345b1daea26db98aa184afcbd216

Observation cb4b85c9-31a7-49a8-a768-9ea4b3c326fe · inbound

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? cites this paper.

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? Vision language models are blind: Failing to translate detailed visual features into words

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T23:19:11.621843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:19:11.621843Z digest=sha256:b342d5fec7143f9174007fe533e704bd0d3faa755688df59cb78bae0182c367f

Observation 39761a1e-ab19-4615-912c-6fc91ee58bac · inbound

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models cites this paper.

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:53:22.538799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:53:22.538799Z digest=sha256:0e62d10c1a7937f4ad05570948158d116e3883fd8f306278f686092a8e250a41

Observation 1ab5e6f0-e331-4e55-beb0-0b0d50c531e6 · inbound

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision cites this paper.

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision Vision language models are blind: Failing to translate detailed visual features into words

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T21:34:35.450573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:34:35.450573Z digest=sha256:1bb68bb0845bf6a304c5dbc2160de93616784d1cc33e8dc5344707e32c4713d2

Observation f2fc676e-e135-4764-98bc-4997bc800d64 · inbound

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models cites this paper.

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:47:58.990688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:47:58.990688Z digest=sha256:23962849ed93a71abe39a6b61e6fa9a11a5b15bcd20bb374f18d9ef548ef08aa

Observation 4492f9b0-c03f-4c34-9738-dd928133df44 · inbound

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models cites this paper.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.238672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.238672Z digest=sha256:3ea4f9d5dad64ca360688e1cf2a5e9b7d097d33408693bbfe98768df0f7b94cf

Observation 85364f4f-4887-461d-a6f9-59f4f65d3781 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Vision language models are blind: Failing to translate detailed visual features into words

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:13.161806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:13.161806Z digest=sha256:7bb56857877237d93909148bd2478d000876b1ed05489eaf8f56926c3435ac00

Observation 2d172e97-d16d-4b23-96e8-07f47b5f7817 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:57:13.399915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:670390879a0c8c14e591e07051504caadf57ec9f0a19068ec2a9b932f47cdc17

Observation 22733819-6813-4eb6-b3fb-90380aba5c39 · inbound

VizTA: Enhancing Comprehension of Distributional Visualization with Visual-Lexical Fused Conversational Interface cites this paper.

VizTA: Enhancing Comprehension of Distributional Visualization with Visual-Lexical Fused Conversational Interface Vision language models are blind: Failing to translate detailed visual features into words

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:51:35.338457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:51:35.338457Z digest=sha256:9e57de341a76fd0622f5f76ac7e0323e23b7a80d5d13c91e2d3539a06dfe31bd

Observation 90ab0e07-92d8-469e-a697-85fb8720ac3f · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.234599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c7b03ea039be1a4ba212b84b71a286ea9142b1ffea6820125a9165dc7770bc68

Observation 78df9661-6317-4dbe-a284-58cccc982189 · inbound

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels cites this paper.

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels Vision language models are blind: Failing to translate detailed visual features into words

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:16:58.822948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:16:58.822948Z digest=sha256:c55466b66cf96ee6eb42e00cb939d217970961ae057f4c77309180793b67846b

Observation fea50c66-695f-4e0d-8199-6795462b794a · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Vision language models are blind: Failing to translate detailed visual features into words

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.991717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:31.991717Z digest=sha256:fe5f85675d274f7f260c35a433f344caeaaacaa3cc6f7e0c372e992977378116

Observation 25b7d359-8125-4274-9e8c-390a39552d4b · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.218019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:33f80e0d35a4ed2d86c512604ab4c3a94b313e39b21377ff274e44e6884dd397

Observation 4a7c6114-3b17-4b08-a388-ef2a0bdc0a10 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.204341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.204341Z digest=sha256:6d0707c623d3efc99afcca5c01bf7ebb5370b394cf3d343904ecd3605ebcb80c

Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · inbound

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation cites this paper.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.423858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.423858Z digest=sha256:d6e3cc35d73f3f0563a49dc2a89026aeffd2ddfcbffc78ddd4edac9e763dc55a

Observation 8bd60bd1-6dd6-41cf-8065-10bf0bc2c12c · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:07.991005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:fd0d301f0e7fe51c9ec6d469f19bdd8ab0c51deffc17d09697f4a9b8ca17c13a

Observation 54beb4b6-eab4-4c6a-a62f-4932d1ec6ffa · inbound

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production cites this paper.

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:25:34.429111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:25:34.429111Z digest=sha256:95c413af0789e4de525a3653c1373cb3ce1724e49dd05d3d843af96ea4041cd4

Observation 7f484ea5-da3b-4100-8470-099fa43f2cb0 · inbound

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks cites this paper.

Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:16:15.699872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:16:15.699872Z digest=sha256:bf414cea1d26d5cda47f816a93cae1e0e58c585e367172f165be9dcec3c2bcad

Observation f940b53a-8cb9-47ba-ba9b-7bc8d974cbef · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:10.881826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:10.881826Z digest=sha256:89f0b252a40776cd6e3ddf2aa54c3f7b64191348c88249a61c772ddd43ffdc2c

Observation 50c24374-7d38-43d7-8e26-b3ad50485149 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Vision language models are blind: Failing to translate detailed visual features into words

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.704948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:c2f549829a1e24af0da3061402a9939bb9890ff3459d33cc25b301830471abcf

Observation fb0aee7d-01c0-4023-a49a-babad6dda0c5 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:75523775d1d359d8d02f74e67539a41642ba2c9f12e74650c887e615afbe6643

Observation 5f717ea5-af40-4dee-a8c1-6ba866d69223 · inbound

ReflectCAP: Detailed Image Captioning with Reflective Memory cites this paper.

ReflectCAP: Detailed Image Captioning with Reflective Memory Vision language models are blind: Failing to translate detailed visual features into words

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:11:02.305599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:11:10.598413Z digest=sha256:9cc94ca79c9e028a8b51174ea453dcd62f99bf165dd417fffbb8e8e7e841d505

Observation 507e832f-3a9e-4d07-b76f-eac62eab3c70 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.274527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:ee9af0d291656bc3a4c0921eb50f12274be6c4fe0584a12e9739e429d23a5375

Observation 53b4db02-75ec-40cf-9ebc-67c356f64060 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.871090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:0118d350878fd448521c9b9e4be48c2d949edddc06378515cc96a9cca0774ac9

Observation 74437cb0-9d1c-4e63-bed6-ff4a7af29f26 · inbound

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? cites this paper.

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All? Vision language models are blind: Failing to translate detailed visual features into words

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:36:18.598684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:35:39.884350Z digest=sha256:b8c5b9ce7ac67fd0eace309c2c5a843f8e57c58d8eb258b6539e07cb4574f4b1

Observation fdb723b4-7ba4-4341-8e5a-1cc979e0f73d · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:59:45.707454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T06:55:04.657347Z digest=sha256:df15020c4fe7c738d0ed5b80a9360f5010391016aae1c28850d069e20394c749

Observation a0e4a19a-dfff-4ede-84aa-bc715d538aab · inbound

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? cites this paper.

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:54:57.786081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T17:52:26.785086Z digest=sha256:d8fcc8dbf0e7ac679f5e419d692e69d6eea7ab0c1d92fd1623fda3728208f888

Observation e5f8806d-6710-47e2-a050-161487d2d9e5 · inbound

Binding Visual Features Point by Point cites this paper.

Binding Visual Features Point by Point Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:03.285291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:56:44.793896Z digest=sha256:02e54751ab1597bb8b8c033ff03d95c0da0a3535d2b068902974f99922ff7552

Observation d137ed69-6c52-4aad-b193-8cee47b04f23 · inbound

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes cites this paper.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Vision language models are blind: Failing to translate detailed visual features into words

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.666857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:020616a45368dbef1c9a95d84747abf2c30737c3768eb18c4f12641d1692c471

Observation e26043c4-e69a-4fba-bf18-736ff1b56518 · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding Vision language models are blind: Failing to translate detailed visual features into words

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.051325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:0a32029c02f0f790cb6d55d1a4f6e52e18a0af8d1bf29c150bb51c1f7e5ae0bc

Observation 8068ad65-ea6a-480c-9149-ba38d04dddc3 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Vision language models are blind: Failing to translate detailed visual features into words

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.978721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:4034111bd0129e3ba4ce3448601616335e0eda3d940516868ab9d6ffdefae4b6

Observation a38d71dd-185c-4860-8360-5c4924cca342 · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.439678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:2796bfa457e6209dd0b357ef0ed399956d8087ea6edbc5e2fc795a9075677001

Observation 58793a2d-0661-4d17-a12c-6c5c5216f1e3 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Vision language models are blind: Failing to translate detailed visual features into words

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.076380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:726ad86ee1c445bb3e9bbd7ed27a49a2f864d144fda9ebffe5f20137c3693515

Observation 287748e6-e590-4333-866d-e815b816a8c8 · inbound

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR cites this paper.

Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR Vision language models are blind: Failing to translate detailed visual features into words

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:31.073799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T17:32:39.976049Z digest=sha256:fa40fcacb317360b3a93360e238f9b5fb1116b73e8692494a55e0e30d9de49a2

Observation 1626bef0-46b0-43d5-9dcb-86301cac682a · inbound

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms cites this paper.

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms Vision language models are blind: Failing to translate detailed visual features into words

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:10:02.337868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-25T22:58:41.991573Z digest=sha256:21591792ef4a83fe3694e529cfb26af8828462a6427fb8e24e8d9770173229eb

Observation fe5085fe-7473-43f1-869e-7b787f611cb6 · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.238996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T05:24:11.228672Z digest=sha256:e62a853327499ac7a48a7f3fb689575079f692fe65af160c6387dffb8b67c936

Observation 6dd3fd32-7674-4264-9f5e-f80b07eefd3f · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T11:58:52.700648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:58:52.700648Z digest=sha256:ed285acde14cb4ed44873c16272e8d86a6f279242aece4992e4ed494e179fdf8

Observation 90a92488-f13b-4b92-a37b-504dac08ccda · inbound

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals cites this paper.

The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals Vision language models are blind: Failing to translate detailed visual features into words

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:42:46.146973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:42:46.146973Z digest=sha256:b20f4083ae2b2b0063eacbbbfc24bc53b324b425a9ec567d02ce22193dfc562d

Observation b9c3781e-bf2f-49f8-8d26-6fbf92dacf31 · inbound

Information-Regularized Attention for Visual-Centric Reasoning cites this paper.

Information-Regularized Attention for Visual-Centric Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.256397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T15:12:43.475802Z digest=sha256:4b20ac918a0016ca2d7160e7bef0214ecd64d86521d8c15c1bfa1aae1f2ca7d1

Observation 64fac733-b34b-4efc-8a60-fa541118694f · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers Vision language models are blind: Failing to translate detailed visual features into words

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:06.631736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:06.631736Z digest=sha256:eab7aa26775844dc4fd5d30acad0daabbc0b12b7e77e6de5116c98a69af2e189

Observation 7310bdba-2adb-4724-b15b-138f74838505 · inbound

Evidence-RL: Towards Evidence-intensive Visual Reasoning cites this paper.

Evidence-RL: Towards Evidence-intensive Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:05.848942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:39:05.848942Z digest=sha256:7746bdc7afdf4e100c392fdd1376f9e33a8091f901793791bfce1746272b85c2

Observation fa21b7df-5558-4e33-8163-a62a77acc62f · inbound

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cites this paper.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning Vision language models are blind: Failing to translate detailed visual features into words

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.194557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.194557Z digest=sha256:4185c83ebb34e328b7a3598aa10389f2719d01d4060adba0a222f4f5cf59272a