Pith. sign in

Paper Citation Record · LEDGER

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

As of 4 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 3 inbound Pith citation observations for arXiv:2604.02486.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.02486 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T21:56:29.924506Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T11:43:17.464276Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:29:41.276218Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4295210a-ea64-4634-a8f4-458c4a8f8c81 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T21:58:19.873755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:b4f02f9b62ab0cba46622ebaa0336fc40983397d2c4396e5744bf360be344988

Observation 3be3dc69-02b2-43c8-ba62-ad5c117f3a99 · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:618831827373c12cbe71981bba1c901cc8963b9076896276fdf0a1fe9545ae90

Observation 97016c81-9cef-4b73-ac5d-a0ee5e861fac · outbound

This paper cites ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:49:50.752568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:ce39bd1c78ba93094575f874a7bed6446dbe1810b94917c6df1063db8822c254

Observation 4825b433-659f-4773-8bee-187497140f64 · outbound

This paper cites Hidden in plain sight: VLMs overlook their visual representations.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Hidden in plain sight: VLMs overlook their visual representations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:58:19.843334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:40efb64463380d0bb0c68bf58ab30f1353a3b9eb82759855a3e77c39e34a779b

Observation 6426b7a2-73a1-413f-8eaf-4ec3f951840a · outbound

This paper cites Accessed: 2026-03-13.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Accessed: 2026-03-13

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.835220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:63f7dc63e31fdec3d39dbaad748fd06e6cee5e79a8581ac618a8d964a4d3c317

Observation f1203dc3-cdbf-4939-9a65-b98ca2e06a09 · outbound

This paper cites Pisapia, Kenji Ikemura, Mert R.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Pisapia, Kenji Ikemura, Mert R

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.840582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:8e2efd81f06e06301b324c38af6c2eca2eb05d87fe0e33d657e6746b0d5dd5ca

Observation 653bfb2b-8800-4429-9574-a79e0b574664 · outbound

This paper cites ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:58:19.868236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:6a1fc155232928c1fdeb5a3492a2a1b6ff4598eaa3ff3d2bc09b7b3a59b9b3b6

Observation 9ea23948-aff0-4e0e-84a5-0843887fd71d · outbound

This paper cites LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-12T02:08:25.203802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:522233d1e7223e5fa0fb46830b9f2a062de50b2c8bf689c020acb4fea65e8393

Observation 49479763-0788-4b1d-8280-03678206a840 · outbound

This paper cites Visual representations inside the language model.arXiv preprint arXiv:2510.04819.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Visual representations inside the language model.arXiv preprint arXiv:2510.04819

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.859370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:213a332980fa48d10c7e0635fca6a16e5581fd0061417472067430c2f52cbadd

Observation f41b1c30-9f69-4cde-8704-ce62ef659d5d · outbound

This paper cites Linearly Mapping from Image to Text Space.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Linearly Mapping from Image to Text Space

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.846073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:613da0219d034227e5abaea9c04e9bbcbc332c4c2bdad8f50b12de34062dd11d

Observation 7b4c1eec-7a33-4f31-9091-beeecad501a6 · outbound

This paper cites SPair-71k: A Large-scale Benchmark for Semantic Correspondence.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors SPair-71k: A Large-scale Benchmark for Semantic Correspondence

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.856641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:bee4108b0be70d0cc0cc49d0e3e84b6ec3033548c9ee1f0e49444c17be208428

Observation 371aaf5a-e7ce-4e94-a9bb-1f0b438d264e · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.838229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:eaa54f0426232f433e08e8b40257fabbf572be423397a1b5f93b609d3aa9016e

Observation 09d240dc-382a-4b3a-960b-7816365a41d2 · outbound

This paper cites Same task, different circuits: Disentangling modality-specific mechanisms in vlms.arXiv preprint arXiv:2506.09047, 2025a.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Same task, different circuits: Disentangling modality-specific mechanisms in vlms.arXiv preprint arXiv:2506.09047, 2025a

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.871151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:1a589b3b83e0f09d2052310aac94d8bb3d849ad1c30920161571908f04a3e6bd

Observation 18e68879-6792-4449-8486-0ba08abb84e7 · outbound

This paper cites GPT-4 Technical Report.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors GPT-4 Technical Report

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T21:58:19.865213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:8dd0aac73a58a4643cee114856d614639807ac7ebe5dd8efeb644d78df929e6b

Observation c470afb2-a3ab-46ec-9af8-360ac8356a35 · outbound

This paper cites IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:58:19.876482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:636e1fdde89e4cc1ca90cfd01beafba736afa8047a946daa19d1eab2fd7fa9dc

Observation 7b7107e0-08f7-4343-a956-fe77e65f4de8 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:29:30.364997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:61598e1c920fe5b8d20764124ef20717700937bb66cecff04adf5ef2eb438828

Pith citing papers

Observation cecb2cf8-1aec-4423-a2b5-8ace25ee3c44 · inbound

3D-Anchored Lookahead Planning for Persistent Robotic Scene Memory via World-Model-Based MCTS cites this paper.

3D-Anchored Lookahead Planning for Persistent Robotic Scene Memory via World-Model-Based MCTS VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:11:06.924970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:05:39.801857Z digest=sha256:d4c9d542e9c6841566089e802164f84d215a49542172d427d1e9077c32ccaf77

Observation 7decb56d-ba39-4c72-82bc-4103d31c0b67 · inbound

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models cites this paper.

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:35:26.781973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T13:25:27.762910Z digest=sha256:fb03c0a2a072ba26dd973848a664bc99bbfa30de63db8c49514940052e9bed72

Observation c91128ae-5053-4e51-a1b2-d4b02d2bc15e · inbound

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR cites this paper.

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:29:41.277734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T11:43:17.464276Z digest=sha256:bd36dd99d33170739b78095c4067427a7c61008b5360b5b60dce29716c17cedd