Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T21:56:29.924506Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 3 inbound Pith citation observations for arXiv:2604.02486.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T21:56:29.924506Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T11:43:17.464276Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T08:29:41.276218Z
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4295210a-ea64-4634-a8f4-458c4a8f8c81 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3be3dc69-02b2-43c8-ba62-ad5c117f3a99 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors BabyVision: Visual Reasoning Beyond Language
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 97016c81-9cef-4b73-ac5d-a0ee5e861fac · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4825b433-659f-4773-8bee-187497140f64 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Hidden in plain sight: VLMs overlook their visual representations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6426b7a2-73a1-413f-8eaf-4ec3f951840a · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Accessed: 2026-03-13
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f1203dc3-cdbf-4939-9a65-b98ca2e06a09 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Pisapia, Kenji Ikemura, Mert R
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 653bfb2b-8800-4429-9574-a79e0b574664 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9ea23948-aff0-4e0e-84a5-0843887fd71d · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 49479763-0788-4b1d-8280-03678206a840 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Visual representations inside the language model.arXiv preprint arXiv:2510.04819
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f41b1c30-9f69-4cde-8704-ce62ef659d5d · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Linearly Mapping from Image to Text Space
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b4c1eec-7a33-4f31-9091-beeecad501a6 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors SPair-71k: A Large-scale Benchmark for Semantic Correspondence
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 371aaf5a-e7ce-4e94-a9bb-1f0b438d264e · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 09d240dc-382a-4b3a-960b-7816365a41d2 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Same task, different circuits: Disentangling modality-specific mechanisms in vlms.arXiv preprint arXiv:2506.09047, 2025a
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 18e68879-6792-4449-8486-0ba08abb84e7 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c470afb2-a3ab-46ec-9af8-360ac8356a35 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b7107e0-08f7-4343-a956-fe77e65f4de8 · outbound
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cecb2cf8-1aec-4423-a2b5-8ace25ee3c44 · inbound
3D-Anchored Lookahead Planning for Persistent Robotic Scene Memory via World-Model-Based MCTS VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7decb56d-ba39-4c72-82bc-4103d31c0b67 · inbound
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c91128ae-5053-4e51-a1b2-d4b02d2bc15e · inbound
When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.