Pith. sign in

Paper Citation Record · LEDGER

Do Vision-Language Models Really Understand Visual Language?

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2410.00193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.00193 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:12.945758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:13:21.512791Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8bcc6c06-589e-4811-b439-600c4ccc9304 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Do Vision-Language Models Really Understand Visual Language?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:12.945758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:12.945758Z digest=sha256:997499ac0a13d0679499071bbf835ec8f28a1c2a7a72a2a85a6e85df06c3f3bb

Observation 3ef1a41b-ba4f-41a4-bf3f-c49dd05ac35d · inbound

WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis cites this paper.

WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis Do Vision-Language Models Really Understand Visual Language?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:00:32.189282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:00:32.189282Z digest=sha256:a0e5df13626b2cfd72763b7345f66449f80d35577d0dcded8b406017258e1d71

Observation aee70df0-c41f-468b-ad14-146a75c38881 · inbound

ViStruct: Simulating Expert-Like Reasoning Through Task Decomposition and Visual Attention Cues cites this paper.

ViStruct: Simulating Expert-Like Reasoning Through Task Decomposition and Visual Attention Cues Do Vision-Language Models Really Understand Visual Language?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:06.467280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:06.467280Z digest=sha256:d5f6e6f7acfde672cc4121296cb65c1d8868732ece9ee7a0565e71240be6e261

Observation a9ba543e-5102-49ad-bd2c-1d24ad5fe75d · inbound

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models cites this paper.

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models Do Vision-Language Models Really Understand Visual Language?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:27.125186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:27.125186Z digest=sha256:9ceebc48f5f80cb277674fb959aec134a9dcdb5b64536096f464d48c80f9b6eb

Observation 9f01fb84-53a7-4d57-8174-b08f4d95fdad · inbound

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features cites this paper.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Do Vision-Language Models Really Understand Visual Language?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.649412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.649412Z digest=sha256:125cbe349ea77d3eb2309b4fb81f4e12d77e5e4c46ab33ebe92a2896044110dc

Observation bb536c98-020c-4ac6-8c8b-7f2f360deeef · inbound

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions cites this paper.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Do Vision-Language Models Really Understand Visual Language?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.284722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:cca552d692053909eaa8a3e8e059e470d95b38b73036417dc69d0c746e7f71bf

Observation 26b2dbf2-6d71-4bae-9ccb-acd527940f04 · inbound

Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding cites this paper.

Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding Do Vision-Language Models Really Understand Visual Language?

Reference 26

Resolution
verified exact
orphan_title_repair, observed 2026-05-13T17:18:02.458267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:15:31.368264Z digest=sha256:36e86528ec1ea4253090e14314cb0a22e4a56bebdba3aa228e48a35b4b9aab54

Observation 51ea26a9-8f75-4e2d-bd1e-165529af6f56 · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Do Vision-Language Models Really Understand Visual Language?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.514602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:9807fcc533062f64f4156302acf2b7785b69db089ac42f3ae3c67c42af8af00f

Observation 8b54b4b5-1c0a-46a6-b413-a2b356e5ee53 · inbound

Prior Bias in Vision Language Models on UML Diagram Interpretation cites this paper.

Prior Bias in Vision Language Models on UML Diagram Interpretation Do Vision-Language Models Really Understand Visual Language?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T06:36:12.761601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T06:36:12.761601Z digest=sha256:5da6d7eadff551ac565062eb7a265bbf8acad76c297d10a0c635024d517d092a

Observation 8e712cc3-8472-41e2-b623-e29e0c4283d1 · inbound

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams cites this paper.

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams Do Vision-Language Models Really Understand Visual Language?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T07:19:14.207385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:19:14.207385Z digest=sha256:20eb8dc8a473f0c060926c76cd8d9a03dbe01e079a6b5c80f37b1cd76ec28c82