Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:13.435547Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2502.04389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:13.435547Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 69974e3e-79c9-463f-91b3-905e1340c4a1 · outbound
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259ff82d-5221-4dc5-95be-99ba4072062f · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f04d3b80-dc59-46f7-b6fa-5b3e6379b48d · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions PaLM 2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f8e264-b9a8-46d0-8995-dd93f79f24a2 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Safeguarding Decentralized Social Media: LLM Agents for Automating Community Rule Compliance
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 373299ba-c648-48fb-b2d2-8624e27acd24 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions How far are we to GPT - 4V ? Closing the gap to commercial multimodal models with open-source suites
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a539b88b-8b20-47aa-b76c-801293357724 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented Generation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c39ac3-b6c8-47a8-8297-8a833e4bf144 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bcc6c06-589e-4811-b439-600c4ccc9304 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Do Vision-Language Models Really Understand Visual Language?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4695d3-f338-45ad-9755-b8c3ef070c0c · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Introducing Gemini 2.0: our new AI model for the agentic era, December 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 576538dc-7b39-4c28-867a-49080da2e95a · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Layoutlmv3: Pre-training for document ai with unified text and image masking
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c8a2effa-a012-434d-9ae5-c944cac9172d · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions ISO / IEC 29500-1:2016, 2016
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99452ea1-9ee4-4289-9077-fe367dc07d56 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f7b4706-7df6-47ec-8b9a-afe353b5659c · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions ( Security ) Assertions by Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d15e50ef-e078-4125-a7c7-98b07b7d8955 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d002fb36-271c-4e91-b220-b6bdc1313401 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Ferret-ui 2: Mastering universal user interface understanding across platforms, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5343c688-3fed-4cae-a254-edb358d51da1 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46a944f-5cdb-4f8a-b1f6-0d92fa01b048 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d38526-95cd-4a30-8825-f859a2f9c39e · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Unraveling the Truth : Do VLMs really Understand Charts ? A Deep Dive into Consistency and Robustness
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9a610d-cec7-4780-a594-25a95d3e5b22 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions GPT-4 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 708da264-625b-477a-b10e-33fe10f5a56d · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions GPT-4o System Card
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d61a7b8-b479-4174-a363-d73d9d86409d · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85364f4f-4887-461d-a6f9-59f4f65d3781 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Vision language models are blind: Failing to translate detailed visual features into words
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18df3a35-225f-4cf7-8ccb-bd1af4aa9f08 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions FlowVQA : Mapping Multimodal Logic in Visual Question Answering with Flowcharts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eca1d1fa-eee1-425b-afd0-a23070db800e · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions LLM4VV: Exploring LLM-as-a-Judge for Validation and Verification Testsuites
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 631d5d0f-34fc-46de-8529-24b46f3513f4 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2f8e72e-f6ad-474f-87c8-239ff489b7a6 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Feighelstein, Jasmina Bogojeska, Joseph Shtok, Assaf Arbelle, Peter W
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f6c368e-81c1-41ef-b40a-3f78581f2390 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions LaMDA: Language Models for Dialog Applications
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14e2e0e-8f85-40ef-a7cf-0fef0c888f50 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7206a5bf-b538-4c66-aee6-8b6a5350f007 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8e8ba3-4849-42b3-817a-bc022ce0aa3b · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f27af81-5143-40f1-bea6-3a9c2504b3b9 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions MMMU : A Massive Multi -discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bfe125d8-adb9-4275-87c1-bfc3a8d37400 · outbound
Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.