Pith. sign in

Paper Citation Record · LEDGER

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding

As of 21 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2412.16420.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16420 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:38:13.638865Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:09:13.307268Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:48:52.562583Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad7f72ba-a4e4-402d-963a-1a5b83cae79d · outbound

This paper cites Llama 3.2: Vision at the edge with mobile devices.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Llama 3.2: Vision at the edge with mobile devices

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:14.056536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.543664Z digest=sha256:fad1a3661c09e2968c743d728f6a90b26e11ec1b6ed854f91fedec8f37e7f845

Observation 368ab642-513a-48f9-bb3a-cbf5c55546f7 · outbound

This paper cites Introducing claude 3.5 sonnet, 2024.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Introducing claude 3.5 sonnet, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:14.042050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.548798Z digest=sha256:addc0c0f849096acddd49ad7a29679afa0b2de5bb42dc6ea5bc79e5809c97dfb

Observation d49b3753-18a3-41d5-b6b8-654fbb811489 · outbound

This paper cites Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T10:38:13.866245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.553624Z digest=sha256:cc192183a66a83837d13aaee25d0c7ba6c48f42e5f4cf8ba483b895dbcb901f2

Observation d6661da2-246c-42f8-99c9-3ce0323167fa · outbound

This paper cites Modularized zero-shot VQA with pre-trained models.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Modularized zero-shot VQA with pre-trained models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:14.027684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.558722Z digest=sha256:61a65d970a0bb209b1799a72ea1420bd853b7ae8240008914d23bfd5f3273426

Observation 4277e715-9ce3-452a-a7d9-76f440b5a08b · outbound

This paper cites The Llama 3 Herd of Models.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.565043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.565043Z digest=sha256:2220f7b23d87ecaa381bb8e185c4a4a95333edf48ade24a1c4eef6a35d733c90

Observation c3a49f58-010b-4fb4-b064-985d72e79c62 · outbound

This paper cites Exploring question decomposition for zero-shot vqa.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Exploring question decomposition for zero-shot vqa

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:14.012757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.569989Z digest=sha256:96d07e15ed8eb2e3112bf273c5c3a07d420bfcedaa31a49e36ccdaac9ecac85b

Observation 81ed77fe-6323-4081-b5d6-73f8fd9526b3 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.574892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.574892Z digest=sha256:c0fcaaba7354cb84679e41a8e2b291c54cc067962bfcce941d0d27a884b300fe

Observation 073f6a9d-1df5-41ea-9497-72e3531543f2 · outbound

This paper cites Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:13.985726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.579169Z digest=sha256:b78777e4f09de648c916d591e1c267403b9a7f5c2141cbd8fdc9c7ec8ada2b2a

Observation fc5d96f9-82bb-445d-9f75-c081e179efc9 · outbound

This paper cites Hello gpt-4o, 2024.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Hello gpt-4o, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:13.968824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.583556Z digest=sha256:16cee6ba7a3ce01e0cd5e15205592aef4f0bccd4a4e98af05a328aede703b967

Observation a76b5c08-c337-4da2-ac1b-e6b8b9454b59 · outbound

This paper cites FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.588334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.588334Z digest=sha256:f63d5897570dcd9e0d56b9178c4b64eaaad89278b4322b149c535f3e8b6715ec

Observation 00c1fba5-9401-4404-87b8-82a9cad67a99 · outbound

This paper cites F low VQA : Mapping multimodal logic in visual question answering with flowcharts.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding F low VQA : Mapping multimodal logic in visual question answering with flowcharts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.593301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.593301Z digest=sha256:65027b0541767350208a46e675b0b148ddb265fb068140dc6c354b49325c1ebd

Observation ae5b42dd-aef2-4c36-b9a6-ee1d4ea67ba7 · outbound

This paper cites Feighelstein, Jasmina Bogojeska, Joseph Shtok, Assaf Arbelle, Peter W.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Feighelstein, Jasmina Bogojeska, Joseph Shtok, Assaf Arbelle, Peter W

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:13.953678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.598100Z digest=sha256:707af7932141ad15f03dfe84ced8748fc589cc9f2ed9efb6c823b440d066a2ea

Observation 5e6a8077-0786-4f75-be2a-47188452bfbd · outbound

This paper cites Mixtral 8x22b: Cheaper, better, faster, stronger, 2024 a.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Mixtral 8x22b: Cheaper, better, faster, stronger, 2024 a

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:13.936915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.602902Z digest=sha256:a864f4b363f30350691eed7b80ae24e893a524de4314f517f4b23c35d0d953fe

Observation d305a17b-54e1-40fb-80e4-d8175ffdf348 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024 b.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Qwen2.5: A party of foundation models, September 2024 b

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.607676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.607676Z digest=sha256:293110213bea8fba33f1a513e2a6aeab9564e818cdf09d489a323274231a1a29

Observation 0a187496-ae91-41a6-b584-7d68cb25a39b · outbound

This paper cites Discover the new multi-lingual, high-quality phi-3.5 slms, 2024.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Discover the new multi-lingual, high-quality phi-3.5 slms, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.612034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.612034Z digest=sha256:622ace6688c52734d9f6650201d38afc027289298488ec87aa0fa14468c6bafb

Observation 0618911a-367a-4201-af64-81e52c18f49e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.616000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.616000Z digest=sha256:a3a235dd9dc94924ef3dec179c7882fa9d245b5a6491d8a6a3d7ee2403a82293

Observation 3ef79cf3-4abd-4921-9ae1-1c4160a53d39 · outbound

This paper cites Ayyubi, Kai-Wei Chang, and Shih-Fu Chang.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Ayyubi, Kai-Wei Chang, and Shih-Fu Chang

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:38:13.910340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:38:13.620521Z digest=sha256:8cdf9f36f747bb45c960094890378c183ea63cdb205192fc13128f10c76692f9

Observation 6b0d8317-1199-4acc-a88d-4391c16bc1ac · outbound

This paper cites write newline.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding write newline

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.624309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.624309Z digest=sha256:4ddbc88f36c2e7731b871bd3347318212eee4a783e9138d1c02097978ba95efa

Observation ba12dd5b-5a3a-46b5-8ad4-9f57605b181e · outbound

This paper cites @esa (Ref.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding @esa (Ref

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.629742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.629742Z digest=sha256:dff7286afa042bbdc46c78285cbc06c6eb4953a21b0f4f89f814e830149a2031

Observation 7b5ec85b-01e4-47b7-a97c-07ab6d2ca185 · outbound

This paper cites an unresolved cited work.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.634478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.634478Z digest=sha256:2cbc104c7dc40056f461814bbc875b9d88a9a0e9d28778075461c96864e23e4e

Observation 175cdee9-cb93-4670-b43c-429b2f305391 · outbound

This paper cites blue The master students are still working on the annotation and will get the results for those empty entries by today.

Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding blue The master students are still working on the annotation and will get the results for those empty entries by today

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:38:13.638865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:38:13.638865Z digest=sha256:e497d8a255cd58614a4be339d255835258f4d271880252847cb080f06a048db8

Pith citing papers

Observation 7206a5bf-b538-4c66-aee6-8b6a5350f007 · inbound

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions cites this paper.

Overcoming Vision Language Model Challenges in Diagram Understanding: A Proof-of-Concept with XML-Driven Large Language Models Solutions Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:09:13.307268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:09:13.307268Z digest=sha256:2f63e8bf32da971a626641bb1c57b0ca7994fd778f9786f2f096f260b793b4d8

Observation 0320b4fa-3aa1-4bd2-b947-5371c8880304 · inbound

Survey of GenAI for Automotive Software Development: From Requirements to Executable Code cites this paper.

Survey of GenAI for Automotive Software Development: From Requirements to Executable Code Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:52.671511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:48:49.568986Z digest=sha256:1bd3d201ed6784560db3a9b5bb4b344c2afe43d1a50cb687a3bd48f80c4326de