Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing

As of 5 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2509.04469.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04469 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:20:26.273267Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8c3042f-6252-44b7-8a55-71f5ec398a58 · outbound

This paper cites From Theory to Practice: Real-World Use Cases on Trustworthy LLM-Driven Process Modeling, Prediction and Automation.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing From Theory to Practice: Real-World Use Cases on Trustworthy LLM-Driven Process Modeling, Prediction and Automation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:20:28.336247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:23.118070Z digest=sha256:26b11207785433b94648dc7d864d54beed5c2c6198fc2135c9e37a42ed2382bc

Observation 5a4f03fa-640c-45c2-9896-a7fbdbc26a68 · outbound

This paper cites Towards automated auditing with machine learning,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Towards automated auditing with machine learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:31.555716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:23.226947Z digest=sha256:4182fc6cc5c05b75bf898f194889aec1a926e5fbac2007d2f8951db09dbda5dc

Observation c881b657-495f-4979-ab75-5612ad701b3c · outbound

This paper cites Advancing risk and quality assurance: A rag chatbot for improved regulatory compliance,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Advancing risk and quality assurance: A rag chatbot for improved regulatory compliance,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:31.274844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:23.386519Z digest=sha256:9485124d95f2fb6571e4761a7e95d70525747ca65495aaadf5cd00016b291254

Observation d3fd2493-b4f4-4333-bbf0-38a0d62f60d7 · outbound

This paper cites Fine-tuning large language models for compliance checks,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Fine-tuning large language models for compliance checks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:31.033404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:23.486011Z digest=sha256:e51d0681522fc5916948d1cc3840bc523f3b0ba5e331ac31efdc7a1ecb11f123

Observation bb6c00df-3c0b-4bb5-9eda-4d0ee44c3118 · outbound

This paper cites An overview of data extraction from invoices,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing An overview of data extraction from invoices,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.695448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:23.624840Z digest=sha256:3ff7b7bd10f951188ceec71491eb8906febba45c1f38f45e81da14037da1b003

Observation a5f28a53-32a9-4176-afea-91452a2ba6fd · outbound

This paper cites A Survey of Deep Learning Approaches for OCR and Document Understanding.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing A Survey of Deep Learning Approaches for OCR and Document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:23.707156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:23.707156Z digest=sha256:a525a62f5091867ec1f1fc7468eb67bd0b43dc51131619a0d20918d118969c2e

Observation a577dc6c-28e4-4526-bf59-61180c98fc14 · outbound

This paper cites Deep Learning based Visually Rich Document Content Understanding: A Survey.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Deep Learning based Visually Rich Document Content Understanding: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:23.881572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:23.881572Z digest=sha256:d251db2c6e6888e5bedc2d7157ce39728be4d1a326a62c9274e4fc58c0bb8bf6

Observation 3f2573d4-a449-446b-831a-a0bee46287a9 · outbound

This paper cites Memory-augmented agent training for business document understanding,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Memory-augmented agent training for business document understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.461236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:24.050023Z digest=sha256:cb0a58b4c9273e3dbe9ef8a60fc21fa7a82b85776916c7c5f626c996a80c8208

Observation a025c8de-59a3-43eb-bcae-a02bba2427cf · outbound

This paper cites Wonderbread: A benchmark for evaluating multimodal foundation models on business process management tasks,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Wonderbread: A benchmark for evaluating multimodal foundation models on business process management tasks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.301934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:24.224802Z digest=sha256:85cee586a7b890b7a105f4e61f6c4903a314e5739f05e1a981b2a4a54e29a136

Observation 8f82b2bb-18c4-4e96-9a2a-b11a4d3fad16 · outbound

This paper cites MMLONGBENCH-DOC: Benchmarking long-context document understanding with visualizations,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing MMLONGBENCH-DOC: Benchmarking long-context document understanding with visualizations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:30.148048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:24.436236Z digest=sha256:86ad9eec81f3fd49f524074ec5eb9679c0f9e085502968454f29d0d9adad641e

Observation 5b09afa9-9cd9-4903-9a5a-e813cffb0543 · outbound

This paper cites M-longdoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing M-longdoc: A benchmark for multimodal super- long document understanding and a retrieval-aware tuning framework,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.998179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:24.581017Z digest=sha256:202f76eacbfd8a37f8bec9d17d9bd8effbdbe647ee8b6ab25bd4f2a3f5cd104f

Observation a7068012-6c16-455a-ba9a-a227db41d7bb · outbound

This paper cites DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.762308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.762308Z digest=sha256:c362b748021702b699533a418d57971083f818fe77e1c10b53af0d90191475a1

Observation 463c174f-422f-4bf7-8a24-ff1346cb8e08 · outbound

This paper cites Gemma 3 Technical Report.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Gemma 3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.883307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.883307Z digest=sha256:7db3e36f53096553133478b64e6faf072dfa5a2f5abc18ca4faf62d9d2b3d2b6

Observation 63cf2eeb-754b-40e3-829e-c38033740bdd · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.672980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.672980Z digest=sha256:d71411901eb27826fb0043aaa011a859656287fbfb57902bf4f9c369aef1a997

Observation 0b8afe44-8a1a-46f0-bb21-6192ce1ac89a · outbound

This paper cites SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.170342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.170342Z digest=sha256:216d590d10b35acddf6b97d0038989da6a8e7c21aa583b6be68f85efd8485fb0

Observation 63da7747-6544-42c2-b6f3-28763bbb382e · outbound

This paper cites Pdf data extraction benchmark 2025: Comparing docling, unstructured, and llamaparse for document processing pipelines,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Pdf data extraction benchmark 2025: Comparing docling, unstructured, and llamaparse for document processing pipelines,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.714799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:25.310390Z digest=sha256:8e1a94ef45dc99ed64b699c8441bed9dcf09f98d44774183c1a89a2354e0def8

Observation d89c6c31-da03-41d0-9765-b874d46061af · outbound

This paper cites Docling Technical Report.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Docling Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.043109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.043109Z digest=sha256:48be4c64cfd38f9465e0e02014173a9e30afc43dd83fcaf6ae6204e96727bcd9

Observation 28516716-8c73-4d53-8e84-be9d4e3feef8 · outbound

This paper cites Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.494652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:25.574862Z digest=sha256:d81cf1b3e245eab6861bdafca2cda3e5bbddbaf5afa23f955679a6e7219a262c

Observation d495aad1-eb43-4047-83c5-ac40093a04a4 · outbound

This paper cites Field extraction from forms with unlabeled data,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Field extraction from forms with unlabeled data,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.253485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:25.713563Z digest=sha256:5988d57ecd99596bee3da761242d6bb11fea936cb23d15e6130af4a4d3ba72f2

Observation 420f57b5-134c-4055-91b6-f8da0a8e243e · outbound

This paper cites Ocr-free document understanding transformer,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Ocr-free document understanding transformer,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.424347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.424347Z digest=sha256:0a8f919dc2c480f2d593dfe5d00da8c813dc8d4edde3ff00c45eb00977807d05

Observation 8f899c6e-b581-4a55-bb0e-b5dc7e5d3b86 · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Layoutlm: Pre-training of text and layout for document image understanding,

Reference 21

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T14:20:27.117317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:25.924804Z digest=sha256:22593f9f1451f2610e928734af29d8413c63ed505b59c9e634e6f8f0f47bda5e

Observation cd29b14b-566c-4a13-8616-c2885530940f · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Layoutlmv3: Pre-training for document ai with unified text and image masking,

Reference 22

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T14:20:26.734768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:26.035589Z digest=sha256:0de334cd6af2aa85ba5a972724dc74bd187c765f1b9e2bb84faab7cf78056790

Observation 8f6fbf9d-d63a-42fb-8116-d30b7f2a2fe9 · outbound

This paper cites Truth tobacco industry documents (formerly legacy tobacco documents library),.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Truth tobacco industry documents (formerly legacy tobacco documents library),

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:29.089939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:25.840205Z digest=sha256:4daa243fdc7086ad7e5af3b8e72bd3f8fb53560e71965c6cbf89e51982ccede7

Observation 1453c3ad-843d-464a-9243-9ea03ac71c51 · outbound

This paper cites Lilt: A simple yet effective language- independent layout transformer for structured document understanding,.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Lilt: A simple yet effective language- independent layout transformer for structured document understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:20:28.848561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:26.134753Z digest=sha256:d7b9050f07825955faa84ee115d4fc751d66e8906e79e7bab042f5fe6e99af67

Observation 4c7db77f-c7e1-43c0-9a40-0aa40ab351d7 · outbound

This paper cites Available: https://doi.org/10.1145/3342558.3345421.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Available: https://doi.org/10.1145/3342558.3345421

Reference 2019

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T14:20:28.028635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:23.281575Z digest=sha256:c55c1f027bf40f30162d57ecb865ff9a7ab593102fd403b08d59df451390332d

Observation 9fb6ba9c-b98d-43be-a961-0e4f1390640f · outbound

This paper cites an unresolved cited work.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:20:28.634779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T14:20:26.273267Z digest=sha256:27008d26bbfcba8ca1032b3aa3b64efdf201385e8dc416093f0b16fd7261b8e8

Observation 21ee41fd-fcae-48f7-a742-f5ac024f7e33 · outbound

This paper cites WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:24.320091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:24.320091Z digest=sha256:ea7351cd7cc268a04bd74a82c88722b9f02ba61b7bfd67395e6c93b252cb8f40

Pith citing papers

No inbound Pith citation observations are available.