Pith. sign in

Paper Citation Record · LEDGER

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

As of 5 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2603.15118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.15118 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T10:34:21.033424Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact7
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ad9980c-87cf-4bda-ba42-3e7578c0f03b · outbound

This paper cites Qwen3-VL technical report.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Qwen3-VL technical report

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.159707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:dca60fdcdcf3960d80c4931289fa82bd0019590407acdf1102de76ca84751b68

Observation cffe2b4d-f93e-4825-8c9e-692f2993f27d · outbound

This paper cites So-bench: A structural output evaluation of multimodal llms.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents So-bench: A structural output evaluation of multimodal llms

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.445604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:1f2bc76edc5b92decf3afd6792f538434934de2e6742a17e8153f8db43461ed7

Observation 0cdfcef6-42dd-482b-9edd-bed48ef1c315 · outbound

This paper cites ExtractBench: A benchmark and evalu- ation methodology for complex structured extraction.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents ExtractBench: A benchmark and evalu- ation methodology for complex structured extraction

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.420029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:da26d2f1443fdda4eebf0380b6a392796d0def034664f076ebe93727d4345678

Observation 179f8895-3953-49a7-a276-a064a88d6c40 · outbound

This paper cites Gemma 3 Technical Report.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Gemma 3 Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.425028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:46d4c683b7d6e0739546338edcbb343cd3bba957daf0e46c1925f02633778ac2

Observation a5a0d67a-5b89-4e0c-abd2-e28059c898bf · outbound

This paper cites JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.414369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:dbd9c2f0e970be3cce062121224aea5df592b3dea7dddd09c3f87977d16510ab

Observation e982eef9-5206-4a15-a018-300faaa05ca3 · outbound

This paper cites Gemini 2.5: A new family of highly capable multimodal models.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Gemini 2.5: A new family of highly capable multimodal models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.161620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:0b832a2e41a7446a02df31bbdc7cfdabd46ce153e5d38e5c9b53eb0ab62ab11b

Observation 92a79c60-6e8b-4068-9333-fe815981ac04 · outbound

This paper cites H2O-VL mississippi.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents H2O-VL mississippi

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.165359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:dbe6b24ed26372852444a659c1f5ec54a359d6874c151a0587ba991941fefc82

Observation b2fa3e21-8463-4984-a86b-e5f4d1aa9303 · outbound

This paper cites LayoutLMv3: Pre-training for document AI with uni- fied text and image masking.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents LayoutLMv3: Pre-training for document AI with uni- fied text and image masking

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.135275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:dcf151a2fc533ab8335ccbbb9ad2e5815a543f377db2fcd3bbf354a90825a021

Observation 4acb5945-8059-4545-81ec-e9fc985a8975 · outbound

This paper cites an unresolved cited work.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-15T10:35:28.137104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:2b89ff3a646fcf4c71a5638e309d2bcf972f15cbf73369325a14875e2bcdc56a

Observation fa70b0cc-aba8-4594-bacf-d2667433626f · outbound

This paper cites GPT-4o system card.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents GPT-4o system card

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.154084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:9870c628b2c290521716cc96b2b4da4c508ef02f7f4032efb5d23272f1af11df

Observation 9e587e8a-819c-469d-96ad-4431396e2a69 · outbound

This paper cites FUNSD: A dataset for form understanding in noisy scanned documents.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents FUNSD: A dataset for form understanding in noisy scanned documents

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.133240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:19768dd0dc9d8be00907318e3bcb7e246dbbc5ba8d6eba5e7626b84fd2d0928c

Observation 21c99078-50b3-4866-8ffe-803decac8f3d · outbound

This paper cites Donut: Doc- ument understanding transformer without OCR.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Donut: Doc- ument understanding transformer without OCR

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.155921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:71cfe58c55f24aec76cbd96f8a644772a67a5a8d7f504471e778f01624c95291

Observation c785d5db-33cd-4c4b-9c13-670dfc349b0f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Gonzalez, Hao Zhang, and Ion Stoica

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.150359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:ecf78f6ce077ddbcd102325ab6f10a1944b15c2e4605b423d539fad51bfa11a8

Observation 336fa79f-c0e9-4990-b115-c7019b1a003a · outbound

This paper cites Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.430189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:8931ba27e57f26a9a676af73c165592a601d6119638002e23a63427b16fef6e2

Observation 64fb9e01-59a4-43de-a415-df65569f04ae · outbound

This paper cites LayTextLLM: A textual layout percep- tion model for visually-rich document understanding.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents LayTextLLM: A textual layout percep- tion model for visually-rich document understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.152322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:57b916d74bb40305597530823f9286544c1505bbbceaefd5f4a0186ebbd2bebd

Observation 3c948b04-e2d4-4d60-a074-c3772d352e1e · outbound

This paper cites Llama 4: Maverick and scout.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Llama 4: Maverick and scout

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.146657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:797f9534e8043d4e964510ccaf5b7bddd4cea6f9461d85bb48aea83db39026cb

Observation 9084fece-81a6-44bd-862b-f46124df0404 · outbound

This paper cites Mistral small and ministral.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Mistral small and ministral

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.140979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:0e34b94d69489b8d385ea57ef7d0831b08520f8420baebce10f73e1de4173424

Observation c5a2ad45-e71f-46aa-b5e5-114bfb83ae77 · outbound

This paper cites NuExtract 2.0: A specialized model for structured extraction.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents NuExtract 2.0: A specialized model for structured extraction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.169141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:7d496e7a8eb2b4b8bb280c1e7b33ed1d11d8b061893282bc0c6cb29c70234630

Observation 9846f4ba-040a-4114-b615-72742a40ca4f · outbound

This paper cites OmniDocBench: Benchmarking diverse PDF document parsing with comprehensive annotations.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents OmniDocBench: Benchmarking diverse PDF document parsing with comprehensive annotations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.171149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:63cf2087d1b7ca6c137bdb6eff95c39e0d8fec474048df8dbb3428dd7385292f

Observation a049f000-fa77-45a7-9e31-ee1b143227e4 · outbound

This paper cites CORD: A con- solidated receipt dataset for post-OCR parsing.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents CORD: A con- solidated receipt dataset for post-OCR parsing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.148589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:7a9d44ade2b3b385da9627fed34bc013bc9cc5b926e4b2bbf3fd33c970900998

Observation 10e21ea5-b0b6-41c7-a66d-9241691aff83 · outbound

This paper cites LLMWhisperer: Layout-preserving text extraction for LLMs.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents LLMWhisperer: Layout-preserving text extraction for LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.139170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:0f2027bf39147e6a92833a6ea5c8b2a59ae4a77266d1703819c3ffca08f42dd2

Observation 9fdf5a07-6aa2-4a42-aed6-ff4baf79d8c7 · outbound

This paper cites DocILE benchmark for document information localization and extraction.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents DocILE benchmark for document information localization and extraction

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.142868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:3db61545a40e14c0a7a24c6aa96f76328e51d6ab271a23dd036c6a5118486b0d

Observation f918b11e-c8c0-4f0e-80e9-0fbb584862ee · outbound

This paper cites LiLT: A sim- ple yet effective language-independent layout transformer for structured document understanding.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents LiLT: A sim- ple yet effective language-independent layout transformer for structured document understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.144784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:64b20cbf9a323c1ddd7e51f6dca77a54c3951894685b9d06944f4d44527c9183

Observation af81ae58-b04f-4684-8bb3-55701c95be5b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.440400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:aa0979af4fe3b71c4e6044abe7e79b1cf639357d66d3fe113bde855f87ba67a7

Observation de4c98f8-cac6-4f19-b6b3-89b89b492be9 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.435589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:49bcda1b69ae880d2c3b268f25a2e730966b303ccc04a4a344de13ab19876938

Observation 8f738e47-96cf-4f61-8bd9-51e66e39d1c9 · outbound

This paper cites VRDU: A benchmark for visually-rich document under- standing.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents VRDU: A benchmark for visually-rich document under- standing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.157879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:704701ebfb26b1a731ef397e583ede18b62ede626ea082e91e80a52e228a9f0b

Observation 9e9ee6e2-3e54-4854-9b85-8ec9c0d512d3 · outbound

This paper cites LayoutLM: Pre-training of text and layout for document image understanding.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents LayoutLM: Pre-training of text and layout for document image understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.173016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:9f52aef3bb5c16c6e0d070922061dd3395222af2ddecc9ed6caa9bf451dd1847

Observation 57bb642c-a190-4e37-bd04-cd7cb34168a0 · outbound

This paper cites LayoutLMv2: Multi-modal pre- training for visually-rich document understanding.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents LayoutLMv2: Multi-modal pre- training for visually-rich document understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.167323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:bf9a1cf81eeb0742dea905660328ae2b796bf199eb02355cbe45723269ddadf6

Observation 9c39ca75-f44c-4356-97b6-05ae9791d3f1 · outbound

This paper cites type": "object.

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents type": "object

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.163541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T10:34:21.033424Z digest=sha256:996a559e6e44ab5a76aa84ffdcd6edda16a5091a6f491954332eb813cd975f74

Pith citing papers

No inbound Pith citation observations are available.