Pith. sign in

Paper Citation Record · LEDGER

DocVLM: Make Your VLM an Efficient Reader

As of 18 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2412.08746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08746 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:42:14.915935Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T21:55:02.656863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T21:59:06.400938Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d95e635-eb90-46bc-8478-43366a9a697e · outbound

This paper cites Sequence-to-sequence contrastive learning for text recogni- tion.

DocVLM: Make Your VLM an Efficient Reader Sequence-to-sequence contrastive learning for text recogni- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.847032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.660000Z digest=sha256:2fc51a7899a47e73f3d6a669d94b0564e3ad63e18ad18559d80b106ed7a01d31

Observation ce335f3e-c1c6-4b33-a514-a18942a47dd9 · outbound

This paper cites Multimodal Semi-Supervised Learning for Text Recognition.

DocVLM: Make Your VLM an Efficient Reader Multimodal Semi-Supervised Learning for Text Recognition

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:42:15.286747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.664037Z digest=sha256:e306b937ebdb6215a76ce2b55bb858657777c3bff09414b98e2721a34d9d9735

Observation a83da128-c292-4c5a-82f0-a0123d3c23a5 · outbound

This paper cites Clipter: Looking at the bigger picture in scene text recogni- tion.

DocVLM: Make Your VLM an Efficient Reader Clipter: Looking at the bigger picture in scene text recogni- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.832160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.667866Z digest=sha256:8367da62882adb17242a6dec653fd61365d15ca012775c0fb67f16a6af72b69d

Observation 97adfa2a-0880-4d59-84bf-f52aa06e0634 · outbound

This paper cites Visfocus: Prompt- guided vision encoders for ocr-free dense document under- standing.

DocVLM: Make Your VLM an Efficient Reader Visfocus: Prompt- guided vision encoders for ocr-free dense document under- standing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.819896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.671218Z digest=sha256:996161fb7bd26e1b10e761b77370d4d675951f1bce18cc85d12ecb060818a00b

Observation 0dbb5cf3-6f5c-4557-bb3e-4aa26e03ea85 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

DocVLM: Make Your VLM an Efficient Reader Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.674632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.674632Z digest=sha256:7647a1682886b3c8d9597b11fb0ef2ae60233ebcb93a2fa9143a53fa562caaa3

Observation abc368ed-db23-432e-b4c8-6a799a689b42 · outbound

This paper cites Docformer: End-to-end transformer for document understanding.

DocVLM: Make Your VLM an Efficient Reader Docformer: End-to-end transformer for document understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.801855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.678011Z digest=sha256:134556cdd7cd57e400321b0d3d1104a11c9afedb950392d5066a0d08feef558b

Observation 33997b18-557c-491d-bfe4-53f9caf214f8 · outbound

This paper cites Docformerv2: Local features for document understanding.

DocVLM: Make Your VLM an Efficient Reader Docformerv2: Local features for document understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.777421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.681409Z digest=sha256:c23fc9ab58cf281779cb9dc986a46bfb416f404e7aca11ac5f1d8b2b70412707

Observation 603ac7e8-1290-4fd2-acee-b26cfdd395dd · outbound

This paper cites ScreenAI: A Vision-Language Model for UI and Infographics Understanding.

DocVLM: Make Your VLM an Efficient Reader ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.684571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.684571Z digest=sha256:4377afb8b96d707e8f0fd6168a4f18c199c1b9d6a42d45da02112b01f4581d14

Observation 86291faa-31e1-4997-9aed-7da95e13ebe8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DocVLM: Make Your VLM an Efficient Reader Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.688335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.688335Z digest=sha256:a568c4aa7013e8f203dd6abc7e70b4306742992141d1af76e29cb0e1d4becf9a

Observation b4053255-729c-4b5c-9baa-ecd98b1ba15f · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

DocVLM: Make Your VLM an Efficient Reader PaliGemma: A versatile 3B VLM for transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.691914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.691914Z digest=sha256:91fceb2cf3beade28cae82e83d60b4d4b3e05413626d576645842e2928d96b40

Observation 6e5bfb32-75f0-4d42-90f3-71b82931c738 · outbound

This paper cites Scene text visual question answering.

DocVLM: Make Your VLM an Efficient Reader Scene text visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.761696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.695774Z digest=sha256:f5bb2add0a025e9cc5ff6241c97645d58371b7bffdbc1956fe68635846f03d20

Observation 680bee12-b608-4c91-9c1a-8a1c292b6423 · outbound

This paper cites Latr: Layout-aware transformer for scene-text vqa.

DocVLM: Make Your VLM an Efficient Reader Latr: Layout-aware transformer for scene-text vqa

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.739553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.699038Z digest=sha256:94d76050aa45f143bc17d98c6a8c0a014d14253761ff8ef5062c5e034acd5167

Observation a19bad3c-e51f-4dd0-8e1b-9e85c7880300 · outbound

This paper cites OCR-IDL: OCR Annotations for Industry Document Library Dataset.

DocVLM: Make Your VLM an Efficient Reader OCR-IDL: OCR Annotations for Industry Document Library Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.702449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.702449Z digest=sha256:e2f71b351f623a317bc4d0e1be6ab72b03b23d9c27e5e8db4017f42eb0783f2c

Observation 4b3582b1-5aba-429e-9051-a6e8b13c226c · outbound

This paper cites Gram: Global reasoning for multi-page vqa.

DocVLM: Make Your VLM an Efficient Reader Gram: Global reasoning for multi-page vqa

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.706094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.706094Z digest=sha256:69f9b179fd997822e5ce3d3309f6ab125bfbc43a12a28b28b3fd8d49916f23c4

Observation 263231ec-21af-428e-bcf0-b8591b249bcc · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

DocVLM: Make Your VLM an Efficient Reader Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.709581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.709581Z digest=sha256:092dd1eb92f2dcc33843f65fbc46281504fcd2e55a97e3ddd9612cdca550814e

Observation 70b2f9db-b7f6-457e-9d1f-2638aeb6eebb · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

DocVLM: Make Your VLM an Efficient Reader PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.713185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.713185Z digest=sha256:1df254472e6124d9c9b59e2509d3182dd5a2520f7fa408d1db4a6db17a210d35

Observation 0db02ade-4d5c-4fc9-80c0-dd8a8cc24eb9 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

DocVLM: Make Your VLM an Efficient Reader PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.716518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.716518Z digest=sha256:41df98086b204a43446d8bfb4a22a40d09eae4b376feb4520fb75c7ef86ad65d

Observation 3809b3e9-1c73-4ba8-b5c8-1daba241f17c · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,.

DocVLM: Make Your VLM an Efficient Reader Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.720325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.720325Z digest=sha256:0d18eae6e4a95e0a405d39c7e2fd2b33463996b18b84f367e651bbf0fe1ea6b4

Observation e34e2c5d-740a-430a-bb97-3b315914d230 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

DocVLM: Make Your VLM an Efficient Reader InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.723706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.723706Z digest=sha256:8380c2c52bb03a4333926c5436effc3506ee33c490e89fadbc0f44eaf3b69402

Observation a8a84652-7a5b-4304-b070-6d2d8702753e · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.727112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.727112Z digest=sha256:389814a1aa91907718bcc2d2c0b1b812ee066d721092cefac7ddb718cdb5c6d3

Observation 5d08ab50-a9f8-47a5-a6ea-1497d6abc638 · outbound

This paper cites Dtrocr: Decoder-only transformer for op- tical character recognition.

DocVLM: Make Your VLM an Efficient Reader Dtrocr: Decoder-only transformer for op- tical character recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.700195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.730476Z digest=sha256:f948c889ea9ea4b929fd9bea343620e5b1e079d2ed9ddf41c643d4c0729c6266

Observation 8e22a9ff-be1b-49c4-b305-6c5d0c8262c3 · outbound

This paper cites Towards models that can see and read.

DocVLM: Make Your VLM an Efficient Reader Towards models that can see and read

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.676824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.733829Z digest=sha256:6cce272f0e45d2e659e1659db4fc4f0310cc77f1f4cdd945caadf99d2f347574

Observation bb305f88-a348-4121-8b18-6bdf6dbe89d1 · outbound

This paper cites Question aware vision transformer for multimodal reasoning.

DocVLM: Make Your VLM an Efficient Reader Question aware vision transformer for multimodal reasoning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.661519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.737048Z digest=sha256:5057c6fb5cf2815b4ec22d5b461cc17b5f125f4a9a5fb88af5d28d4b2ea5c0eb

Observation 420b070a-ff4b-4341-bc40-f9ab7456889e · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

DocVLM: Make Your VLM an Efficient Reader Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.649027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.740285Z digest=sha256:5be7ccc0f39574208f4fc76259b93778e6649793cdf632f0bde9269860be35c3

Observation 74329bad-5407-48c5-8423-7176f78b40f2 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

DocVLM: Make Your VLM an Efficient Reader Funsd: A dataset for form understanding in noisy scanned documents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.637499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.743453Z digest=sha256:d1ffb2c6e9eb9151077873f0d24bfbed3b55d9c22eac509322e0bc9983def499

Observation c256a160-35cd-4c4f-85dc-b43b8871cd0f · outbound

This paper cites M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation.

DocVLM: Make Your VLM an Efficient Reader M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:42:15.161999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.748106Z digest=sha256:ca8d04b52c288168f87c8ed4414eea7dc7b8b33e06c07af38ad44cea6eff78d4

Observation 06682577-5388-4b29-b7fa-a806ecdc8dc3 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

DocVLM: Make Your VLM an Efficient Reader mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.751688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.751688Z digest=sha256:e3c1ce687c642ae821339fbea29f59f01ec1df375ce1d7a511966cb0e8e21a76

Observation 33ef0390-6e67-4dfa-9864-6e3c4f3dac4d · outbound

This paper cites Towards unified scene text spotting based on sequence generation.

DocVLM: Make Your VLM an Efficient Reader Towards unified scene text spotting based on sequence generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.625573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.755209Z digest=sha256:3d950a16915b5c24691d64d9f2ec042d757fe82ff1b770537730aaca43a34b5e

Observation e2e9d1b4-b0b2-40de-981a-93cbddac54b1 · outbound

This paper cites OCR-free Document Understanding Transformer.

DocVLM: Make Your VLM an Efficient Reader OCR-free Document Understanding Transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.758608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.758608Z digest=sha256:da30fbbf26b9453c22c2954ca66d907fd070fae6a52cdd9bde375fa46d90a8f6

Observation 7185aa3e-4c12-40d5-a219-ce5e089d5ed4 · outbound

This paper cites What matters when building vision-language models?.

DocVLM: Make Your VLM an Efficient Reader What matters when building vision-language models?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.762029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.762029Z digest=sha256:378c5a44b40d6c13101280e54c3f845167e3c2d23a88dd2b60b0c02a5c340d7a

Observation cccefa27-388d-4548-a17e-43f726d64b16 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DocVLM: Make Your VLM an Efficient Reader LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.765303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.765303Z digest=sha256:99b55906c2b3203a646a3b6efba9350d3c0f057466f62430a15585bbe86d1777

Observation 808b0bc4-6e05-4254-bdce-04d498a0186a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DocVLM: Make Your VLM an Efficient Reader Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.768805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.768805Z digest=sha256:55fdc238d3cbb3545c77f4627dfa91d1116278b1496318282965dc9b8e475288

Observation a0873528-fedb-49f4-88f7-8c9689af5dca · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DocVLM: Make Your VLM an Efficient Reader TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.772001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.772001Z digest=sha256:175a6ba43275c21e70f3f6abb6d2eb84b2787da8aa90c610967bfb54fb7d8536

Observation a1cb8b79-8268-4011-8271-c5f017051d93 · outbound

This paper cites Scatter: selective con- text attentional scene text recognizer.

DocVLM: Make Your VLM an Efficient Reader Scatter: selective con- text attentional scene text recognizer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.601849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.778183Z digest=sha256:2fe7e858160a5239a14e552e77310a304d33aea2f602d516b74f6f7c7b7345bb

Observation 9fdad15d-55ce-4431-9fe5-4d0bf66f134d · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

DocVLM: Make Your VLM an Efficient Reader Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.781530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.781530Z digest=sha256:0f7efca4122cfe8f7d3c1345dcb878d3720ad410e52b1cfbe7bfec02560f6c67

Observation b9272f10-53a6-48da-b75c-b78946a74760 · outbound

This paper cites Visual instruction tuning.

DocVLM: Make Your VLM an Efficient Reader Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.784614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.784614Z digest=sha256:42b41ae9ccf145c2c1ff0ae03ca0df0ad5050641ce40bcbc9db28daa5e62d018

Observation dd5f3872-6ef9-4f77-8e34-ccbb466a6e86 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

DocVLM: Make Your VLM an Efficient Reader Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.569569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.789406Z digest=sha256:8cd1a833eaabcd6404dc2962860ba1eafdbf424d0ef70ace6899857348d26e8f

Observation 4be20eb9-54f9-4a2e-8cb7-53b5e0f58641 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

DocVLM: Make Your VLM an Efficient Reader ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.792528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.792528Z digest=sha256:c171e2acd1de4cadd2bb34fa7ddf3100aca9fb35b473e9eebcf7ceda6855f9e1

Observation 83b1b8fb-4943-4c57-931b-5ba6817448e9 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DocVLM: Make Your VLM an Efficient Reader Docvqa: A dataset for vqa on document images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.553179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.796956Z digest=sha256:9153bffac43e0cda9ef617598321229dd1d6d0f0a21afdce8ffd9d004eef9aec

Observation 49244a03-65ff-401b-8ff1-1f842a332346 · outbound

This paper cites Infographicvqa.

DocVLM: Make Your VLM an Efficient Reader Infographicvqa

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.534751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.801218Z digest=sha256:7664ea09618590ce96cf9af0114a4ce7bb4624010bfbfb97ada926dc4631179f

Observation 19674d70-1e3f-49a8-88e9-4b1ee35b7cd7 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

DocVLM: Make Your VLM an Efficient Reader Ocr-vqa: Visual question answering by reading text in images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.804446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.804446Z digest=sha256:d4bfcb65131cef9e76a85c79e56eb41485af0f1a7f91635a7078e76795e69cbc

Observation aec45507-a783-4669-86b1-4295e6e6e4fd · outbound

This paper cites Textadain: Paying attention to shortcut learning in text recognizers.

DocVLM: Make Your VLM an Efficient Reader Textadain: Paying attention to shortcut learning in text recognizers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.493123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.807710Z digest=sha256:7f7b7769ec0ea79b0ce3c6a45bb50e6db2d446f10e1e67f58dc0a4d31339e056

Observation eca49a30-8850-46bd-a733-9f620a43f6f5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

DocVLM: Make Your VLM an Efficient Reader Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.810883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.810883Z digest=sha256:6c9ec43d31b65862ff550175a1a9a61b47ede5565e530f5fa173c9a6f7018313

Observation a1acf826-880f-4844-87cb-0403900ba6dd · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

DocVLM: Make Your VLM an Efficient Reader Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.814493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.814493Z digest=sha256:c85ba66dd0a4abbad9224f6077f7f945257011aabdf8e4fb5015429d84f957b8

Observation b4c32e6e-f4a5-4dfa-b566-d7031b1c6888 · outbound

This paper cites GLASS: Global to Local Attention for Scene-Text Spotting.

DocVLM: Make Your VLM an Efficient Reader GLASS: Global to Local Attention for Scene-Text Spotting

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.818997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.818997Z digest=sha256:a8f954c716d920d623c41267794ccf509f19024e2d5bf8bccf36eead9bf8f414

Observation 22912a0c-9d61-4879-b8fe-13698debcdd3 · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

DocVLM: Make Your VLM an Efficient Reader Textcaps: a dataset for image caption- ing with reading comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.822578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.822578Z digest=sha256:13e035847a047e1f4524e03a3319253ba7ad5a7d4505beb459e2c5ccbc972d45

Observation 3d30a0a9-8f37-44cb-99c9-20d469c48152 · outbound

This paper cites Towards vqa models that can read.

DocVLM: Make Your VLM an Efficient Reader Towards vqa models that can read

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.438395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.826725Z digest=sha256:1a2630641e7fdfa0739452935bac5e77df3ab34a39a25c3d56a6fb8753feb22a

Observation e2a6e898-08b5-4869-b4ea-79fc57b94d29 · outbound

This paper cites Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions.

DocVLM: Make Your VLM an Efficient Reader Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.421608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.832432Z digest=sha256:368054c90fc5add911b39dd0d4b5c92430ee7a91bfdc645bfca4b3169bdaf4fa

Observation 4fc6056f-ba81-4d87-9163-bd3fbb4aa3c4 · outbound

This paper cites Hi- erarchical multimodal transformers for multipage docvqa.

DocVLM: Make Your VLM an Efficient Reader Hi- erarchical multimodal transformers for multipage docvqa

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.391807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.836449Z digest=sha256:a8072119b251ef13c159e191b7875a7b50c9501f252583bad1e8bf8dddfc4a7d

Observation f2f38ca2-8ee0-429e-a95b-d9d29bfbb79a · outbound

This paper cites Document understanding dataset and evaluation (dude).

DocVLM: Make Your VLM an Efficient Reader Document understanding dataset and evaluation (dude)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.370799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.850835Z digest=sha256:5cb5557be5bbfb3f548098a87f6467eac4026d4bb18b893f4d1f86be74416f90

Observation 8a1ab35e-680a-4395-8b0d-539f5af5fb83 · outbound

This paper cites DocLLM: A layout-aware generative language model for multimodal document understanding.

DocVLM: Make Your VLM an Efficient Reader DocLLM: A layout-aware generative language model for multimodal document understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.858346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.858346Z digest=sha256:b5df9666bb282c4576d6eed021612307e6b6a53b9e74c0a7be2f19f16713574a

Observation c6e054a9-6d96-4299-8b1f-da5b50f2d127 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DocVLM: Make Your VLM an Efficient Reader Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.862395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.862395Z digest=sha256:3efebb6961afcc8dd6546f9eda6cff4595252f2eeade4f5e4d026250cb79f4cc

Observation c96d11ff-42d2-4c2a-9ca8-414d0cfd81f1 · outbound

This paper cites Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering.

DocVLM: Make Your VLM an Efficient Reader Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.867758Z digest=sha256:fdffdbbb44ffd2bf00e1c8c302d73b4cd4fbc9d439cb261ea3a08c59984acdc6

Observation b02716ba-9456-4415-9e20-88f23408f67f · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

DocVLM: Make Your VLM an Efficient Reader Layoutlm: Pre-training of text and layout for document image understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.359252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.873685Z digest=sha256:dbb1639eeb7664ec213525ba7729454babb6bdcc625f5f05169e882486767773

Observation 9f687888-e6cb-4ea7-8640-113138fd9b41 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

DocVLM: Make Your VLM an Efficient Reader UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.877803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.877803Z digest=sha256:741c47b27efddd9f12a6f0640f2262b17c2188bc9cbc49b1f59f1dbc9c68179e

Observation c9cca9cb-c68d-4cc0-8bb8-54592e6afc9e · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

DocVLM: Make Your VLM an Efficient Reader Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.347097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.881568Z digest=sha256:f0a73394d4b84859aa5055597a267fc8d579f2a3e9aa9146db8e3f1007689410

Observation a6f9ff17-c632-449a-b020-fb0fb9f84129 · outbound

This paper cites Deepsolo: Let transformer decoder with explicit points solo for text spot- ting.

DocVLM: Make Your VLM an Efficient Reader Deepsolo: Let transformer decoder with explicit points solo for text spot- ting

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.331773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.888590Z digest=sha256:8d5316cd5772d7d4a10c4fb8a87ea5287ad52dbc63f3f1f6c473dafbd750e054

Observation 4d194bfd-b861-4324-b4ef-628cc0342fc5 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

DocVLM: Make Your VLM an Efficient Reader mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.892017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.892017Z digest=sha256:0ae88c9fca5148abb24e9ed3a497359837e30f9d65854697eac110297456b5d5

Observation 381c7c14-7bc2-468b-adad-b45847ef488a · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.897792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.897792Z digest=sha256:c72c854651d7c09413eedac0b534dd89df4ab684c3bcb4837f6efa4b86e85b58

Observation f3afa750-7b12-474b-ae8e-56f25805e8af · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DocVLM: Make Your VLM an Efficient Reader MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.901281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.901281Z digest=sha256:66613263b754e7fd132d7998ecf1530227c41a1e256409b1845d540205e7b846

Observation 23c2e6cd-2f15-4de5-b2ee-35443e785d9d · outbound

This paper cites Towards complex doc- ument understanding by discrete reasoning.

DocVLM: Make Your VLM an Efficient Reader Towards complex doc- ument understanding by discrete reasoning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.318670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.905475Z digest=sha256:095922744879c2a94688e4c8707b80ed5101d927b505ab5496527e912952a78f

Observation c79542d8-7131-404a-813b-9090a688b5c9 · outbound

This paper cites The encoder is initial- ized with pretrained weights from DocFormerV2, which was pretrained on the Industry Document Library (IDL) dataset [13].

DocVLM: Make Your VLM an Efficient Reader The encoder is initial- ized with pretrained weights from DocFormerV2, which was pretrained on the Industry Document Library (IDL) dataset [13]

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.302121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T17:42:14.915935Z digest=sha256:c88e7cba4776765b704b1924f01029c378227e69b5f3e019ae94ec58a99bf052

Pith citing papers

Observation 83f55883-6cbf-4800-a317-78e8c3ccad98 · inbound

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production cites this paper.

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production DocVLM: Make Your VLM an Efficient Reader

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.404665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T21:55:02.656863Z digest=sha256:d7e6f5915914982e9a2c35a0c4a4605cc0cb6026e9052e4c5b7ab70f985277d9