Pith. sign in

Paper Citation Record · LEDGER

DocVLM: Make Your VLM an Efficient Reader

As of 14 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2412.08746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08746 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:42:14.915935Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T21:55:02.656863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T21:59:06.400938Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d95e635-eb90-46bc-8478-43366a9a697e · outbound

This paper cites Sequence-to-sequence contrastive learning for text recogni- tion.

DocVLM: Make Your VLM an Efficient Reader Sequence-to-sequence contrastive learning for text recogni- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.847032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.660000Z digest=sha256:24248fa23685dee8de25580e9338eedf4af57ff74416f8b86cce0a5b5502d5dd

Observation ce335f3e-c1c6-4b33-a514-a18942a47dd9 · outbound

This paper cites Multimodal Semi-Supervised Learning for Text Recognition.

DocVLM: Make Your VLM an Efficient Reader Multimodal Semi-Supervised Learning for Text Recognition

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:42:15.286747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.664037Z digest=sha256:e0ac10af717b8de8734b0cedccf2b97580c2e8030149c3671d8052394f768b37

Observation a83da128-c292-4c5a-82f0-a0123d3c23a5 · outbound

This paper cites Clipter: Looking at the bigger picture in scene text recogni- tion.

DocVLM: Make Your VLM an Efficient Reader Clipter: Looking at the bigger picture in scene text recogni- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.832160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.667866Z digest=sha256:9f2ae172c8b118375f35e2bd5260da96069d8363f71da9cf964e36ebdc044ea6

Observation 97adfa2a-0880-4d59-84bf-f52aa06e0634 · outbound

This paper cites Visfocus: Prompt- guided vision encoders for ocr-free dense document under- standing.

DocVLM: Make Your VLM an Efficient Reader Visfocus: Prompt- guided vision encoders for ocr-free dense document under- standing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.819896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.671218Z digest=sha256:fec45a74693a4b95c5c9151ff1194dc9819509a12cf97ab2da18adf7dd6d83fa

Observation 0dbb5cf3-6f5c-4557-bb3e-4aa26e03ea85 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

DocVLM: Make Your VLM an Efficient Reader Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.674632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.674632Z digest=sha256:cf3538101fdb1b7754b873df0853ec1edd1047c8b3d91073a4a696c79678592a

Observation abc368ed-db23-432e-b4c8-6a799a689b42 · outbound

This paper cites Docformer: End-to-end transformer for document understanding.

DocVLM: Make Your VLM an Efficient Reader Docformer: End-to-end transformer for document understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.801855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.678011Z digest=sha256:425bb721019bbf0962e26facfb87fb259dcbd3ae137afdc1566136213b9fbe7b

Observation 33997b18-557c-491d-bfe4-53f9caf214f8 · outbound

This paper cites Docformerv2: Local features for document understanding.

DocVLM: Make Your VLM an Efficient Reader Docformerv2: Local features for document understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.777421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.681409Z digest=sha256:b7b21e7b238c2d19e29eb3853262f4cd29f9028bc38765757f4e9270613286e4

Observation 603ac7e8-1290-4fd2-acee-b26cfdd395dd · outbound

This paper cites ScreenAI: A Vision-Language Model for UI and Infographics Understanding.

DocVLM: Make Your VLM an Efficient Reader ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.684571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.684571Z digest=sha256:71d06618c031f343579e8b0d9665d8c32b95051bbe34b57b08589675c46f16f7

Observation 86291faa-31e1-4997-9aed-7da95e13ebe8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DocVLM: Make Your VLM an Efficient Reader Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.688335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.688335Z digest=sha256:93d6bc5132221281b0e3e59a0352f215cecab5d39c4ce108895c635b54f0ea9f

Observation b4053255-729c-4b5c-9baa-ecd98b1ba15f · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

DocVLM: Make Your VLM an Efficient Reader PaliGemma: A versatile 3B VLM for transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.691914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.691914Z digest=sha256:351b517002039ec0ada3167cfcfc2a196a5606454de8fc538c234cc88a42c23c

Observation 6e5bfb32-75f0-4d42-90f3-71b82931c738 · outbound

This paper cites Scene text visual question answering.

DocVLM: Make Your VLM an Efficient Reader Scene text visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.761696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.695774Z digest=sha256:961620663d4f9b416142760944ad0429a114bfdc59cab0a69f683c6feeaf7788

Observation 680bee12-b608-4c91-9c1a-8a1c292b6423 · outbound

This paper cites Latr: Layout-aware transformer for scene-text vqa.

DocVLM: Make Your VLM an Efficient Reader Latr: Layout-aware transformer for scene-text vqa

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.739553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.699038Z digest=sha256:2a56813517b86df40e94742d811915d372781b5206f2d114e14e64692cab2d34

Observation a19bad3c-e51f-4dd0-8e1b-9e85c7880300 · outbound

This paper cites OCR-IDL: OCR Annotations for Industry Document Library Dataset.

DocVLM: Make Your VLM an Efficient Reader OCR-IDL: OCR Annotations for Industry Document Library Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.702449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.702449Z digest=sha256:b2b7c564a4e89fbf7ab2fde46713850e9ed85d818c0126bb9db729cff2c07b8a

Observation 4b3582b1-5aba-429e-9051-a6e8b13c226c · outbound

This paper cites Gram: Global reasoning for multi-page vqa.

DocVLM: Make Your VLM an Efficient Reader Gram: Global reasoning for multi-page vqa

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.706094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.706094Z digest=sha256:b6cab2b49780b6165cfd0665b4aefececfb08b6310bebd14b15a904fd949bf77

Observation 263231ec-21af-428e-bcf0-b8591b249bcc · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

DocVLM: Make Your VLM an Efficient Reader Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.709581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.709581Z digest=sha256:9e1aa42489063bf8ade77d65b9f7ab1faf5c80d8512abd768daea1900955ac2d

Observation 70b2f9db-b7f6-457e-9d1f-2638aeb6eebb · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

DocVLM: Make Your VLM an Efficient Reader PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.713185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.713185Z digest=sha256:fcf1166e6d5fc66e42555b2566fbe46e364b0f4e5f826b4eb5a663a93e359fcc

Observation 0db02ade-4d5c-4fc9-80c0-dd8a8cc24eb9 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

DocVLM: Make Your VLM an Efficient Reader PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.716518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.716518Z digest=sha256:96c27f8d36f29d5a01dcbf58be493aeff9ac8f074ef7a3c6662acf012773c264

Observation 3809b3e9-1c73-4ba8-b5c8-1daba241f17c · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,.

DocVLM: Make Your VLM an Efficient Reader Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.720325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.720325Z digest=sha256:3e1b2f267c383c1173c417025d40be81df9a795780ea5a7efc0f95fbba84c168

Observation e34e2c5d-740a-430a-bb97-3b315914d230 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

DocVLM: Make Your VLM an Efficient Reader InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.723706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.723706Z digest=sha256:aa3f2e295c12fe8c7da7f379d60245dfffa8a6afb09532c06811db3862823c8a

Observation a8a84652-7a5b-4304-b070-6d2d8702753e · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.727112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.727112Z digest=sha256:b7e21123ddbcbf31ce3bf90cc28655fb6c8e9acadd3d87aacd2c5fb65b5f8623

Observation 5d08ab50-a9f8-47a5-a6ea-1497d6abc638 · outbound

This paper cites Dtrocr: Decoder-only transformer for op- tical character recognition.

DocVLM: Make Your VLM an Efficient Reader Dtrocr: Decoder-only transformer for op- tical character recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.700195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.730476Z digest=sha256:4717c5a4ce4be56df4a89445dee2e5e29baaadb8f05c0d46cd960fcef03e3edd

Observation 8e22a9ff-be1b-49c4-b305-6c5d0c8262c3 · outbound

This paper cites Towards models that can see and read.

DocVLM: Make Your VLM an Efficient Reader Towards models that can see and read

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.676824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.733829Z digest=sha256:5d1e9186a8a438b76ffeed2497925d7debca5fe2efff06ca10f9190be9cea8e6

Observation bb305f88-a348-4121-8b18-6bdf6dbe89d1 · outbound

This paper cites Question aware vision transformer for multimodal reasoning.

DocVLM: Make Your VLM an Efficient Reader Question aware vision transformer for multimodal reasoning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.661519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.737048Z digest=sha256:24ec1778ce4b0446b1c8a5b5caf4d64c257b86b30f35ce07699c1a68539a5957

Observation 420b070a-ff4b-4341-bc40-f9ab7456889e · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

DocVLM: Make Your VLM an Efficient Reader Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.649027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.740285Z digest=sha256:38afcfd482747d1c63b01afc45293b2c0aabd78fd321f5d8692920d5158a0de8

Observation 74329bad-5407-48c5-8423-7176f78b40f2 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

DocVLM: Make Your VLM an Efficient Reader Funsd: A dataset for form understanding in noisy scanned documents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.637499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.743453Z digest=sha256:67ef36d340d9a0af22da586ca6faff31d8dfc74cf302f9c1690750c73f06fad1

Observation c256a160-35cd-4c4f-85dc-b43b8871cd0f · outbound

This paper cites M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation.

DocVLM: Make Your VLM an Efficient Reader M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:42:15.161999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.748106Z digest=sha256:9bf1a4e9271de4ca0f90cd0cd243b0d92bba1cc9d2f9818b9b9e672a5288fb32

Observation 06682577-5388-4b29-b7fa-a806ecdc8dc3 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

DocVLM: Make Your VLM an Efficient Reader mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.751688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.751688Z digest=sha256:1c04b2df4af2d0ccec047d384f43c23f654f44ffa897f40a8df5fd39b6f4643a

Observation 33ef0390-6e67-4dfa-9864-6e3c4f3dac4d · outbound

This paper cites Towards unified scene text spotting based on sequence generation.

DocVLM: Make Your VLM an Efficient Reader Towards unified scene text spotting based on sequence generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.625573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.755209Z digest=sha256:34cf32a98f677e86d1320de1f59d55f6f06e3a65db853d38a7a540febef001ce

Observation e2e9d1b4-b0b2-40de-981a-93cbddac54b1 · outbound

This paper cites OCR-free Document Understanding Transformer.

DocVLM: Make Your VLM an Efficient Reader OCR-free Document Understanding Transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.758608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.758608Z digest=sha256:1137d8a083e1fa8390f42af9aa9b64ea1127ccff592980b9ba567aca881ff675

Observation 7185aa3e-4c12-40d5-a219-ce5e089d5ed4 · outbound

This paper cites What matters when building vision-language models?.

DocVLM: Make Your VLM an Efficient Reader What matters when building vision-language models?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.762029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.762029Z digest=sha256:d2bab85bb655e8e11ed1f2471e13df095c0b4deac1e3dd32672865418290753d

Observation cccefa27-388d-4548-a17e-43f726d64b16 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DocVLM: Make Your VLM an Efficient Reader LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.765303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.765303Z digest=sha256:2bf55248a8d8ca82c5cf0c5d09942598d5519cb90348bc415469f0daaa5d3b5d

Observation 808b0bc4-6e05-4254-bdce-04d498a0186a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DocVLM: Make Your VLM an Efficient Reader Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.768805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.768805Z digest=sha256:cc2c3637db0ed195a18ce071998a5a59abd07e01e0676811f3cb7f429a93a563

Observation a0873528-fedb-49f4-88f7-8c9689af5dca · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DocVLM: Make Your VLM an Efficient Reader TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.772001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.772001Z digest=sha256:b9f5e8dcea86f66c2e03dc046ba4a4b02da3d6bdc38154013a8ca806cbd5ecb5

Observation a1cb8b79-8268-4011-8271-c5f017051d93 · outbound

This paper cites Scatter: selective con- text attentional scene text recognizer.

DocVLM: Make Your VLM an Efficient Reader Scatter: selective con- text attentional scene text recognizer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.601849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.778183Z digest=sha256:50c1eea395c8fb0032ed3154a0542a29869b943c5fd658d2387bf60094b3c624

Observation 9fdad15d-55ce-4431-9fe5-4d0bf66f134d · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

DocVLM: Make Your VLM an Efficient Reader Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.781530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.781530Z digest=sha256:47c5f43d5c5f5faf84219460172f1f3b3c81f7b0badf6a205a6217ef9552a650

Observation b9272f10-53a6-48da-b75c-b78946a74760 · outbound

This paper cites Visual instruction tuning.

DocVLM: Make Your VLM an Efficient Reader Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.784614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.784614Z digest=sha256:9357e0c0a9aba9716d3dad0ca3a6bdb942404fbb8a432405ea6d0c643f0dfb2c

Observation dd5f3872-6ef9-4f77-8e34-ccbb466a6e86 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

DocVLM: Make Your VLM an Efficient Reader Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.569569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.789406Z digest=sha256:8857416589bb441a3c5ffc6296d5d625d9a0fc938ebd2d1ea10847e1ec6af758

Observation 4be20eb9-54f9-4a2e-8cb7-53b5e0f58641 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

DocVLM: Make Your VLM an Efficient Reader ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.792528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.792528Z digest=sha256:0c3268712b7228362b27a31b4e3d5240a91060ea79091fd5fd5acd14ba1455d5

Observation 83b1b8fb-4943-4c57-931b-5ba6817448e9 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DocVLM: Make Your VLM an Efficient Reader Docvqa: A dataset for vqa on document images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.553179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.796956Z digest=sha256:591b8219b1579b45b24f1b991475d927cc1b373bd81ba12f620ff1276025e13a

Observation 49244a03-65ff-401b-8ff1-1f842a332346 · outbound

This paper cites Infographicvqa.

DocVLM: Make Your VLM an Efficient Reader Infographicvqa

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.534751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.801218Z digest=sha256:2ce69d78ce497a8cb4149df188d437e20150500a310c5daf58cc0532dba9df9d

Observation 19674d70-1e3f-49a8-88e9-4b1ee35b7cd7 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

DocVLM: Make Your VLM an Efficient Reader Ocr-vqa: Visual question answering by reading text in images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.804446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.804446Z digest=sha256:41ecbacc7912651bff23fa09d267788271ddc1c64374c9f386f98c097bdb8b5c

Observation aec45507-a783-4669-86b1-4295e6e6e4fd · outbound

This paper cites Textadain: Paying attention to shortcut learning in text recognizers.

DocVLM: Make Your VLM an Efficient Reader Textadain: Paying attention to shortcut learning in text recognizers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.493123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.807710Z digest=sha256:fa0caf7d7e3b17a82cfa7dd4fb4f4e7111a8adaa147cd78754e0de251eff0273

Observation eca49a30-8850-46bd-a733-9f620a43f6f5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

DocVLM: Make Your VLM an Efficient Reader Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.810883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.810883Z digest=sha256:4ee7277c3912ea44ccc8de3446e1d5313d3f42c885afb82ebd173c301547dae2

Observation a1acf826-880f-4844-87cb-0403900ba6dd · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

DocVLM: Make Your VLM an Efficient Reader Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.814493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.814493Z digest=sha256:3539d80e2d6f66608299b45df006e3f0251f00c951785d212a9314f47310f903

Observation b4c32e6e-f4a5-4dfa-b566-d7031b1c6888 · outbound

This paper cites GLASS: Global to Local Attention for Scene-Text Spotting.

DocVLM: Make Your VLM an Efficient Reader GLASS: Global to Local Attention for Scene-Text Spotting

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.818997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.818997Z digest=sha256:d116cd41a3ae205000d3c076360357d03353bbf0938ee6718ec6e9a4296fcb5e

Observation 22912a0c-9d61-4879-b8fe-13698debcdd3 · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

DocVLM: Make Your VLM an Efficient Reader Textcaps: a dataset for image caption- ing with reading comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.822578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.822578Z digest=sha256:1522bcccc5cc2ad478d29239700655db7337d72c4cdae17f39c8d76c66f05292

Observation 3d30a0a9-8f37-44cb-99c9-20d469c48152 · outbound

This paper cites Towards vqa models that can read.

DocVLM: Make Your VLM an Efficient Reader Towards vqa models that can read

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.438395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.826725Z digest=sha256:40c2adbc9547c5bde665c072078b1430c5128c6a417d210da23e836793b57187

Observation e2a6e898-08b5-4869-b4ea-79fc57b94d29 · outbound

This paper cites Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions.

DocVLM: Make Your VLM an Efficient Reader Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.421608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.832432Z digest=sha256:7886f2580523c69fa6ab43821fdc190a74ee1dc08b3b42c36128f99b64c8d53c

Observation 4fc6056f-ba81-4d87-9163-bd3fbb4aa3c4 · outbound

This paper cites Hi- erarchical multimodal transformers for multipage docvqa.

DocVLM: Make Your VLM an Efficient Reader Hi- erarchical multimodal transformers for multipage docvqa

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.391807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.836449Z digest=sha256:53bf4d646e73634408008556e359620ccc9e00a5301c2ebd889e1cf7291588a5

Observation f2f38ca2-8ee0-429e-a95b-d9d29bfbb79a · outbound

This paper cites Document understanding dataset and evaluation (dude).

DocVLM: Make Your VLM an Efficient Reader Document understanding dataset and evaluation (dude)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.370799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.850835Z digest=sha256:d42ce39d7a89ff4289fce03e22cd756f880adaec19a5f9ecbfc26c8e41a33620

Observation 8a1ab35e-680a-4395-8b0d-539f5af5fb83 · outbound

This paper cites DocLLM: A layout-aware generative language model for multimodal document understanding.

DocVLM: Make Your VLM an Efficient Reader DocLLM: A layout-aware generative language model for multimodal document understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.858346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.858346Z digest=sha256:abdf6ef242ab30984cf71ead84e8df5103652c6a0cc3ba4e1c087f5ad6328c83

Observation c6e054a9-6d96-4299-8b1f-da5b50f2d127 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DocVLM: Make Your VLM an Efficient Reader Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.862395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.862395Z digest=sha256:2de3eccad52fc0c3f4c318d8423e56a67919fe260ae475b4f29b361ef66ca35b

Observation c96d11ff-42d2-4c2a-9ca8-414d0cfd81f1 · outbound

This paper cites Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering.

DocVLM: Make Your VLM an Efficient Reader Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.867758Z digest=sha256:1768b41189c5c006d3b5af2d8a6f96e0bb1ce4932a635ec073d8c028cbb46983

Observation b02716ba-9456-4415-9e20-88f23408f67f · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

DocVLM: Make Your VLM an Efficient Reader Layoutlm: Pre-training of text and layout for document image understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.359252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.873685Z digest=sha256:d812f4e821f64ee69bce95157489ed88c6546bd9b97b9ddef6988678cec25176

Observation 9f687888-e6cb-4ea7-8640-113138fd9b41 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

DocVLM: Make Your VLM an Efficient Reader UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.877803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.877803Z digest=sha256:bdf99045d73e7c6ec9684ad282ad36fbe00233dff046c67e28fe0f0d3837f77d

Observation c9cca9cb-c68d-4cc0-8bb8-54592e6afc9e · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

DocVLM: Make Your VLM an Efficient Reader Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.347097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.881568Z digest=sha256:c5d0f4b6a4fe8cd64776fc5b56791738bd4031bb34057e0c1d8a5ab4e5da4cf1

Observation a6f9ff17-c632-449a-b020-fb0fb9f84129 · outbound

This paper cites Deepsolo: Let transformer decoder with explicit points solo for text spot- ting.

DocVLM: Make Your VLM an Efficient Reader Deepsolo: Let transformer decoder with explicit points solo for text spot- ting

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.331773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.888590Z digest=sha256:9163529ae778560e9f7f044096f1335f8954f3279fa793c3936a2cf340a87df6

Observation 4d194bfd-b861-4324-b4ef-628cc0342fc5 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

DocVLM: Make Your VLM an Efficient Reader mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.892017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.892017Z digest=sha256:c806a460a6dcd0c8eb303a3e0a3ff78ae8a945a5e582585abca0719febda1b1b

Observation 381c7c14-7bc2-468b-adad-b45847ef488a · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.897792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.897792Z digest=sha256:5271a0b590822d0e6e1fd9b6dfcffa9c1b8a06962aa2b6961152cbfa0614f19e

Observation f3afa750-7b12-474b-ae8e-56f25805e8af · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DocVLM: Make Your VLM an Efficient Reader MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.901281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.901281Z digest=sha256:7fbf6cf1cda6477a5aa6af4536b481fd9812c56784d55aa4cfec81ab448872d2

Observation 23c2e6cd-2f15-4de5-b2ee-35443e785d9d · outbound

This paper cites Towards complex doc- ument understanding by discrete reasoning.

DocVLM: Make Your VLM an Efficient Reader Towards complex doc- ument understanding by discrete reasoning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.318670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.905475Z digest=sha256:7c032ebcb830660395fdc9cd5b0e4fcc3dd90c92c9ba78b51d260854fdb77252

Observation c79542d8-7131-404a-813b-9090a688b5c9 · outbound

This paper cites The encoder is initial- ized with pretrained weights from DocFormerV2, which was pretrained on the Industry Document Library (IDL) dataset [13].

DocVLM: Make Your VLM an Efficient Reader The encoder is initial- ized with pretrained weights from DocFormerV2, which was pretrained on the Industry Document Library (IDL) dataset [13]

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.302121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.915935Z digest=sha256:fc801068a5ca5cf852fc569babf4c218acabc8cea80081f24adce595c36d4896

Pith citing papers

Observation 83f55883-6cbf-4800-a317-78e8c3ccad98 · inbound

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production cites this paper.

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production DocVLM: Make Your VLM an Efficient Reader

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.404665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T21:55:02.656863Z digest=sha256:ee755ff3c865f4103fb6ce643cbf76910fb1f7abe947db48a274809f185bc26e