Pith. sign in

Paper Citation Record · LEDGER

DOGR: Towards Versatile Visual Document Grounding and Referring

As of 22 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 6 inbound Pith citation observations for arXiv:2411.17125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17125 v3

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:34:31.918078Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:11.055077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:35:07.388169Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26bfbaeb-8f0a-4d97-81af-e54ea2c53471 · outbound

This paper cites Qwen2.5-VL Technical Report.

DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.648211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.648211Z digest=sha256:8eff3ee686c990bfa63fb25d5762084aafa20e8dd050dd76ca84f120831d975a

Observation 24a4114f-9514-43c0-99d1-d27c2a3f1906 · outbound

This paper cites Jawahar, Ernest Valveny, and Dimos- thenis Karatzas.

DOGR: Towards Versatile Visual Document Grounding and Referring Jawahar, Ernest Valveny, and Dimos- thenis Karatzas

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.916172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.653462Z digest=sha256:8e8ac78df3789a609404d38b60e5e66cfc9025e1e4b44ec2c8e0106bd26d443d

Observation 3bdda016-7e96-44d9-82cc-85788fd7cd7a · outbound

This paper cites Textocr-gpt4v.

DOGR: Towards Versatile Visual Document Grounding and Referring Textocr-gpt4v

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.902951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.657301Z digest=sha256:92e40a9a64d17c2468d4c0d64842cbb1c83aaee49b02ce245e049c81bd391251

Observation 17f76eeb-cb04-45da-ac8d-ea8681731f84 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

DOGR: Towards Versatile Visual Document Grounding and Referring Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.661116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.661116Z digest=sha256:b6bb3e6b3905a0e6fa3eb308154a5f4ce8762e9b5eb0a72ce3cb6fb47063affe

Observation 6bd12d6c-c3d4-4aa0-aeb2-f7bcec147bfb · outbound

This paper cites Tabfact : A large-scale dataset for table-based fact verification.

DOGR: Towards Versatile Visual Document Grounding and Referring Tabfact : A large-scale dataset for table-based fact verification

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.890189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.665898Z digest=sha256:e2eff6eb27340bf5f2f7b9ea5185a69cff48488685e35cf0f974096f0695780c

Observation 6a3282b8-7a6b-4543-9a1f-523f3ca78d66 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

DOGR: Towards Versatile Visual Document Grounding and Referring InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.669626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.669626Z digest=sha256:084592f23a0bacb3e9908ec354b9f863b979524ddd2da185bde36a51892c15fd

Observation 11801deb-4f35-47ee-946e-a8a415b08cbe · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DOGR: Towards Versatile Visual Document Grounding and Referring Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.673783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.673783Z digest=sha256:624d2db7b41e623eba542db4fd300c7a9365676ce97ec4d60f71b75a8c0d100e

Observation 1ae2b981-6d89-453b-96ca-01ddd952001c · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

DOGR: Towards Versatile Visual Document Grounding and Referring How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.678432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.678432Z digest=sha256:8f4269c2fac76e65c6895e3339efb15adcbad08af3b5a3600fa7987f602be2ff

Observation b5be5ab8-9bc1-44a1-add0-7135206b6aa0 · outbound

This paper cites Hitab: A hierarchical table dataset for question an- swering and natural language generation.

DOGR: Towards Versatile Visual Document Grounding and Referring Hitab: A hierarchical table dataset for question an- swering and natural language generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.877411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.682552Z digest=sha256:69a5f29bdb7e3a46be7385cf0b3b274ad289902636ef766ec1d3c6d55eac8401

Observation 0833509a-d22a-43bd-b5b2-dc5b3badfb5a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

DOGR: Towards Versatile Visual Document Grounding and Referring Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.686379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.686379Z digest=sha256:73f374840ca667702b643dc344b66ae90629fadf2d9839afa5946b47047f04fa

Observation 86c18ed8-c0ac-4dd9-8bae-979d92f363c7 · outbound

This paper cites Pymupdf: Python bindings for mupdf.

DOGR: Towards Versatile Visual Document Grounding and Referring Pymupdf: Python bindings for mupdf

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.864954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.690243Z digest=sha256:67bb8193880ff0fb4bf063bf230f5f020b947eb86326174974579b725fbc0aee

Observation 02369ccb-e3bd-494a-a328-777d4ed14c00 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

DOGR: Towards Versatile Visual Document Grounding and Referring InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.694114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.694114Z digest=sha256:f6125151c50983963811b8460adec64a5ab4556379bd560f5ec4c457266ff728

Observation ff194cb4-9bee-4e45-ab8c-d9081b52d5a2 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

DOGR: Towards Versatile Visual Document Grounding and Referring mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.702415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.702415Z digest=sha256:eff872e0891f606332d58bab8c9f7bae842f5d645a2e4a2aa988f1154fcdae48

Observation 3c2d8526-7b32-4d4b-ab24-bd8dbacfbb5b · outbound

This paper cites mplug- docowl2: High-resolution compressing for ocr-free multi- page document understanding, 2024.

DOGR: Towards Versatile Visual Document Grounding and Referring mplug- docowl2: High-resolution compressing for ocr-free multi- page document understanding, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.851190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.706461Z digest=sha256:e4f649da7be670469f1e8cc6a9e90a7e40ac5b57385d052bd3955fa6a457b999

Observation c4b5a9a1-23d7-4426-9464-4172398ab6e7 · outbound

This paper cites GPT-4o System Card.

DOGR: Towards Versatile Visual Document Grounding and Referring GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.710462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.710462Z digest=sha256:0fc56fde1f5c3a5f13caba7286765ed87d20f5c942198fd56f2917503add7ff5

Observation 685b00ea-5278-4467-b538-07ddaa96579e · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

DOGR: Towards Versatile Visual Document Grounding and Referring Dvqa: Understanding data visualizations via ques- tion answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.838579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.714509Z digest=sha256:066ab068a693cb3863b4f5ecd385951f91c31a9a7e535f882a81cff1f29dd214

Observation bf19eda8-5938-4e02-9bbe-f4aa97156073 · outbound

This paper cites Fig- ureqa: An annotated figure dataset for visual reasoning,.

DOGR: Towards Versatile Visual Document Grounding and Referring Fig- ureqa: An annotated figure dataset for visual reasoning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.718388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.718388Z digest=sha256:c771d807f5431fdfb25d71e6a7d8255ffefdb94c6375d4a74851ae3e3e2aef1d

Observation abc404f9-3788-4835-88f4-690e59522ef0 · outbound

This paper cites A diagram is worth a dozen images.

DOGR: Towards Versatile Visual Document Grounding and Referring A diagram is worth a dozen images

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.818590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.722309Z digest=sha256:6b6ff6e17e34794d56f4a5b712440b55b5a28598861d048e1704cb01866f1388

Observation 3201579a-ddad-4b98-bd1a-cf897857c69a · outbound

This paper cites Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension.

DOGR: Towards Versatile Visual Document Grounding and Referring Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.806395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.726598Z digest=sha256:883586ba646e4897c5f177c58ef485055b039a8f092d747a9a6c0bb7082de4b6

Observation 8254ade0-206f-4b5e-bed7-4b9587714bb1 · outbound

This paper cites Ocr-free document understanding transformer.

DOGR: Towards Versatile Visual Document Grounding and Referring Ocr-free document understanding transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.793430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.730781Z digest=sha256:e5507d4e37d3eceeed535d099651ff555d1089a4fdad047aa5e9e324a38dc31b

Observation bd41f984-f577-42d4-a053-51f817060538 · outbound

This paper cites TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation.

DOGR: Towards Versatile Visual Document Grounding and Referring TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.734719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.734719Z digest=sha256:0038161b47f069b3a5f838538dbd6ae2a21d830a64565cfbb61ad7fb356d702c

Observation 233f6a0a-37ed-4caf-b440-9f7b4dbe29bb · outbound

This paper cites Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,.

DOGR: Towards Versatile Visual Document Grounding and Referring Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.781648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.738941Z digest=sha256:0f030cf33bedd1484d586722e611e2dc9679bca1e68eae13ca347412398c5d38

Observation f8c0dc91-a888-4c5e-8823-ec4ddd7802b6 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

DOGR: Towards Versatile Visual Document Grounding and Referring Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.746755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.746755Z digest=sha256:641461bb7b5ef74594da1b5f98e76c3fc0bf7da60590155e110e329ae2c151aa

Observation 82070d9c-33c0-4d3d-a63d-994a987b1975 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

DOGR: Towards Versatile Visual Document Grounding and Referring Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.750312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.750312Z digest=sha256:2a7d1e0558c57ce856f1e77cba182766bce61f92b753d180f833bd24db72da23

Observation b86853e7-b33e-4e6d-a01f-1ae64ff2d84f · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

DOGR: Towards Versatile Visual Document Grounding and Referring MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.754314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.754314Z digest=sha256:5ed6aab6f751e2873531a1659aaa73cc4e92b9d79bd32042d906d77b8fe62a54

Observation 3a49c185-310c-466b-9d49-ddf16efa2cb9 · outbound

This paper cites Visual instruction tuning, 2023.

DOGR: Towards Versatile Visual Document Grounding and Referring Visual instruction tuning, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.758827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.758827Z digest=sha256:f5c13edd3947b9a0838702f1af3d1997e6ea6c9960e218d0b900fadcd9c66639

Observation cf4bc3c0-0fc8-4a09-808e-3e2358f57ebe · outbound

This paper cites Textmonkey: An ocr-free large multimodal model for understanding document, 2024.

DOGR: Towards Versatile Visual Document Grounding and Referring Textmonkey: An ocr-free large multimodal model for understanding document, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.762840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.763097Z digest=sha256:9e219a935d132d8acdf19194699963b785f99fd08c0755bb48c61ed26b764a99

Observation 55824c0c-9478-463f-8219-ebe694609b52 · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

DOGR: Towards Versatile Visual Document Grounding and Referring KOSMOS-2.5: A Multimodal Literate Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.768116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.768116Z digest=sha256:3a65f2857b11c66ab946632fb0fd4b45633b0bbdeb9bcd6078eace69550fe8a7

Observation b46eb2d2-7fa6-4847-8a65-a189fd7ff92f · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

DOGR: Towards Versatile Visual Document Grounding and Referring The iam-database: an english sentence database for offline handwriting recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.750409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.772448Z digest=sha256:2122e6b17f8675b84a089defe764c4fbd9462c47f124051b48f1570e3ec070a5

Observation b971a719-8ed5-4cd5-801f-d4ba6aa2f63c · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

DOGR: Towards Versatile Visual Document Grounding and Referring ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.737825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.776555Z digest=sha256:0b5222b603b627752360d3586418adafacc8dd201ab286a98186dd757691d18b

Observation cfb35cac-fa5c-485d-8c02-57a3c615389e · outbound

This paper cites Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning.

DOGR: Towards Versatile Visual Document Grounding and Referring Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.723847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.780990Z digest=sha256:6ee146ba374e37f7cc139d52a2cc11f3ad69a76938cac83388ef7d709356eb60

Observation 4d854a3b-a335-4c5b-b400-fb0d838884a2 · outbound

This paper cites an unresolved cited work.

DOGR: Towards Versatile Visual Document Grounding and Referring Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:34:32.710777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.784819Z digest=sha256:d4eebaa5811f214a3487811f1e56aed2cb1271267493c10a181eb75375929ce6

Observation fc59c0a9-032d-41cc-917b-9dcc68d1af8a · outbound

This paper cites an unresolved cited work.

DOGR: Towards Versatile Visual Document Grounding and Referring Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:34:32.698912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.789514Z digest=sha256:8d917b54f9c7ec55dceaf1ae8060d297e1bed6a6344bd05711508bba29d897f2

Observation ff135251-6a65-4408-b668-52e9ca751260 · outbound

This paper cites Mishra, K.

DOGR: Towards Versatile Visual Document Grounding and Referring Mishra, K

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.686897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.793645Z digest=sha256:15a89620b093ad24627f7684019ecb256674d6898c47fcaada99ff54d9c2c865

Observation f8b951c5-e9b6-4b9e-ad33-16912428e8d7 · outbound

This paper cites Chart-to-text: Generat- ing natural language descriptions for charts by adapting the transformer model, 2020.

DOGR: Towards Versatile Visual Document Grounding and Referring Chart-to-text: Generat- ing natural language descriptions for charts by adapting the transformer model, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.674676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.798015Z digest=sha256:8455bb1a8c2745240b7a2f82458e645593d9b0aec8cd781223e4f2b37bb6bf34

Observation 73c958f6-3054-4c68-aedd-73337ca00cd2 · outbound

This paper cites Compositional semantic parsing on semi-structured tables.

DOGR: Towards Versatile Visual Document Grounding and Referring Compositional semantic parsing on semi-structured tables

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.662272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.801879Z digest=sha256:48408b1591ec871cddc13c6e0d59003c9ee2bdf438f7b32ccbb667cbf2597958

Observation bcb6e8df-9dc2-4b8c-85fc-5b4f0032ac27 · outbound

This paper cites Kosmos-2: Ground- ing multimodal large language models to the world.

DOGR: Towards Versatile Visual Document Grounding and Referring Kosmos-2: Ground- ing multimodal large language models to the world

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.805555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.805555Z digest=sha256:3ffe56d4b9c206bbbd56dfe67bf52848c2dc1c69bb496f1f65626ef4a983668d

Observation c4006fc3-7d27-4cd1-94da-e727d5633d65 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

DOGR: Towards Versatile Visual Document Grounding and Referring Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.641077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.809155Z digest=sha256:3d623636fa7c7dcd7fd96c48cf6c1e52d059d92d01b34a219a186bd5055fe547

Observation 139bd0b3-a96d-496f-9f22-13e542788837 · outbound

This paper cites Textcaps: a dataset for image caption- ingwith reading comprehension.

DOGR: Towards Versatile Visual Document Grounding and Referring Textcaps: a dataset for image caption- ingwith reading comprehension

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.627247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.812360Z digest=sha256:97dc39987f0a05ed2ee9918d00b6da1c713791d2afcc9815a367e9528bfa8415

Observation 6aee4a9f-96da-4d56-91db-0d9ac003b27e · outbound

This paper cites Towards vqa models that can read.

DOGR: Towards Versatile Visual Document Grounding and Referring Towards vqa models that can read

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.614643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.815734Z digest=sha256:886ea86342b573a599196a8a266ad985a67415ddf7ca54ffef86a17e249cbeaf

Observation dd376bc9-496b-41e7-9f7c-a4a872091e98 · outbound

This paper cites Kleister: Key in- formation extraction datasets involving long documents with complex layouts.

DOGR: Towards Versatile Visual Document Grounding and Referring Kleister: Key in- formation extraction datasets involving long documents with complex layouts

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.468596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.819457Z digest=sha256:e3dc7a87e114b26e430189aed1dbcc36a2c3e80e77e137a31b93b1f356dc3cfc

Observation 53f13a80-5f57-42a4-b9d7-25d681224026 · outbound

This paper cites Deepform: Understand structured docu- ments at scale.

DOGR: Towards Versatile Visual Document Grounding and Referring Deepform: Understand structured docu- ments at scale

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.454889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.823183Z digest=sha256:4d962684b99f916dc1f11162e695fdde19009cbd9ece37b393a059815be0acc1

Observation b6e509b7-96b3-4561-9974-20b284045ab9 · outbound

This paper cites Vi- sualmrc: Machine reading comprehension on document im- ages.

DOGR: Towards Versatile Visual Document Grounding and Referring Vi- sualmrc: Machine reading comprehension on document im- ages

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.436897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.827900Z digest=sha256:c73bc7c7efc98b2af85a4b1f588ac68dd671d057ac6ac791ca4cdc25f086e652

Observation 16a1f5ae-9a57-492d-ab18-c1ea2f67bf5d · outbound

This paper cites Tang, Angie Boggust, and Arvind Satyanarayan.

DOGR: Towards Versatile Visual Document Grounding and Referring Tang, Angie Boggust, and Arvind Satyanarayan

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.831664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.831664Z digest=sha256:4fbe0b8ff8727ee1eb8f09c60ba92b1a4f7c3643e43692b9704de3e85bd684ee

Observation 4a88d28f-3ed9-43c3-9fab-4cf55b9231c2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DOGR: Towards Versatile Visual Document Grounding and Referring Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.836067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.836067Z digest=sha256:a1d74f58c6ddc4f53f16b1c97f870c2e76b3651fdd05fd756eb04b93cd18c64b

Observation c97f8457-39f0-4a02-b8e9-5a0f22fd0dcb · outbound

This paper cites Cc-main-2021-31-pdf- untruncated.

DOGR: Towards Versatile Visual Document Grounding and Referring Cc-main-2021-31-pdf- untruncated

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.412470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.840254Z digest=sha256:74a25c7732347d3899c2c75561fbec487ab4909038672d5220592298b3a50e54

Observation 73d5a102-9789-401d-ad47-2de3a29363ea · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

DOGR: Towards Versatile Visual Document Grounding and Referring Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.844044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.844044Z digest=sha256:437f746cfc57ea426c947126fcb6f0cef209cddc777f222fdaaa905062eeced9

Observation 002e889c-ada0-4abe-8e71-2fd052a4fbfe · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

DOGR: Towards Versatile Visual Document Grounding and Referring Lawrence Zitnick, and Devi Parikh

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.399990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.848176Z digest=sha256:dbfd42f72d6ca80de7f33269d3b48fd89ef2b79a1fcea102c756651c3f76c283

Observation 833c0e39-612c-45fd-992d-e3ab4ab6ccd2 · outbound

This paper cites Screen2words: Automatic mobile ui summarization with multimodal learning, 2021.

DOGR: Towards Versatile Visual Document Grounding and Referring Screen2words: Automatic mobile ui summarization with multimodal learning, 2021

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.388039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.851720Z digest=sha256:36e333812e1da5cb4aba9b580251b00a654dfdd9f202b23fa10b4769cc98bec2

Observation a6ae5d35-1d18-43a3-8b4e-ec6e9f6e336c · outbound

This paper cites Mineru: An open-source solution for precise document content extrac- tion, 2024.

DOGR: Towards Versatile Visual Document Grounding and Referring Mineru: An open-source solution for precise document content extrac- tion, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.375706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.855418Z digest=sha256:ae7192237989bca1c19e0f1b19d724e7927196a6ae51fc68f0507220bf9d7594

Observation fb01f736-d8c8-492b-8506-5086074fa2e6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.859533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.859533Z digest=sha256:268da0bdf837d42c4591fbec570af4184e8b7fb7c6bbf9b8b55d79882f82036a

Observation 03780143-2616-445c-ba74-cfe102edf024 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

DOGR: Towards Versatile Visual Document Grounding and Referring Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.863650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.863650Z digest=sha256:63e9c0a565e0a7fc3f067c4250c4a039a34d65fc98f7b8515b90fce5b6ea37fd

Observation 91f13e26-c975-4900-a338-4ae045535b86 · outbound

This paper cites wendlerc/renderedtext, 2023.

DOGR: Towards Versatile Visual Document Grounding and Referring wendlerc/renderedtext, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.362511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.868240Z digest=sha256:3af2bb9636ea9cf29e6021d43ad42ae26be63f3470a11cff8ac96c2da2100895

Observation 39990f8d-4f61-4327-acb1-f0f1310396da · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

DOGR: Towards Versatile Visual Document Grounding and Referring Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.873021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.873021Z digest=sha256:d7db860964f00d3db5d5c61cd38b8b21a16a124321f72f7fbb94c5a2020de531

Observation 397fcff5-cc26-4214-9c8a-d9b81bec51da · outbound

This paper cites Canvasvae: Learning to generate vector graphic documents.

DOGR: Towards Versatile Visual Document Grounding and Referring Canvasvae: Learning to generate vector graphic documents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.877150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.877150Z digest=sha256:1b260e170808fb023ebdcfd135373a19efd52edb7f14ca27d6447a60d3ffa27e

Observation 2f46af81-460d-4109-9a0e-016a452736b2 · outbound

This paper cites Qwen2 Technical Report.

DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.881839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.881839Z digest=sha256:5f82bb199c8772a9293aa062506a6be3f426e46c3bb155a34b447138f1d9660d

Observation e472e1d7-8b69-4fb1-8f6c-ecbb9889acab · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

DOGR: Towards Versatile Visual Document Grounding and Referring Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.340509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.886664Z digest=sha256:30a2e6ec31bd8108b8543ec5cbd5d3644940a531d58cdba708b5557d8925347b

Observation 1852f1e3-5581-4968-921f-5172e00bd02c · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

DOGR: Towards Versatile Visual Document Grounding and Referring Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.890922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.890922Z digest=sha256:b8e86ac767332cd79d0c04f35cce7667e6871d04d968bbdebb4345d088511ad0

Observation 84c914e4-b932-470c-b125-8178f240c966 · outbound

This paper cites Syntax-Aware Network for Handwritten Mathematical Expression Recognition.

DOGR: Towards Versatile Visual Document Grounding and Referring Syntax-Aware Network for Handwritten Mathematical Expression Recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.895427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.895427Z digest=sha256:cf86c6ebd399a217555ada3c5d59b645f00781b35832e09ea71acce559355d1f

Observation a68251d5-6c94-4426-9053-cb39539666b5 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

DOGR: Towards Versatile Visual Document Grounding and Referring MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.900255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.900255Z digest=sha256:1f1b3aa2829af74c03e7250be29e4e3269eb2645a47d0238ec74060c8905c0e0

Observation 9e1a1a81-6f22-4224-ab42-dcc186d4266d · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

DOGR: Towards Versatile Visual Document Grounding and Referring Llava-grounding: Grounded visual chat with large multimodal models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.326936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.904576Z digest=sha256:82123419a8b0a71369a36c6b69ac470cbaf9c6ae3abfc8389a7eb2e67c86d1fd

Observation 85c0c9ca-6ee4-471d-8341-845c55f471cf · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DOGR: Towards Versatile Visual Document Grounding and Referring InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.909206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.909206Z digest=sha256:c5dc0067652afcfe3d93261992fbaf6d8e02dc1bf8ea3d16f639515246daf4ce

Observation 430b7b93-879e-45fb-ab31-b9193ec284fa · outbound

This paper cites RobuT: A systematic study of table QA robustness against human-annotated adversarial perturbations.

DOGR: Towards Versatile Visual Document Grounding and Referring RobuT: A systematic study of table QA robustness against human-annotated adversarial perturbations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.314568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T12:34:31.913697Z digest=sha256:ee320459d8aca495f381b3e7c6d97440fd7603d19cd79e41f1d4303069f663b2

Observation 75aee65c-07e5-4e9d-900c-654db90d476e · outbound

This paper cites Scale Up Composed Image Retrieval Learning via Modification Text Generation.

DOGR: Towards Versatile Visual Document Grounding and Referring Scale Up Composed Image Retrieval Learning via Modification Text Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.918078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.918078Z digest=sha256:975604e8da02e7221243b55dda73f7b895928a26dc24840b49542e29cb3d7b04

Pith citing papers

Observation a83651ba-fafa-43c7-916a-56f4d0b78fbb · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:11.055077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:11.055077Z digest=sha256:ff0d047bcf98b0f4332565c8c626a8d81f9350d2802dffc3ffd161a2e29e47ba

Observation accb5211-bd9b-4273-870a-f2ef7a501548 · inbound

DocVXQA: Context-Aware Visual Explanations for Document Question Answering cites this paper.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.593869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.593869Z digest=sha256:a04b49b261259495482dd0f485b1c2e1a8caf1ee23248247bc2e13b4cb0a575a

Observation df04bb59-303a-4b86-b1d3-0706cfdbad84 · inbound

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning cites this paper.

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:40.427096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:40.427096Z digest=sha256:2832af40b114144fe4f050589238fa190c3a84e0275a594b5313630b6bbbafd2

Observation 83781985-7053-40ae-a698-8e9887e53953 · inbound

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents cites this paper.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:38.340134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:38.340134Z digest=sha256:17ca447548732a7299f69e5a1e5ceccd6c2c96a55194f2465402f2a5499dc018

Observation 626704e1-01be-4ffa-8bbd-5cea8d2f4de8 · inbound

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring cites this paper.

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:56.645771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:48:09.517875Z digest=sha256:30ce26776244ed5c010b1e0b49988c44eace2550b3642429288b869189b50781

Observation f816d463-e89f-46cf-b5f2-8a7105fdf6d0 · inbound

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring cites this paper.

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.389489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:31:19.715414Z digest=sha256:665f0c47ecc718ef85988e63bf41668fe33a9467d47e39b8c4fc54094b4e3219