Pith. sign in

Paper Citation Record · LEDGER

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding

As of 21 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2505.05446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05446 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:08:52.342090Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44de7fcd-6490-404c-aa00-b3099dce6f5a · outbound

This paper cites Au- tomaTikZ: Text-guided synthesis of scientific vector graph- ics with TikZ.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Au- tomaTikZ: Text-guided synthesis of scientific vector graph- ics with TikZ

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.449902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.360963Z digest=sha256:3601798d3a997648e6c90d633c2f5c6d6dbe701910936c9b72fe8dfa7d9bd63a

Observation b4f7802e-dcef-4357-babd-25f8bb87b9ce · outbound

This paper cites DeTikZify: Synthesizing graphics programs for scientific figures and sketches with TikZ.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding DeTikZify: Synthesizing graphics programs for scientific figures and sketches with TikZ

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.436810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.444090Z digest=sha256:542b1a8b492aa2b6dd48e99daa56c705952a4d804e3a5d8a3c61a7af3892184e

Observation a04f5739-81b5-4c47-a927-96adfbdcfaff · outbound

This paper cites Scene text visual question answering.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Scene text visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.332181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.485954Z digest=sha256:fa71ad8a2bdf534d1517549160eee840318fc5a211cbc671cd50362d0ec7713e

Observation 941fc196-2803-45a1-9ab8-cae16487696a · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.491089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.491089Z digest=sha256:53a0d5bbb771509545c2f8e36dd261b6a71751e0f1acad59a82e0acff67f28fb

Observation c90c241d-7996-4298-acb4-4583b69951b8 · outbound

This paper cites Onechart: Purify the chart structural extrac- tion via one auxiliary token.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Onechart: Purify the chart structural extrac- tion via one auxiliary token

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.318027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.496304Z digest=sha256:fc78786ff45043e6a4a24567976025542f6f1802ef847fbfa490de589a39eb04

Observation 2c1ff4e6-b7b4-4ae7-8400-3044113cfc7a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.501111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.501111Z digest=sha256:75c652bef57bb9674ab890f279cfa180055d5d41048b2ad52d1441d6d69917aa

Observation f00b2977-ba1c-4df6-a042-a5141caa90f3 · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.505215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.505215Z digest=sha256:ca2d378bd426368d30144686f28d0b66049de1c564895ded1cb7f2bfabb7bc87

Observation 6e24ccaf-f037-40a1-a15e-e4c4903d3014 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.510896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.510896Z digest=sha256:d25ac06f44040b02d65d255b5a60e66194b8327ddc5aedbb1dad3b2ca120acf8

Observation 84a9a66b-0bb7-4def-9a8d-d123e137dccf · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.304245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.567479Z digest=sha256:5567ab81251803fee469dcc396e3bc94d0a6e9798d08ffbd688be9bab79222cc

Observation 3a87248b-98d1-4384-9a5d-f1bfb82ca3bc · outbound

This paper cites Complicated Table Structure Recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Complicated Table Structure Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.588436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.588436Z digest=sha256:876e92d8325a5590deb5b9fd92c4ea5d652ce0852c9ee9440bbdc6b61a0f64ca

Observation 93c99fa7-5177-4118-a41f-7018fc45724c · outbound

This paper cites DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.701331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.701331Z digest=sha256:6e1dd1170dc7448b55201bdae5a4ffe0d5d5b06122f7d1cbaf2993da7df64723

Observation 9ccfac39-359c-43f5-ac9d-c6d97d994414 · outbound

This paper cites G-llava: Solving geomet- ric problem with multi-modal large language model, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding G-llava: Solving geomet- ric problem with multi-modal large language model, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.247784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.740427Z digest=sha256:d50d3b48805f045fbcea2e9123672f8727c4a1f527303cfb75c0ed6008cfc1af

Observation acf0312e-3aea-4063-a99e-eb6897865239 · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.745849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.745849Z digest=sha256:197e0169739359cef1ca80838faa83cc59b9dc8441a21f6774fb1a2c59912767

Observation 9ae419b8-c3a0-4a49-bc84-717668730182 · outbound

This paper cites MathWriting: A Dataset For Handwritten Mathematical Expression Recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding MathWriting: A Dataset For Handwritten Mathematical Expression Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.751241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.751241Z digest=sha256:cd26ccab5e7dc9dda4ac70d78ffeda0d7ecbd107922bff85c8c5132749967ef9

Observation 9469f8fe-955f-4044-a005-7086ea73b5cb · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding ImageBind-LLM: Multi-modality Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.755574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.755574Z digest=sha256:313b1c0252e49056704e58e7b683148771d9a5e6462a9d3bcd6f71a5406264d0

Observation 0882036b-164f-4497-b06a-85f854dc7042 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Cogagent: A visual language model for gui agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.233788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.759741Z digest=sha256:6ab5f9beb1611c36b7d7fbd0314d8d77c03a4ea2233d8ac5eccaee9190c82b89

Observation 423749f3-1393-4285-a447-57066666117c · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.763889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.763889Z digest=sha256:9224b4d93af8dd9d1470c45d359205437371010a2985e1cab135015d80fe2fa8

Observation 35a35a60-4480-43da-9239-1340ba8b725d · outbound

This paper cites Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.218188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.864930Z digest=sha256:cbe2f0b4fbeef8cd4892caf4716edde3093f9adbee66f253948e18e560306491

Observation 537c7fa5-3bf6-4da7-9eeb-275ef31a95b0 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Funsd: A dataset for form understanding in noisy scanned documents

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.178053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.869149Z digest=sha256:8c65bc91894ec01831be051465053e80f42c2c7d4f192ffe6b314e81d214a0b7

Observation 8fca41c6-d9be-48c6-b7b3-ed7f004ba0ca · outbound

This paper cites Revisiting scene text recognition: A data per- spective.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Revisiting scene text recognition: A data per- spective

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.131556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.873974Z digest=sha256:0872d68e0e45ce35e273b514c34f9ae868d0f7026aa18f14bfbaab524f3188af

Observation dbcb6e44-73f9-4d68-9a63-2ceebab4812a · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Dvqa: Understanding data visualizations via ques- tion answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.118741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.878908Z digest=sha256:36b0962d2ba0280d06c2c6abd981616fa3177c6f0dfaf77023a7c284f0bdcffd

Observation 7adfcb41-ad6f-4f8b-9ac8-af84783e55d8 · outbound

This paper cites Icdar 2013 robust read- ing competition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2013 robust read- ing competition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.105042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.884108Z digest=sha256:66625021c7a2820059374fc91071696fa29f796cd066bff23c815a4433131258

Observation c963a7f8-e702-4c4c-a539-633f4b09a945 · outbound

This paper cites Icdar 2015 competition on robust reading.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2015 competition on robust reading

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.090467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.887834Z digest=sha256:4b5525e25940f1e44868077ab8de00f5d2573285351806c56861c3fad3be0a99

Observation 0aa04442-fd9a-4ecb-93be-248b41e0fd04 · outbound

This paper cites A diagram is worth a dozen images.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding A diagram is worth a dozen images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.892593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.892593Z digest=sha256:2ffe2d0c2182d09f6312d04ad39d573b9de23cda8431cd817de9f3cad725e8b5

Observation ef09898d-750d-46ca-9d18-02d7410f8be8 · outbound

This paper cites Openassistant conversations-democratizing large lan- guage model alignment.Advances in Neural Information Processing Systems, 36, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Openassistant conversations-democratizing large lan- guage model alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.039697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:50.912468Z digest=sha256:2d5f730e5c6c7dcaacad4289bad4112090dd7875bc82b630072a2f54562208df

Observation 0b92f8a9-5590-471d-b182-e0510bdb4375 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solu- tion.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visual information extraction in the wild: practical dataset and end-to-end solu- tion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.026666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.001750Z digest=sha256:21d0e7e3f4788acacba2ff4b30eca18e9854b87f371db51bc520696af9f26945

Observation eacbb7c1-5e47-4b45-9a7b-dbc794caf5b4 · outbound

This paper cites Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.006221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.006221Z digest=sha256:09f6e62f2b4827e7c2ca71641274f2039e23bda540dc1d82d325130c92cec138

Observation 3740c690-0340-41b4-bba2-d560e50f4f95 · outbound

This paper cites Docmatix dataset.https://huggingface.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Docmatix dataset.https://huggingface

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.991853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.010835Z digest=sha256:1b7ea409f5dca157e5063ad27725b992972a531c48543870ac4df797367340a4

Observation 0a090194-7485-4d18-93ae-fbe83f66bed8 · outbound

This paper cites When counting meets hmer: counting-aware network for handwritten math- ematical expression recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding When counting meets hmer: counting-aware network for handwritten math- ematical expression recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.958651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.015778Z digest=sha256:85701b7c485f524b7f2c8a6faf52589f5fac8b74a023be941c59e39221443733

Observation fc9495b3-c0bf-4d05-b34f-b9b57932b061 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.020428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.020428Z digest=sha256:51c2187660dd73b6118f5a95e40622cb4bdc5e13f9771e25655800c5d6854129

Observation 87faa7b8-5261-4eda-ad00-a04cb3e8ab59 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.024767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.024767Z digest=sha256:94e0f59e6d893b807b19392bc8dd88e5e13214193a293e3f66f3d1b2f7fb73fa

Observation 56e408e2-84b5-4f09-bf76-53378ba7cd81 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.057587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.057587Z digest=sha256:6667e49c1f7878e3d091bf2f059936c4004a1427cd35463c3f82df44f3320e39

Observation f3c3310f-88d5-47a6-accf-79aaf5186120 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.943284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.107858Z digest=sha256:51a3aff1afdd139e83c1fa8850ae7419ef270c1f1e5c28abc3ef16bed9f2b018

Observation 16cdadcf-dcaa-4ee2-8d23-365128d973fa · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.128928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.128928Z digest=sha256:5bfa5b01ec4394bb12462608a50994c36e918466b08c950e005f4fd8b5d4cb1c

Observation 5f31cc37-4a91-4ef2-8cec-9edded724a85 · outbound

This paper cites Visual Instruction Tuning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.133780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.133780Z digest=sha256:7456115e851f29188f92b8c824619bc6442985b2ae4edabb7c92c82c56f9b1d6

Observation a4f4326a-ca18-4762-b33c-7be770ea2b88 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.186321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.186321Z digest=sha256:144c5624de15d5989d82f887912aea573af90d05bcd4433fe4998c9011da8fc5

Observation c3850ba8-bba0-4961-96f5-8c67cdf312e2 · outbound

This paper cites Visualwebbench: How far have multimodal llms evolved in web page under- standing and grounding?, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visualwebbench: How far have multimodal llms evolved in web page under- standing and grounding?, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.907864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.224727Z digest=sha256:c80c845790081d19e7939dd31013c3083b0eb1404f0fe8df4c0b6341beff078d

Observation 22bdc781-3999-4107-9311-c625b02c7084 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.229797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.229797Z digest=sha256:6715727506674062aeb9a69fbd71e3633887f0615f3eb89a65c1e84d490ac3ba

Observation 3d98969a-46df-4552-8f3e-845d90a7d4ac · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.234125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.234125Z digest=sha256:3a0059046a4ac5f6cf4afa9cf0380c20417b04248970739a9924d9b04af1c57b

Observation f0e294dc-0f64-4f08-8c8f-492a9028c9f1 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Towards end-to-end unified scene text detection and layout analysis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.880831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.238255Z digest=sha256:64e08e4584d6d60448d04a01e7d6a52ccda327ded943146d18fb926c05b654b6

Observation 26c5c5f3-e426-4c74-9b09-b033ce70db73 · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.851414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.242047Z digest=sha256:3dcf4f2ae64d446cf43a58980abd88d1160715d6254678eb6253d94f0b882843

Observation e38818c9-4342-4872-9d5d-88f6529e1f4a · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.818038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.352094Z digest=sha256:5b3aad7786acdbf93f321ad50b587443a8aa3d812b5d211a5caf859b463c3d50

Observation f2d84bb7-0985-4d1f-82e4-af39afe7b3db · outbound

This paper cites V Jawahar.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding V Jawahar

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.803621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.431640Z digest=sha256:392702bcebfd17a8d9ce3782e8fc76405e39d1a21fb09b0db325f3cb0e8be7a0

Observation fab00ed6-d167-49a8-b0b8-87ff19e275a0 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Docvqa: A dataset for vqa on document images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.732117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.435477Z digest=sha256:692919c324fb37a884fec572f1945e0b7e8bee3b0175c71ea24e3b4caac9f84c

Observation c58d1bd8-7718-44c6-a068-62ea7207e501 · outbound

This paper cites Plotqa: Reasoning over scientific plots.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Plotqa: Reasoning over scientific plots

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.439981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.439981Z digest=sha256:6a3641e2906b26cbc5a7861e694f7078cc8dcf47274d222d0ccdc6f62c84cb7d

Observation 5a30488f-ad31-4d6c-9b82-c12b060e3bbe · outbound

This paper cites TableFormer: Table Structure Understanding with Transformers.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding TableFormer: Table Structure Understanding with Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.444111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.444111Z digest=sha256:c19ca172648f49235ae7cd1dc9448f08f515df7fde5695fde977d85218662137

Observation 841dcd17-c17b-4980-b4f1-ec88b48646cb · outbound

This paper cites Chatgpt.https://chat.openai.com, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Chatgpt.https://chat.openai.com, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.708357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.448954Z digest=sha256:6fbd151be657e639c5b00e1db6226d822e1ad80fd5b7298f759b4c4c81364664

Observation 162f744d-64d1-433b-bedf-63d2467bddd6 · outbound

This paper cites GPT-4 Technical Report.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.453832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.453832Z digest=sha256:99049fb1c9aec589bf2e9f555f8c040826f0d24760a79c537561f8ed551310f2

Observation 5c5577d2-9d1e-4491-b448-e12047bacfd2 · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Training lan- guage models to follow instructions with human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.673222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.503220Z digest=sha256:719f513be2d3b92a1f7501af91a98e02a12bb3b1ea8cea5f65043a8422e53929

Observation 6d581baf-0a6a-4422-ab2f-232eebd5f580 · outbound

This paper cites Cord: a con- solidated receipt dataset for post-ocr parsing.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Cord: a con- solidated receipt dataset for post-ocr parsing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.644893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.539927Z digest=sha256:674e8b5731ed1730216f2af4b8197fcd9cbfb5467d9cf0b78bc7acef74de9e27

Observation 9783b180-e7fe-4fc2-a1cc-867a29e12868 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.545097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.545097Z digest=sha256:60bb4b88b9a995fa3bca890cf8c297ad8481f50df6cfbe5fc9e1ab03a3080b56

Observation 9d718dc0-ddd2-4652-bd37-83cac6334546 · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.611521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.627950Z digest=sha256:f71b02f06acfea813c555b23529040f4270d362e94949cfaff90103b296391bf

Observation 58d14f95-aa7a-45f8-bb40-7e2f252b0480 · outbound

This paper cites Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models.arXiv e-prints, pages arXiv–2403, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models.arXiv e-prints, pages arXiv–2403, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.595885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.748557Z digest=sha256:0330d4ea5bad7b7577c5b76e55ea65bde8b3bac3c703c88b103bff973ba0dd90

Observation 2a6bd541-baf3-48fd-b6c0-6323a2ff5c15 · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17).

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar2017 competition on reading chinese text in the wild (rctw-17)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.523375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.795572Z digest=sha256:bb9f2d172c40b9bbca311182405cb79b4268a243c277ef1303f99d680c46cb54

Observation 00ce7920-1e11-49b4-ba66-e0ea3c37dcd8 · outbound

This paper cites Towards vqa models that can read.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Towards vqa models that can read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.505733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.800403Z digest=sha256:ee706329ae2a40b25727d062c0cb369a8aae7c052d19da249f6aaeba50d78e84

Observation 81387647-952a-479e-867d-044615ede585 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.473925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.805429Z digest=sha256:679ebf53a895ccce1dc91e5cdc89c297a6097d580c9d3d35b7dbfaeaf32ee298

Observation f64efca0-3d6f-4ff7-aca3-99c6fa340b25 · outbound

This paper cites Spatial Dual-Modality Graph Reasoning for Key Information Extraction.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Spatial Dual-Modality Graph Reasoning for Key Information Extraction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.809034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.809034Z digest=sha256:6069294873a15072ebe3cc682ce8b1d6498d03e17af4070fea515e5a8ddc6314

Observation 7ee44980-8101-4545-9080-8d4175fe7c0f · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.459557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.813894Z digest=sha256:b8cb208efa9669ba929b3c99946557dbb17b929df3eeb9814801043e233673e9

Observation 35ffffdd-eff8-4d0f-8ba9-996962eda127 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.818507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.818507Z digest=sha256:afb3ea78c1fe9886717461213cfdf79b84d9bf8f2bb244446c3ffc9a2a78cc93

Observation cac7b022-3036-4974-babc-824a47373cf1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.878962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.878962Z digest=sha256:997ce31e57122bae4a79d3e7db061a3fcbf1febcd8a2c46914636bdf3471389a

Observation c2a191ff-1344-40a5-a6b8-ca3f7b1a630a · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.969524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.969524Z digest=sha256:212e7ccbf07036e6d581ec0417d3868018397d6fe602771e9508169ee5b5cc6d

Observation 2f7fb109-156f-4f24-b39a-101539bb3ab4 · outbound

This paper cites Measuring multimodal mathemat- ical reasoning with math-vision dataset, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Measuring multimodal mathemat- ical reasoning with math-vision dataset, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.356298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.974634Z digest=sha256:f5255c4c7c1cdd4c148b8b34dc828ec907eeebaa3e0f3ce0867dd958c5bf5bea

Observation 7a1efaed-b0d0-4f4d-b5ad-a3abd983b7ea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.979528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.979528Z digest=sha256:5fb04790f718d17778020ced1d526e17f75975237875705ff040e83d94e72c96

Observation 2c687a36-62e2-4d01-a1bd-69b66d8fe781 · outbound

This paper cites On the general value of ev- idence, and bilingual scene-text visual question answering.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding On the general value of ev- idence, and bilingual scene-text visual question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.343783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.984044Z digest=sha256:df64fcf0b12d7ce8c1dae75e8f10eda6b97f235b91c711fb6ebb452976a82cda

Observation 5dcf2f51-9e6b-4896-8c6b-a2070646d065 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision-language model.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Vary: Scaling up the vision vocabulary for large vision-language model

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.276269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:51.988895Z digest=sha256:97ab7dc78e442a9ee8b90f9d5e4d83651dc92c45a05894f887e008fa16a84482

Observation cf6d6f5b-fe80-4f4f-a99c-2df0ec142db4 · outbound

This paper cites Toward understanding wordart: Corner-guided transformer for scene text recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Toward understanding wordart: Corner-guided transformer for scene text recognition

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.259690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.076260Z digest=sha256:32df02043f39c65b8d5724719ffed08950609cd6199182af50119b111c0e4bf4

Observation ab7e3ebd-126c-4a2c-be6e-4037f1143f18 · outbound

This paper cites Xfund: a benchmark dataset for multilingual visually rich form under- standing.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Xfund: a benchmark dataset for multilingual visually rich form under- standing

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.243919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.080316Z digest=sha256:38739cf4b7d3127ba8c609e78f6ef36527cf110997b9804cfda269694643e3fb

Observation 1e2dd512-7d28-422b-9972-c10ec4d8d51b · outbound

This paper cites Tgrnet: A table graph reconstruction net- work for table structure recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Tgrnet: A table graph reconstruction net- work for table structure recognition

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.141050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.084510Z digest=sha256:1ba238c127b8d2445ac0976168d63c532a653b3c49b655027aa5c8b9c9e67fe9

Observation 1ef12314-b89f-4532-8fcf-fbde5f03a3d6 · outbound

This paper cites A large-scale dataset for end-to-end table recognition in the wild.Scientific Data, 10(1):110, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding A large-scale dataset for end-to-end table recognition in the wild.Scientific Data, 10(1):110, 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.127344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.088527Z digest=sha256:938a6be9385c8b5fe049d536ad3f80efa3dc15cbbb558dcb5eb5405097926fa4

Observation 95a8e5eb-970f-4e92-ad66-8d2e24d4f79a · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.092596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.092596Z digest=sha256:034884b26f1967faecb6ebc8f668f72a1256971c494c05c420c196eb252f3e87

Observation 102e7b8a-ffbc-4d6e-be24-f0867ae3c2c6 · outbound

This paper cites Icdar 2023 competition on structured text extraction from visually-rich document im- ages.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2023 competition on structured text extraction from visually-rich document im- ages

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.987500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.097699Z digest=sha256:867130b78179a8a7d39c64a506bfc5f56e2c24c23de5ec91ae18d5f18c279f81

Observation 555b5dd8-002c-4868-ac4b-d776363d8406 · outbound

This paper cites Syntax-aware network for handwritten mathematical expression recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Syntax-aware network for handwritten mathematical expression recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.974402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.101990Z digest=sha256:8e26941c9d7df195297eb06b3912922267a34ecff49c9b2869bf7c3311081e26

Observation 11aefa47-b1fa-445e-bf97-aa85ada1a6b7 · outbound

This paper cites Detecting Curve Text in the Wild: New Dataset and New Solution.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Detecting Curve Text in the Wild: New Dataset and New Solution

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.197635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.197635Z digest=sha256:7c51f7a47b23fea96de89c8ea41a1757cb030bb42708ee7ff1eacb5ed3407168

Observation 97b5a255-71c8-4c62-9a18-46fd0f6bdf04 · outbound

This paper cites LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.203082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.203082Z digest=sha256:0d4162a6a2baa7ac76f827b8f6ec7e55c853e2e271c70a995f29857005998dd2

Observation 3f52d260-71b8-4827-8cbb-3806f05b94e7 · outbound

This paper cites Image-based table recognition: data, model, and evaluation.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Image-based table recognition: data, model, and evaluation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.207933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.207933Z digest=sha256:536567ad9a7bbb1b94d5b4f914b984bc84e41597ced9d638e0335b522212045c

Observation c5ee0e39-02ff-4ed2-b880-fa60c8a1a524 · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.881171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.211898Z digest=sha256:09f61a51d323f8fc3c4a53e4207ab6f356a9e68aa77d898b8c670601582a49da

Observation dae84381-e88b-4ee2-bb1a-8dd3f0c4e0ba · outbound

This paper cites - Answer: The known answer to the question.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding - Answer: The known answer to the question

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.866566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.216345Z digest=sha256:d6c2902e258f5d11918e07e35b309528d159e0e01fb84e403f447356a0642bcd

Observation 50b3169c-79a7-4055-905b-18333b55a3e7 · outbound

This paper cites - `<txt_gd></txt_gd>`: Text with coordinates for context.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding - `<txt_gd></txt_gd>`: Text with coordinates for context

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.779914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.328887Z digest=sha256:958f59f949e34907f802239866d8d2643e60e75e79eac918210262d625745093

Observation eaf472a1-940f-4aa9-8f29-aac0b996b54e · outbound

This paper cites - Ensure that the extracted content retains its original formatting.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding - Ensure that the extracted content retains its original formatting

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.766246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.337274Z digest=sha256:2abe72405ad4163f373b36f1868ed9b53f68525ba3edc388accc86666876627f

Observation 481d3b8f-4f52-41e7-860a-0a68407ee8f3 · outbound

This paper cites title":.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding title":

Reference 80

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T23:08:52.752277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:08:52.342090Z digest=sha256:ce5f3b4917362b65d1d4b20927ebe074d750c06677f5c25a23568b8505732613

Pith citing papers

No inbound Pith citation observations are available.