Pith. sign in

Paper Citation Record · LEDGER

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding

As of 21 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2505.05446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05446 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:08:52.342090Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44de7fcd-6490-404c-aa00-b3099dce6f5a · outbound

This paper cites Au- tomaTikZ: Text-guided synthesis of scientific vector graph- ics with TikZ.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Au- tomaTikZ: Text-guided synthesis of scientific vector graph- ics with TikZ

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.449902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.360963Z digest=sha256:80032f771f67dbf28667b41b2cb6dbbf0d5796d18ed841338909b4c3cf053552

Observation b4f7802e-dcef-4357-babd-25f8bb87b9ce · outbound

This paper cites DeTikZify: Synthesizing graphics programs for scientific figures and sketches with TikZ.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding DeTikZify: Synthesizing graphics programs for scientific figures and sketches with TikZ

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.436810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.444090Z digest=sha256:207fc0a88ddfb3db2d49fb9fed4ac02f5ff1dbd365f301029bb02b3ab1cb538c

Observation a04f5739-81b5-4c47-a927-96adfbdcfaff · outbound

This paper cites Scene text visual question answering.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Scene text visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.332181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.485954Z digest=sha256:0bdd56f70055972c62f2193f7ae46b0dce05d79ba5f8ad2db085b16236541ff7

Observation 941fc196-2803-45a1-9ab8-cae16487696a · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.491089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.491089Z digest=sha256:53a0d5bbb771509545c2f8e36dd261b6a71751e0f1acad59a82e0acff67f28fb

Observation c90c241d-7996-4298-acb4-4583b69951b8 · outbound

This paper cites Onechart: Purify the chart structural extrac- tion via one auxiliary token.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Onechart: Purify the chart structural extrac- tion via one auxiliary token

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.318027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.496304Z digest=sha256:23cae0d0918ac6f59f4c53e6e6f1ca00810b3a045a1fe777bffc1be65359a7ee

Observation 2c1ff4e6-b7b4-4ae7-8400-3044113cfc7a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.501111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.501111Z digest=sha256:75c652bef57bb9674ab890f279cfa180055d5d41048b2ad52d1441d6d69917aa

Observation f00b2977-ba1c-4df6-a042-a5141caa90f3 · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.505215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.505215Z digest=sha256:ca2d378bd426368d30144686f28d0b66049de1c564895ded1cb7f2bfabb7bc87

Observation 6e24ccaf-f037-40a1-a15e-e4c4903d3014 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.510896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.510896Z digest=sha256:d25ac06f44040b02d65d255b5a60e66194b8327ddc5aedbb1dad3b2ca120acf8

Observation 84a9a66b-0bb7-4def-9a8d-d123e137dccf · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.304245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.567479Z digest=sha256:8617b47c941bfb85c304d392dba80567cac4e9f4cf16a67b316750b1db16f7e5

Observation 3a87248b-98d1-4384-9a5d-f1bfb82ca3bc · outbound

This paper cites Complicated Table Structure Recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Complicated Table Structure Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.588436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.588436Z digest=sha256:876e92d8325a5590deb5b9fd92c4ea5d652ce0852c9ee9440bbdc6b61a0f64ca

Observation 93c99fa7-5177-4118-a41f-7018fc45724c · outbound

This paper cites DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.701331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.701331Z digest=sha256:6e1dd1170dc7448b55201bdae5a4ffe0d5d5b06122f7d1cbaf2993da7df64723

Observation 9ccfac39-359c-43f5-ac9d-c6d97d994414 · outbound

This paper cites G-llava: Solving geomet- ric problem with multi-modal large language model, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding G-llava: Solving geomet- ric problem with multi-modal large language model, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.247784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.740427Z digest=sha256:215299476de74dda6dbfd6c17c6cf1c1f10522fd3dd98dbfafbc3461cf836169

Observation acf0312e-3aea-4063-a99e-eb6897865239 · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.745849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.745849Z digest=sha256:197e0169739359cef1ca80838faa83cc59b9dc8441a21f6774fb1a2c59912767

Observation 9ae419b8-c3a0-4a49-bc84-717668730182 · outbound

This paper cites MathWriting: A Dataset For Handwritten Mathematical Expression Recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding MathWriting: A Dataset For Handwritten Mathematical Expression Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.751241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.751241Z digest=sha256:cd26ccab5e7dc9dda4ac70d78ffeda0d7ecbd107922bff85c8c5132749967ef9

Observation 9469f8fe-955f-4044-a005-7086ea73b5cb · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding ImageBind-LLM: Multi-modality Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.755574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.755574Z digest=sha256:313b1c0252e49056704e58e7b683148771d9a5e6462a9d3bcd6f71a5406264d0

Observation 0882036b-164f-4497-b06a-85f854dc7042 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Cogagent: A visual language model for gui agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.233788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.759741Z digest=sha256:785f0d293d95709c63087bbb6beab1f51c65651efd9a99ee131228ecf9e03e73

Observation 423749f3-1393-4285-a447-57066666117c · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.763889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.763889Z digest=sha256:9224b4d93af8dd9d1470c45d359205437371010a2985e1cab135015d80fe2fa8

Observation 35a35a60-4480-43da-9239-1340ba8b725d · outbound

This paper cites Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2019 robust reading challenge on scanned receipts ocr and information extraction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.218188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.864930Z digest=sha256:e631f3373f2556db09ef0265effbf72fb8bc758446be7c11c1d51eca6b6cb7c4

Observation 537c7fa5-3bf6-4da7-9eeb-275ef31a95b0 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Funsd: A dataset for form understanding in noisy scanned documents

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.178053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.869149Z digest=sha256:99b6280e3093ca9f7249b5f2a1ff24a1450add19f705883778166f8bd344154c

Observation 8fca41c6-d9be-48c6-b7b3-ed7f004ba0ca · outbound

This paper cites Revisiting scene text recognition: A data per- spective.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Revisiting scene text recognition: A data per- spective

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.131556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.873974Z digest=sha256:365755478e23e3e5d80e0121c31691eeb43442df6520860f9433281615de50ee

Observation dbcb6e44-73f9-4d68-9a63-2ceebab4812a · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Dvqa: Understanding data visualizations via ques- tion answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.118741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.878908Z digest=sha256:955677dda713a204879147ae67f97365ba778fc988ef1c97759126bbbcedcef8

Observation 7adfcb41-ad6f-4f8b-9ac8-af84783e55d8 · outbound

This paper cites Icdar 2013 robust read- ing competition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2013 robust read- ing competition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.105042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.884108Z digest=sha256:7dea77f6cf255d75e9ef4b0c9adb4a91c41aa349e1f590c55ea10f55004252a3

Observation c963a7f8-e702-4c4c-a539-633f4b09a945 · outbound

This paper cites Icdar 2015 competition on robust reading.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2015 competition on robust reading

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.090467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.887834Z digest=sha256:5f5b3d6253e2c28ac5a80f2cdf7e1999a3606a0b1183118e18a6b4a1bda1f6b2

Observation 0aa04442-fd9a-4ecb-93be-248b41e0fd04 · outbound

This paper cites A diagram is worth a dozen images.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding A diagram is worth a dozen images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:50.892593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:50.892593Z digest=sha256:2ffe2d0c2182d09f6312d04ad39d573b9de23cda8431cd817de9f3cad725e8b5

Observation ef09898d-750d-46ca-9d18-02d7410f8be8 · outbound

This paper cites Openassistant conversations-democratizing large lan- guage model alignment.Advances in Neural Information Processing Systems, 36, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Openassistant conversations-democratizing large lan- guage model alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.039697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:50.912468Z digest=sha256:3f97ca44b82f65dd37e79bbae3fe2af64b9bfe794a079610a828dcef8b1d2b19

Observation 0b92f8a9-5590-471d-b182-e0510bdb4375 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solu- tion.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visual information extraction in the wild: practical dataset and end-to-end solu- tion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:54.026666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.001750Z digest=sha256:1d2c6254e0841a0bcb1f681e703e51682dab139dd7b73580a8475fe1692b5199

Observation eacbb7c1-5e47-4b45-9a7b-dbc794caf5b4 · outbound

This paper cites Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.006221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.006221Z digest=sha256:09f6e62f2b4827e7c2ca71641274f2039e23bda540dc1d82d325130c92cec138

Observation 3740c690-0340-41b4-bba2-d560e50f4f95 · outbound

This paper cites Docmatix dataset.https://huggingface.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Docmatix dataset.https://huggingface

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.991853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.010835Z digest=sha256:15cfb355181d5fea72a42567d2d99046061e96555e7ff3c0b7bc391a6b869496

Observation 0a090194-7485-4d18-93ae-fbe83f66bed8 · outbound

This paper cites When counting meets hmer: counting-aware network for handwritten math- ematical expression recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding When counting meets hmer: counting-aware network for handwritten math- ematical expression recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.958651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.015778Z digest=sha256:e69febc84d525399f569f7d370e5d3e9027a21122aee0c036c2eff38ff5ff14e

Observation fc9495b3-c0bf-4d05-b34f-b9b57932b061 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.020428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.020428Z digest=sha256:51c2187660dd73b6118f5a95e40622cb4bdc5e13f9771e25655800c5d6854129

Observation 87faa7b8-5261-4eda-ad00-a04cb3e8ab59 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.024767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.024767Z digest=sha256:94e0f59e6d893b807b19392bc8dd88e5e13214193a293e3f66f3d1b2f7fb73fa

Observation 56e408e2-84b5-4f09-bf76-53378ba7cd81 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.057587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.057587Z digest=sha256:6667e49c1f7878e3d091bf2f059936c4004a1427cd35463c3f82df44f3320e39

Observation f3c3310f-88d5-47a6-accf-79aaf5186120 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.943284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.107858Z digest=sha256:d0977d13d67b40b8c4f4ff2ff9d0e41cfb5bc1463966bd26b54e9ac19457c1f8

Observation 16cdadcf-dcaa-4ee2-8d23-365128d973fa · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.128928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.128928Z digest=sha256:5bfa5b01ec4394bb12462608a50994c36e918466b08c950e005f4fd8b5d4cb1c

Observation 5f31cc37-4a91-4ef2-8cec-9edded724a85 · outbound

This paper cites Visual Instruction Tuning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.133780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.133780Z digest=sha256:7456115e851f29188f92b8c824619bc6442985b2ae4edabb7c92c82c56f9b1d6

Observation a4f4326a-ca18-4762-b33c-7be770ea2b88 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.186321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.186321Z digest=sha256:144c5624de15d5989d82f887912aea573af90d05bcd4433fe4998c9011da8fc5

Observation c3850ba8-bba0-4961-96f5-8c67cdf312e2 · outbound

This paper cites Visualwebbench: How far have multimodal llms evolved in web page under- standing and grounding?, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visualwebbench: How far have multimodal llms evolved in web page under- standing and grounding?, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.907864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.224727Z digest=sha256:a2f832d9a22cd08ef4fd9b629f083df89fcc5ba565aba3268c9ea007c3cee7ad

Observation 22bdc781-3999-4107-9311-c625b02c7084 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.229797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.229797Z digest=sha256:6715727506674062aeb9a69fbd71e3633887f0615f3eb89a65c1e84d490ac3ba

Observation 3d98969a-46df-4552-8f3e-845d90a7d4ac · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.234125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.234125Z digest=sha256:3a0059046a4ac5f6cf4afa9cf0380c20417b04248970739a9924d9b04af1c57b

Observation f0e294dc-0f64-4f08-8c8f-492a9028c9f1 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Towards end-to-end unified scene text detection and layout analysis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.880831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.238255Z digest=sha256:6c1ac55e4e80822a67a7413a15ec9a4028bc949d943e65bbe6dc6656da724885

Observation 26c5c5f3-e426-4c74-9b09-b033ce70db73 · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.851414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.242047Z digest=sha256:8a3677260471103b83bb8843ca8ab57055d5a561ab75218c0eda79b67d9fdf87

Observation e38818c9-4342-4872-9d5d-88f6529e1f4a · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.818038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.352094Z digest=sha256:47edc187490442ed57389eab72b25c258f5c2e190ab3b2c6917bdfa0ac39ddf9

Observation f2d84bb7-0985-4d1f-82e4-af39afe7b3db · outbound

This paper cites V Jawahar.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding V Jawahar

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.803621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.431640Z digest=sha256:82864b14fe7bc9125affcc1fc727ef284af65ee92a46c321dbd05fb1b5b6fba6

Observation fab00ed6-d167-49a8-b0b8-87ff19e275a0 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Docvqa: A dataset for vqa on document images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.732117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.435477Z digest=sha256:31f760dfca4f495cb95e90afdd70625688284a5f9302988195a3985e16c25078

Observation c58d1bd8-7718-44c6-a068-62ea7207e501 · outbound

This paper cites Plotqa: Reasoning over scientific plots.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Plotqa: Reasoning over scientific plots

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.439981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.439981Z digest=sha256:6a3641e2906b26cbc5a7861e694f7078cc8dcf47274d222d0ccdc6f62c84cb7d

Observation 5a30488f-ad31-4d6c-9b82-c12b060e3bbe · outbound

This paper cites TableFormer: Table Structure Understanding with Transformers.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding TableFormer: Table Structure Understanding with Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.444111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.444111Z digest=sha256:c19ca172648f49235ae7cd1dc9448f08f515df7fde5695fde977d85218662137

Observation 841dcd17-c17b-4980-b4f1-ec88b48646cb · outbound

This paper cites Chatgpt.https://chat.openai.com, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Chatgpt.https://chat.openai.com, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.708357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.448954Z digest=sha256:fd50e2b633a7708d08eccfa35c11980fe3353a2e169dd7568f3bd4a9bdac9ea9

Observation 162f744d-64d1-433b-bedf-63d2467bddd6 · outbound

This paper cites GPT-4 Technical Report.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.453832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.453832Z digest=sha256:99049fb1c9aec589bf2e9f555f8c040826f0d24760a79c537561f8ed551310f2

Observation 5c5577d2-9d1e-4491-b448-e12047bacfd2 · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Training lan- guage models to follow instructions with human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.673222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.503220Z digest=sha256:d405aa1d0798fd613a872ffa792d414a111d213da3f824889bae48c4543c5dc0

Observation 6d581baf-0a6a-4422-ab2f-232eebd5f580 · outbound

This paper cites Cord: a con- solidated receipt dataset for post-ocr parsing.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Cord: a con- solidated receipt dataset for post-ocr parsing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.644893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.539927Z digest=sha256:94fcaaa33283333ee70e27ffb1582a83d0d11f86ba97a36f340a443724e33b0b

Observation 9783b180-e7fe-4fc2-a1cc-867a29e12868 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.545097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.545097Z digest=sha256:60bb4b88b9a995fa3bca890cf8c297ad8481f50df6cfbe5fc9e1ab03a3080b56

Observation 9d718dc0-ddd2-4652-bd37-83cac6334546 · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.611521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.627950Z digest=sha256:67d0af054e529b71449dd22d8fe9105cdbd5e391f3ac918a3669b543bcd22154

Observation 58d14f95-aa7a-45f8-bb40-7e2f252b0480 · outbound

This paper cites Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models.arXiv e-prints, pages arXiv–2403, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models.arXiv e-prints, pages arXiv–2403, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.595885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.748557Z digest=sha256:49a8c78e690a8c6f7ff7bcb54734d9786c5abb8ab6b7742eaef0acc655102608

Observation 2a6bd541-baf3-48fd-b6c0-6323a2ff5c15 · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17).

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar2017 competition on reading chinese text in the wild (rctw-17)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.523375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.795572Z digest=sha256:36c76bef403106b11be1a76a060723fa08ec3fcae732ff4cd800d90e8ace4408

Observation 00ce7920-1e11-49b4-ba66-e0ea3c37dcd8 · outbound

This paper cites Towards vqa models that can read.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Towards vqa models that can read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.505733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.800403Z digest=sha256:fd217ac1d69393fff576f5ef0fc9f659e98c7ebb7ae842357921880ae0960083

Observation 81387647-952a-479e-867d-044615ede585 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.473925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.805429Z digest=sha256:5679fc756b3ca006c19018a055cc0b3c56dc5eaa49d1beb381f086bc550e382c

Observation f64efca0-3d6f-4ff7-aca3-99c6fa340b25 · outbound

This paper cites Spatial Dual-Modality Graph Reasoning for Key Information Extraction.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Spatial Dual-Modality Graph Reasoning for Key Information Extraction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.809034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.809034Z digest=sha256:6069294873a15072ebe3cc682ce8b1d6498d03e17af4070fea515e5a8ddc6314

Observation 7ee44980-8101-4545-9080-8d4175fe7c0f · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.459557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.813894Z digest=sha256:2ed5f6c2f59da3738a35481865186c8ac50a9c0eaf7c5b158607a83d8d08f860

Observation 35ffffdd-eff8-4d0f-8ba9-996962eda127 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.818507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.818507Z digest=sha256:afb3ea78c1fe9886717461213cfdf79b84d9bf8f2bb244446c3ffc9a2a78cc93

Observation cac7b022-3036-4974-babc-824a47373cf1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.878962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.878962Z digest=sha256:997ce31e57122bae4a79d3e7db061a3fcbf1febcd8a2c46914636bdf3471389a

Observation c2a191ff-1344-40a5-a6b8-ca3f7b1a630a · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.969524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.969524Z digest=sha256:212e7ccbf07036e6d581ec0417d3868018397d6fe602771e9508169ee5b5cc6d

Observation 2f7fb109-156f-4f24-b39a-101539bb3ab4 · outbound

This paper cites Measuring multimodal mathemat- ical reasoning with math-vision dataset, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Measuring multimodal mathemat- ical reasoning with math-vision dataset, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.356298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.974634Z digest=sha256:323e1011dca3c9f2cabf4d1073f22c6fff69054cc413766bbb96f3c17260c486

Observation 7a1efaed-b0d0-4f4d-b5ad-a3abd983b7ea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:51.979528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:51.979528Z digest=sha256:5fb04790f718d17778020ced1d526e17f75975237875705ff040e83d94e72c96

Observation 2c687a36-62e2-4d01-a1bd-69b66d8fe781 · outbound

This paper cites On the general value of ev- idence, and bilingual scene-text visual question answering.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding On the general value of ev- idence, and bilingual scene-text visual question answering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.343783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.984044Z digest=sha256:005c6029163dd0c49b530d04986b905d22985486de3d4e911a972f30280307d3

Observation 5dcf2f51-9e6b-4896-8c6b-a2070646d065 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision-language model.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Vary: Scaling up the vision vocabulary for large vision-language model

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.276269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:51.988895Z digest=sha256:bd8c4c150af892de18a568b713cacbb3d6374f2ade71cefdf25333055380dcdd

Observation cf6d6f5b-fe80-4f4f-a99c-2df0ec142db4 · outbound

This paper cites Toward understanding wordart: Corner-guided transformer for scene text recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Toward understanding wordart: Corner-guided transformer for scene text recognition

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.259690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.076260Z digest=sha256:d831081b06ad3410c7222411578e43785a8f5f2c63583d02a40c52398580b767

Observation ab7e3ebd-126c-4a2c-be6e-4037f1143f18 · outbound

This paper cites Xfund: a benchmark dataset for multilingual visually rich form under- standing.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Xfund: a benchmark dataset for multilingual visually rich form under- standing

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.243919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.080316Z digest=sha256:f07bf0d92945d16dd44a6aed56339df8fca4aa39cb5ee91bfce4630bdeb1ee59

Observation 1e2dd512-7d28-422b-9972-c10ec4d8d51b · outbound

This paper cites Tgrnet: A table graph reconstruction net- work for table structure recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Tgrnet: A table graph reconstruction net- work for table structure recognition

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.141050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.084510Z digest=sha256:e54feb1d00f4e29675e81b4aaafccd91991a797ff5e8a3e5000b472e655911f4

Observation 1ef12314-b89f-4532-8fcf-fbde5f03a3d6 · outbound

This paper cites A large-scale dataset for end-to-end table recognition in the wild.Scientific Data, 10(1):110, 2023.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding A large-scale dataset for end-to-end table recognition in the wild.Scientific Data, 10(1):110, 2023

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:53.127344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.088527Z digest=sha256:f4c2a7f55de74bc465d76e09b7ecf00122706fb939854be5947d38ddeeb0b6f3

Observation 95a8e5eb-970f-4e92-ad66-8d2e24d4f79a · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.092596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.092596Z digest=sha256:150499310394a3b43129bbf707a60104e5785f4045db08d1fefcf2559d9660df

Observation 102e7b8a-ffbc-4d6e-be24-f0867ae3c2c6 · outbound

This paper cites Icdar 2023 competition on structured text extraction from visually-rich document im- ages.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Icdar 2023 competition on structured text extraction from visually-rich document im- ages

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.987500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.097699Z digest=sha256:10957e868cc4489453069f63f7384d7dc67c8b68f26a12bb0cb25b2c02364376

Observation 555b5dd8-002c-4868-ac4b-d776363d8406 · outbound

This paper cites Syntax-aware network for handwritten mathematical expression recognition.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Syntax-aware network for handwritten mathematical expression recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.974402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.101990Z digest=sha256:c103114470dbd19e702e15e2a9904d007d2d91931940e31eaf0a3549d68cea86

Observation 11aefa47-b1fa-445e-bf97-aa85ada1a6b7 · outbound

This paper cites Detecting Curve Text in the Wild: New Dataset and New Solution.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Detecting Curve Text in the Wild: New Dataset and New Solution

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.197635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.197635Z digest=sha256:7c51f7a47b23fea96de89c8ea41a1757cb030bb42708ee7ff1eacb5ed3407168

Observation 97b5a255-71c8-4c62-9a18-46fd0f6bdf04 · outbound

This paper cites LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.203082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.203082Z digest=sha256:0d4162a6a2baa7ac76f827b8f6ec7e55c853e2e271c70a995f29857005998dd2

Observation 3f52d260-71b8-4827-8cbb-3806f05b94e7 · outbound

This paper cites Image-based table recognition: data, model, and evaluation.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Image-based table recognition: data, model, and evaluation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:52.207933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:08:52.207933Z digest=sha256:536567ad9a7bbb1b94d5b4f914b984bc84e41597ced9d638e0335b522212045c

Observation c5ee0e39-02ff-4ed2-b880-fa60c8a1a524 · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36, 2024.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.881171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.211898Z digest=sha256:0465e2cee75f060f64d569a27eec96f8e6190585f36ccf075bcf83bb7ddec255

Observation dae84381-e88b-4ee2-bb1a-8dd3f0c4e0ba · outbound

This paper cites - Answer: The known answer to the question.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding - Answer: The known answer to the question

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.866566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.216345Z digest=sha256:4a7477af39f791c7e5493ef7f159cb8b21dee3963ee5d1e851113d68cc174466

Observation 50b3169c-79a7-4055-905b-18333b55a3e7 · outbound

This paper cites - `<txt_gd></txt_gd>`: Text with coordinates for context.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding - `<txt_gd></txt_gd>`: Text with coordinates for context

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.779914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.328887Z digest=sha256:78d79eeb021d028faaf85fed1731eaf2a11be6a951de9cfba18a8bd71aa8f81c

Observation eaf472a1-940f-4aa9-8f29-aac0b996b54e · outbound

This paper cites - Ensure that the extracted content retains its original formatting.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding - Ensure that the extracted content retains its original formatting

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:52.766246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.337274Z digest=sha256:0931c48633254fe0343da214eb1fc41be75c58d6e71ee53f21eba7e01fcf6e52

Observation 481d3b8f-4f52-41e7-860a-0a68407ee8f3 · outbound

This paper cites title":.

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding title":

Reference 80

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T23:08:52.752277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:08:52.342090Z digest=sha256:5296317310616ad51caf0d0b583cbdf813a2d18894b4cb1f4c1b86afe5e5ab43

Pith citing papers

No inbound Pith citation observations are available.