Pith. sign in

Paper Citation Record · LEDGER

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 75 inbound Pith citation observations for arXiv:2409.01704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.01704 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:50:57.814634Z

measured 130 of 130 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:53:52.033271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T20:35:34.407064Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact7
  • verified fuzzy28
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch16

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e42b16ce-f7f9-4e15-9848-74c15374e6b8 · outbound

This paper cites https://huggingface.co/datasets/Teklia/CASIA-HWDB2-line (2024) 6.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model https://huggingface.co/datasets/Teklia/CASIA-HWDB2-line (2024) 6

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.939483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d6047b39e6ed915f2099f853b544c398a3d13ab0b3829bfba033263ea426dedc

Observation a0339d30-0c36-4232-b9d2-d7727883167d · outbound

This paper cites https://huggingface.co/datasets/Teklia/IAM-line (2024) 6.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model https://huggingface.co/datasets/Teklia/IAM-line (2024) 6

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.943760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d9b2dc0486ad1d8032a76dc6c73bd62b06b4843887708c078d4204f618ba7264

Observation 86520682-abb3-41b4-91bf-977a563a714b · outbound

This paper cites https://huggingface.co/datasets/Teklia/NorHand-v3-line (2024) 6.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model https://huggingface.co/datasets/Teklia/NorHand-v3-line (2024) 6

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.945895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:c7481f9354c4fb2304fd24d0beba143b6b8cd91eb353702be4c77afa8cddc1c8

Observation 05cf8137-b9ab-41b8-9bfe-f42bfba1c56f · outbound

This paper cites Qwen Technical Report.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Qwen Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.867470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:8399f78b322b044bb812fe33c40d20349b08396c6f8c9e2c5efb61b8c414ff5d

Observation a8e08e09-0ea6-4040-86fc-2d33094a51cb · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.844042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:08ff75bc6124b603822a66b80ff9116dca92800cf0f603238916df3b3bfb19e2

Observation 9e922595-aaa7-4d6c-a70d-f7aa170edfc3 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Nougat: Neural Optical Understanding for Academic Documents

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.852382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:aac8345b48f38902ce35cfff38b33902ecb9a2972885528557ba70548dc48cde

Observation b77623ec-ee2a-416d-813e-bcf2461ac366 · outbound

This paper cites ACM Computing Surveys (CSUR) 53(4), 1–35 (2020) 7.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model ACM Computing Surveys (CSUR) 53(4), 1–35 (2020) 7

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.947757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:0b4e623c139a36a52e68e2f5ae9df85f2173771b4031fbbb57e1b3785e12cb08

Observation 98246ad0-f46f-404a-b663-842a2a3996c1 · outbound

This paper cites OneChart: Purify the Chart Structural Extraction via One Auxiliary Token.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model OneChart: Purify the Chart Structural Extraction via One Auxiliary Token

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.896433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:de0e957ae6d6dbb799401ca86f78e3698e416684d5184722e89f813bb372ad31

Observation b647b900-f7e5-4a9f-89af-cc4dc5a6b2a8 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.910456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:9fe347c80396d158b62824ba37c10db235590ece2a7418d5339e010fb3a1611f

Observation 7a415611-44e1-4922-97bd-7dcc1c008b6c · outbound

This paper cites PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.840233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:fe89b66922db28b5efc1ce4966c0fbc96a933a6b85a001987e8e70228312dde9

Observation c1fc507a-0fe6-404a-92e2-08f3b1aa7e14 · outbound

This paper cites In: International Conference on Machine Learning (ICML) (2006) 4.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: International Conference on Machine Learning (ICML) (2006) 4

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.949638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:f3cf4f55f70c2dd42a541442696b1e71269362b01ad7df9c4cff779ece84a7f4

Observation 840ff4db-3b79-4c46-8d3e-28ee0e780c4e · outbound

This paper cites Advances in Neural Information Processing Systems 35, 26418–26431 (2022) 5.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Advances in Neural Information Processing Systems 35, 26418–26431 (2022) 5

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.951588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:e7b13ea4c7f508f239c66e3055fe3bc74ca1f74f00dcc0cfd5022dd1da58bbdf

Observation 29680334-4140-4af1-ae06-56f51675dc11 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.864055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:c85ca053925a983ac8befde6f4df6d455573540f4b7b58f793873bd6a30d6745

Observation bcd9fa5e-d377-4893-a8f5-6ae4c9d6c675 · outbound

This paper cites Proceedings of the IEEE 86(11), 2278–2324 (1998) 4.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Proceedings of the IEEE 86(11), 2278–2324 (1998) 4

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.953392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:3dc3cb1307fe6ebedb2e56bebf5ea286d40c6a7399614b9af575802c8cc265de

Observation 30e3625d-7691-41a1-8edc-50a6f1a28613 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.875034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:0ca921f30ae5c459cd453cf74d92c46c8ef02e40973cbb4b3c503906071104bb

Observation d32a899d-f610-43d9-b37c-1605c776805d · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.955229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:575134bfbb5ad50f0ee3af62cb6469db4eb4381630ab13ffd74af9de914695ab

Observation 539bbe91-0bd0-45d7-ae30-623f97fb16ca · outbound

This paper cites In: European conference on computer vision.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: European conference on computer vision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.957199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:c5201e08c255ba93806d93e869516c27556fb454ad90032819b09fa180a6c28e

Observation 0ddbaff4-0cbe-41d7-b9e2-73f17f7aa9e8 · outbound

This paper cites In: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence (2017) 4.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence (2017) 4

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.959034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:a714e1d8a3620b7732141f5ab81d5c218b2755fde743428da2e3cc80333d27f4

Observation 1e7ca359-2f55-4c40-89ea-72244b72c121 · outbound

This paper cites IEEE transactions on pattern analysis and machine intelligence 45(1), 919–931 (2022) 4.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model IEEE transactions on pattern analysis and machine intelligence 45(1), 919–931 (2022) 4

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.960930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:0317ce5c2d65f814209dc31dada07f793af8edd666fdfc82834d5858bef72c7c

Observation 816924cd-6ad3-4340-bfc8-13b08c9031ed · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.848203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:033cf22385bfbcfde7d59aac195a1938cf60a1c8c83ba7e1bc149ed3702f0b53

Observation 9a51ff0f-31ed-4f4a-ab7a-2a40e93e16b4 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.962772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:9fbee1a6cf56d90a13a5ea00d85119bc803e79cae9b3ee404d10681ba06c50e1

Observation 83552657-d82b-473b-94ae-9a1f2ec4f7b5 · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.856462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:68079d189415065c4152066a68a95d471cd0ce36d921352ef2d7404c144460e1

Observation b005dc02-1751-41b1-9734-0949702479f0 · outbound

This paper cites an unresolved cited work.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:50:57.964571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:892353c948563a6e9c5b97f18847c9a06da01f03d9a39f6967a8513946246fcf

Observation 8ea36cf8-f70e-4bfb-aa83-06ddb666f457 · outbound

This paper cites an unresolved cited work.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:50:57.966444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:af548f2f588d5d328c77f6293f9d624a1640f235f3343de3612654d772c2c142

Observation 996bf395-d5f7-471c-8a4d-d67e788c0e1e · outbound

This paper cites ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.871222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:9c1753dae16ec5f7da24f15857086ec4c5d2d45b81400b89927b244594b2c19c

Observation 3016e0fd-e9a8-4736-974f-bfb7a5f3ed67 · outbound

This paper cites Pattern Recognition 90, 337–345 (2019) 4.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Pattern Recognition 90, 337–345 (2019) 4

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.968646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:8f3c6ddf1d09cd01d3527b92b7cba140149cd9661a9796d3e2c63f3aebcaaba7

Observation 3032a2af-cde3-4694-8376-af8e8cfababc · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.878737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:6174094d8eac242e2df5aa21710ec937a3df12499b060899e42ab35e7396a80b

Observation 643a49af-8a9e-4281-8cbe-89752ba8049a · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.885866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:29e9f05224b8195d6bcfaece621fb70ea4394704048ce3e5d95f7ac43540e442

Observation 6bc7756d-b20e-4751-8133-3a5fbc9c09cb · outbound

This paper cites In: ICLR (2019) 8.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: ICLR (2019) 8

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.970572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:45e0f20357531047eb71422f1ebc2d567199967d4f1ab7796b38e6cd65a4050f

Observation 642220e4-7796-4971-8edf-2995b5df6bb7 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.972562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:7c35a72f5e4878d518e956ee3e4c9d297f272decc2f69ef5b722939ad5ca4eb0

Observation 1f4f5004-0d44-4619-a9cc-8b607dd1f329 · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.916817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:c14567c442599515b85ca12afa57c9205ecaa5e9e04dd8179a655b8366c4aef8

Observation 091fe756-3425-42ad-b4b5-28929c4c0ea6 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.919606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:55d7491bc3775cf113df7a565ec9843543f3dba527b7f0f10973c7f38518122b

Observation 38ed69b7-55e1-43c6-87f4-5ad8e28ab190 · outbound

This paper cites In: Proceedings of the IEEE/CVF winter conference on applications of computer vision.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the IEEE/CVF winter conference on applications of computer vision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.974514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:1869b60fa4373f71bce90c37d3af71978edf35c5ff1d185d5b580685185ccab6

Observation 411de043-b84b-436b-b8da-7b1c55a8cdd7 · outbound

This paper cites The PracTEX Journal 1, 1–22 (2007) 7.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model The PracTEX Journal 1, 1–22 (2007) 7

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.976588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d4f99bb6f7e39c79ae7fc7525ea0d604da494d512c3b896896473aec0a7c971d

Observation 9f5463a2-2ea0-4362-a738-d7d4373b7914 · outbound

This paper cites In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.978590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:784dae013a36bb8ca3d11a469952d83900335cd1e6d49052d3f9d5dc9e49f5cf

Observation a796079b-98c9-42ff-b422-f72ca288d10a · outbound

This paper cites an unresolved cited work.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:50:57.980490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:b4bb3bda887eeae8ac3e187d45b582c50967a1cd6789865f685a33dd1d21335a

Observation 4088a510-f470-4d09-a79e-25dcdc8b1e90 · outbound

This paper cites In: International conference on machine learning.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: International conference on machine learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.982628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:072e611587df72b260a983c4647268542be9c8ed2260ae40276507cf872a0a9f

Observation 4db4e10a-7dea-4de5-92c2-a0cdea1b29ea · outbound

This paper cites Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.860246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:8fa4da11e0684529db15687941b4afc6f41a4b3bd2a04d6bb77465dcb233f834

Observation 6aae967e-6667-4ed0-b03b-98b87b0feaa1 · outbound

This paper cites International Journal on Document Analysis and Recognition (IJDAR) 26(3), 347–362 (2023) 7.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model International Journal on Document Analysis and Recognition (IJDAR) 26(3), 347–362 (2023) 7

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.984879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:30c79556a7c5a9cb30d9a79dab473204568ea0c8b6f3285b213354689a51baef

Observation c60d3ebe-c1c5-43df-a439-f73fcd965334 · outbound

This paper cites Advances in Neural Information Processing Systems 35, 25278–25294 (2022) 5.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Advances in Neural Information Processing Systems 35, 25278–25294 (2022) 5

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.922062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:01109f8e049546736824f34caa8f4dba4b918757aacbb5dd8ab3fd822c8370c0

Observation d2acb293-1f50-487c-9330-e23795783204 · outbound

This paper cites In: 2017 14th iapr international conference on document analysis and recognition (ICDAR).

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: 2017 14th iapr international conference on document analysis and recognition (ICDAR)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.924449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d8c6902b4d09a776902e6601ed39b5bc1bed37e2e34fa3254b45b81f171d6786

Observation 1d0224b0-2eae-40ab-90ce-cb177cdc53f4 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.926597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:1186596d6ec287fb1b1178aa567bb3e07a2a56c5283d6fb67c324a592b2e3b6e

Observation f7bf3861-6a08-4b4b-85af-c94b8f851f99 · outbound

This paper cites In: European conference on computer vision.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: European conference on computer vision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.928787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:bb3d45feff984c4dceb932b776e4ba342050921c8c4fa7f4918f6387878d927c

Observation 891e267d-cfe6-4107-95fd-e32739877243 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.882154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:1b973a624d54a0e03dca0a7fd2f875c21fe3f9dffe6be2f1cc97683f0a17eb3a

Observation 0dd7089c-ef9a-4445-b93e-7d6d6abb44de · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.931015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:c4ae6acbe44d41f970197d24b49982c4091afd9b80df456944bf6994d497947c

Observation d04d7704-ece2-451a-9042-66889085de75 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.889473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:645b9ff9b97cb3a5ef2df426e711a27b1e0ae19e877f6ea204ee9a5f2715a9f2

Observation 24c23946-3246-4124-a574-e50af50f122a · outbound

This paper cites Small Language Model Meets with Reinforced Vision Vocabulary.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Small Language Model Meets with Reinforced Vision Vocabulary

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.893024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:9176ad384cfa44965cc067e35dd349d0e7e1932514064dfc85e944274f00890e

Observation 447735a3-1299-4a13-9074-92cfea3d70c4 · outbound

This paper cites an unresolved cited work.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:50:57.933072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:3ed77c639feea8d2ec02ddd61a657bc0fe704661f38bf0d8acad0a0e056d37f3

Observation ca4339bc-d3f9-4dc1-81b3-4fc30fb7f0ce · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.899919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:545afdc605bef11edca3f02f9c09513611f69f3bfbb47793f4f748374b035796

Observation bf3a85f8-994d-4172-92fc-4ad13c8c5fd8 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.903141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:9188d6d9ceedfba1c1038ec03e9ae64c73d2b3ddd0ace79a8f3e9ffd6b52bd6a

Observation bbc42b83-15e6-4f1d-9cb9-a75ed3bc9308 · outbound

This paper cites ShopSign: a Diverse Scene Text Dataset of Chinese Shop Signs in Street Views.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model ShopSign: a Diverse Scene Text Dataset of Chinese Shop Signs in Street Views

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.906823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:fecb887dea85fe79e5077d139442f42ab29ddc96044e36672d33e798cf560ad0

Observation 04376c78-3274-44d4-bcf0-6e34456bff3d · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.935125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d90d327ff0e1c15ff2e9a5eac961fa21a4acfd220a96042b65ff25434b4ac4cd

Observation 297a21fe-7022-469c-9b71-a0740b0ddfe9 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model OPT: Open Pre-trained Transformer Language Models

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.913525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:d02bc73fe288c9842b9f6402104d01ae41e97ee000b0e6e7b64f4e36649f7344

Observation 5e999e3a-48c5-49a6-9493-490274c9b75b · outbound

This paper cites In: 2019 International conference on document analysis and recognition (ICDAR).

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: 2019 International conference on document analysis and recognition (ICDAR)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.937198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:020e563e8964fbbf38bc1302eee96c98d5fd0453240273c843b17ea505a5bd48

Observation ecabadaf-de54-4320-9096-3b59a7570d2f · outbound

This paper cites In: Proceedings of the IEEE International Conference on Computer Vision (ICCV) (2017) 4 19.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model In: Proceedings of the IEEE International Conference on Computer Vision (ICCV) (2017) 4 19

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:50:57.941507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:3faa2aeb16abe6e8ed0c62160b6b881e38559b8078f509bdaa86e1098e7a367c

Pith citing papers

Observation 9b61352c-f9e7-4ba8-93b1-0343acf84d18 · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:3b6f5e56b351d4851c2324754fc6653ee7eeeaa23daa22bb96180443a9667a72

Observation 1e395297-78b8-4680-a6cc-3f4442b7996d · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 256

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:15:47.044086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:d0e298469adbf16260b399e7a8e467c5945ee287c25f92efe62c51c1c34f9c40

Observation 8c19a771-7244-4420-8605-7120f2b21645 · inbound

Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction cites this paper.

Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:33:45.046203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:33:45.046203Z digest=sha256:916e9c01a54a341f41d245f14a1ac0f15172b1539262c75aa931226b942d70dc

Observation ed5a5282-cece-4f82-b510-45b796e7b417 · inbound

MIMIC: Multimodal Islamophobic Meme Identification and Classification cites this paper.

MIMIC: Multimodal Islamophobic Meme Identification and Classification General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:10:10.389994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:10:10.389994Z digest=sha256:10a77f9c602459e81727a9e777fd13e6015d0543281ae7371dff495ad1776c62

Observation f1ddf2ba-4fa9-4ea4-a4de-3b59de155505 · inbound

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy cites this paper.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.254073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.254073Z digest=sha256:0347aaae649be5383c35fb7f2072a31c9c915c76417ade8e71a5dd2c5b6bd5b7

Observation 132e1a50-9027-4c35-8248-00fea3a2b30d · inbound

OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation cites this paper.

OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T23:21:42.107756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:21:42.107756Z digest=sha256:08c59d0d239f2a1a66234b8d88166fc1e0e33a99e84370a29bd8157610bd773a

Observation a24772f8-bc34-485a-bf2c-578634d568be · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.706187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.706187Z digest=sha256:f3dc2f2807301b49a223190db7b9bae3a3ac5def8851ffbb426b1ef4c2812c88

Observation 631c7b43-e1b5-4cbf-b71e-308a40f6fbdd · inbound

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations cites this paper.

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T18:41:26.773703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:41:26.773703Z digest=sha256:104d87363ec87d0ffcbdf3e68cdf6f5d27c14d74072f0c1e078f3307d26ca2ad

Observation 3a165f9f-8af0-404f-9027-4c9205e448a9 · inbound

DocFusion: A Unified Framework for Document Parsing Tasks cites this paper.

DocFusion: A Unified Framework for Document Parsing Tasks General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.777751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.777751Z digest=sha256:d4cd6b062b4224ce994550923e4d054df2df858863a8afd001f20fbfe0fe07b3

Observation c2718c4b-f368-43ac-ba93-365dc2bff340 · inbound

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models cites this paper.

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:58:43.722047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:58:43.722047Z digest=sha256:2140921e6b9b332bd57a49e7be33052299114a00061fbacd6c78e60b945ccc1c

Observation 68c8970f-eb14-4f7e-8ff1-b6d4041993d8 · inbound

Progressive Multimodal Reasoning via Active Retrieval cites this paper.

Progressive Multimodal Reasoning via Active Retrieval General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-11T11:55:10.583946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:55:10.583946Z digest=sha256:eb2bf4632c85222a6d3b3bdd5a3c5ceba74124ef79b5169f7534e3b539c9ac28

Observation 247ff7ab-c972-486c-86ed-def90eea7d84 · inbound

Slow Perception: Let's Perceive Geometric Figures Step-by-step cites this paper.

Slow Perception: Let's Perceive Geometric Figures Step-by-step General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:42.678604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:21:42.678604Z digest=sha256:26412ff605e72b5c6c0078a53036f730459eabb3f7f6c970de1d225c0341384a

Observation 8b2f4c51-28a2-4949-a9d1-5d03b6f70994 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:580c4eb90c2f894f747392d6dd2b871e737b6167e48d70199dedb4cd2f0f1753

Observation b43f16ab-e725-4eaf-bf37-13af65de50bd · inbound

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents cites this paper.

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T23:13:14.003979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:13:14.003979Z digest=sha256:d9347a49d6f04e718f4b790bcb8380220257907cf328c003f477b17429ebefb3

Observation b5814798-dcd5-4b2d-81b7-208ac9a90280 · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.403437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.403437Z digest=sha256:1c6fb0dd6306fda1f61418fcef0f3024cfd22ead4fe29e69b065ad0120687ffb

Observation 7711f7ff-4629-48e4-bd52-a996a0bba7eb · inbound

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition cites this paper.

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T22:57:16.881380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:57:16.881380Z digest=sha256:852ce3cdaae68cf120c6b144c56d0bef4eb31828fab93f68348eacc03d5ac875

Observation 0b73bebb-7e9d-4dae-9ec4-25d3d96d9260 · inbound

Muon is Scalable for LLM Training cites this paper.

Muon is Scalable for LLM Training General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T23:02:51.656353Z digest=sha256:e4ad8e8ab97501ecdc2ccfb862d6fd9a77dbd3478642cb164a5fa94f07055b54

Observation a27c7054-d39e-40cf-bb19-cae973a8309c · inbound

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR cites this paper.

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:32:04.643167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T20:31:34.074705Z digest=sha256:b7575410cdb3cab30959ef1fff2eb56ae7100cf95d55f55d4d3bc3067f9bc415

Observation 33026d3b-32b8-43b2-b4ad-aba31cdb9822 · inbound

Document Image Rectification Bases on Self-Adaptive Multitask Fusion cites this paper.

Document Image Rectification Bases on Self-Adaptive Multitask Fusion General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:53:52.033271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:53:52.033271Z digest=sha256:56183726c894c1fcbf8dae2f981259f5fffe830041c3e1a47b860d17a1f40ce6

Observation 6141f1d2-1314-4115-9085-0329bb18c23e · inbound

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? cites this paper.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.623999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.623999Z digest=sha256:ddeb7e073ef4b279b7b80446e0298c575224f3f195cfc5248bc483831082c82e

Observation 05f74a1d-8b18-407b-b1f3-b588864851e2 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:29.746865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:29.746865Z digest=sha256:c07c81e63b22e1b509d67feee3ce90e668601f35183e9cb95b30257b90d01b86

Observation a060338a-bfa2-405d-91f5-108686d35f23 · inbound

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning cites this paper.

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:37.205231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:37.205231Z digest=sha256:d8bd7030cb1f7fe8add801a547ece594316c2d4cfea3f20afe8c293ad9ee2546

Observation 20ca24ff-9389-43fb-b12d-ccb9e7061bbc · inbound

A document is worth a structured record: Principled inductive bias design for document recognition cites this paper.

A document is worth a structured record: Principled inductive bias design for document recognition General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 8

Resolution
malformed identifier
local_arxiv, observed 2026-05-19T05:02:04.440884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T04:57:35.441758Z digest=sha256:706f70a1ec9602e218822e077390b6eb54bbfa9c8a09010f63061dafd5bb8685

Observation 80793239-78d1-4cc0-950b-73f4e181560d · inbound

Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil cites this paper.

Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:19:33.643325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:19:33.643325Z digest=sha256:61f3ec45f52631698f13aedc18f6104273f0f78cd8f59dea849f561a3640bc33

Observation e79f9c03-8f8e-4163-8f48-d84cf427e9cc · inbound

LMM-Det: Make Large Multimodal Models Excel in Object Detection cites this paper.

LMM-Det: Make Large Multimodal Models Excel in Object Detection General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:41.300411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:41.300411Z digest=sha256:e48945f0f76f3a612b4a04680d67a2c83b2ecab2680a2fb2ced5c3c13659c9ff

Observation ae9a255d-3e03-438a-8fe1-ed02d81dea12 · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.420714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.420714Z digest=sha256:d0fd67ec751fc8ff7820292bac721ba7309ae6078fbd69f48e99ecd70db7d3d7

Observation e7445f10-8aee-4e75-9637-f528bf00acaa · inbound

IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning cites this paper.

IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:20:12.582337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:20:12.582337Z digest=sha256:61d39177a8a64126967e088a474e46e10f10c2aa2cabcfb19a14822e78bd4d9a

Observation f345069e-23ab-46ec-b1b7-be2235b54907 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.591046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.591046Z digest=sha256:66ac90fbb35e40e43104346fd266dc068c0e957ddef0c6bf197b26f871375c17

Observation 71dc6786-089a-44d0-af56-97bc2990d45f · inbound

MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval cites this paper.

MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:26:48.764056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:26:48.764056Z digest=sha256:ff9040029be0553e09e46b0df7844c164a3f36d9b718f734d144cae2d26c37f4

Observation 765ad383-e252-4db4-9a90-bec6d7788331 · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:9bcc40efec004b0e6c68fb9e27237ac7fec127ec314fdebfac78dd6b8f9c7b8e

Observation 756a901f-d8bc-42ea-b6de-7714c6a759cd · inbound

DeepSeek-OCR: Contexts Optical Compression cites this paper.

DeepSeek-OCR: Contexts Optical Compression General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:35:50.647950Z digest=sha256:d7b6d9f0b6ac2f369bedfd9a7c8eb41a0be5bd651024db079ae849a1ae062078

Observation 18d8d4c7-1379-42a6-a82a-31ad7dca5d47 · inbound

Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images cites this paper.

Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:26:35.645223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T13:25:37.515812Z digest=sha256:53447cb0fcb7154ebfa095a1476e2a18cbc8d2e3d89dfc84041a1738bd854e56

Observation 605d2982-70bb-4c1e-8f10-895ff9f74128 · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:6b44ebe35fd5079f09fa12db7c0b80892b95c4cbe4dca809f7d882247c6936fc

Observation 7f16cff5-3871-436d-93a9-2626f4638884 · inbound

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation cites this paper.

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:30.541979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:30.541979Z digest=sha256:cc4e6d70290db26d1f4fe3df4d9e1b682c892acedad5817ac4e07dfa0c6faa5d

Observation b19ebe6a-47ea-4785-8136-98a70a01bf87 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.859133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.859133Z digest=sha256:76389ac5827c64d8b4515e659237775cdba1810fe818314b2a97062069d60b18

Observation 410fd685-5f7b-400b-9dc9-880c8072c599 · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T16:25:03.743594Z digest=sha256:2761998183979f20bb7851da1096af8e7624b2216596c109c34b60e1c34e3a20

Observation d394d786-bcc7-45b0-ae78-a5bd01e36e7f · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:04:14.700455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T16:01:52.150950Z digest=sha256:8a6f1970cfcb9ae1dbca32165d1004d13765a1b51116a32a37af63b560b83b0f

Observation f6ed3d2e-45d9-4b4e-95a2-1727ff1ccdda · inbound

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain cites this paper.

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T14:26:08.188950Z digest=sha256:b9f17713266a6bc663399580c05db46d384d53eeeb2a867e42adf390cde38b59

Observation 8374201a-4916-4d89-8531-ceb6b2fd6130 · inbound

Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding cites this paper.

Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:44:17.198748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:44:17.198748Z digest=sha256:40bc318d317578c0cc5ff05c59da4f08b47cda0e438411024565674ceab0b5bf

Observation 2b592335-508d-43c3-af07-c65e28a0ddd3 · inbound

Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:8d7e52a769696ce7f3c6ad85bdbcbcb1b2fe7a7861f2eb613197ed692d0dbd3c

Observation 45bc09ba-417f-4053-9ca0-68d1e493c820 · inbound

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding cites this paper.

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T23:44:40.099054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:44:40.099054Z digest=sha256:b5fdb28c6f84fa0d4f56be6f3dc2641e45cda95d0d0800c8f1c1595290b57ad6

Observation a9f30a76-cbde-4426-8a68-f639683e31eb · inbound

DODO: Discrete OCR Diffusion Models cites this paper.

DODO: Discrete OCR Diffusion Models General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T22:27:43.628538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:27:43.628538Z digest=sha256:fd1ce716eb181804d60ef44c0b3e91247f1b1b87d519920a41067ac27b0ac92e

Observation 1bba8d35-e6f2-42cb-af78-bdc67548e64e · inbound

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training cites this paper.

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T01:15:26.757215Z digest=sha256:4d36564e3791aac2b234ae80d59feca86bbbcffede93c2582b1fe757a7959cab

Observation 7ec6c456-bd98-46dc-a400-ca6449cd05c3 · inbound

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing cites this paper.

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T00:25:19.782732Z digest=sha256:02ea950d5eb487c9b46f983495e2592d8602756f2213c3e95309520641228b9d

Observation 9c01214c-a45a-4446-a616-1825defcb246 · inbound

OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning cites this paper.

OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T23:17:45.005518Z digest=sha256:8d76cb9434717f9ef29e464fe05a5ab1bf93b150404a8840405fdf7b427958f4

Observation 0cd11a8a-a797-40f9-8401-9826349b4258 · inbound

OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning cites this paper.

OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T15:15:25.977993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:15:25.977993Z digest=sha256:e9e98baf14659ff009d91cde3de208a8044b3e2e3de7844aeb28f66aae88619d

Observation 69a4efdd-10d5-4355-9efb-0c25a196660a · inbound

InstructTable: Improving Table Structure Recognition Through Instructions cites this paper.

InstructTable: Improving Table Structure Recognition Through Instructions General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T20:53:57.029294Z digest=sha256:8e19f6e367350f8319487e80947ca6b0c813de26aaf01828af6def1ee4a082b5

Observation 780c6f79-2f43-4090-be7a-0de66f9ae74d · inbound

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale cites this paper.

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:58:41.377996Z digest=sha256:6619fdef941401cd83e4eeaadcfabdd6e8c2e016274b88f8e4ffe8757fab8f7e

Observation ce28198a-e3fa-4bf3-8ae9-9482a216de9a · inbound

TableSeq: Unified Generation of Structure, Content, and Layout cites this paper.

TableSeq: Unified Generation of Structure, Content, and Layout General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T08:53:14.886099Z digest=sha256:005028b91406e146dce9a2762de58bfed8d7ac974ab9580f636d62373a82cf88

Observation 792ceaf0-50e4-49f6-bf8f-d11706e504e8 · inbound

SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making cites this paper.

SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:05:18.970086Z digest=sha256:5a94e1ede0ccdb5840241085de6caa852f1d0e82b95facd887c7eba4e4051972

Observation 66b6a92f-8620-4852-b73c-b39f129ee88d · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:e9163d0aa90203aa311c8c4d8fdc285caf75c60d504a2eb193ab935a1c6d749a

Observation 44a9747f-c778-4d3a-90e6-0df217f46c00 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T09:54:47.097063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-22T09:51:40.160096Z digest=sha256:9a4835d06626f7b52bd8b54586d27a36f1123d5d2064e39a3857799b223fef62

Observation b19bc4fb-50a4-444a-93ed-d0157391db56 · inbound

FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing cites this paper.

FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:58:24.813176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T14:53:57.376715Z digest=sha256:0ef3524b097d31ff07b44763aed01f17cf85654cf889b6fa6466e7ae585ee1d7

Observation 9151640c-0f97-481a-b94c-4225f20be6b8 · inbound

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding cites this paper.

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:53:05.057867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T05:48:34.771799Z digest=sha256:358b7f24af2df4faf0958bacd8e04142364f59a1a7807fa5d388d757fb361bc7

Observation 53fcbf81-7705-44f2-bdd4-9d5f88fa8717 · inbound

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation cites this paper.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.695715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:1c2d936eeeaeee45eedcbcb1007a93432b4be2ce2941a520dc5a3046fc54a57c

Observation 855fc805-dc42-4dfc-af08-c65114863b62 · inbound

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing cites this paper.

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:01:08.919944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T05:58:04.055855Z digest=sha256:9df2afc820acecd4da83f846bef37bd1e8c6c940bf8f5977df698ce5aa196384

Observation 917c201c-0b89-4539-84c8-18eb88a4b3ee · inbound

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing cites this paper.

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:05:48.197561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T17:37:33.750306Z digest=sha256:7c02a83a0a3fee6db752cfeabcdc4484c8ee0278a239f57067ffbf53e279e1e7

Observation 4197ac5c-f05a-4f4f-bcdc-448003cc3519 · inbound

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding cites this paper.

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:41:15.193383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T07:36:53.347112Z digest=sha256:0045bb532a1e142397e5bad017a9f8cd63a6e68ff0162a512b1ace3836c0e0d5

Observation a817441c-13b4-4a0f-bbfd-6446ee886f0d · inbound

ABot-OCR Technical Report cites this paper.

ABot-OCR Technical Report General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:33:28.133191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T13:29:17.221676Z digest=sha256:e3bdb159ad7c8a5acd9589e62fe6a39530a3db6cdde54a514749f4db9bf5fcca

Observation 58049c56-4317-4cd3-a67c-316ef1a20f74 · inbound

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training cites this paper.

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:28.805474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T10:39:14.444486Z digest=sha256:1e9d24dcf8c2dcc1186ce42b2f7dd235cea40a21fdcf16c3388caed2be7de61b

Observation d0e80d9f-7c30-4861-8632-e3c0ee8ead87 · inbound

EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval cites this paper.

EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:37:35.702362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T15:04:30.757297Z digest=sha256:2f11c62b7ca26bc44fd69cfdf2a6a497b62530c7723673e7344443a65555f85e

Observation 358b28e6-6d8d-4f1f-9519-50d00677e258 · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:38:44.217382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:5f8840031f74ced59e2bdf2b8f5d06c332e344a0e6df1c41fc863fb347e9bf58

Observation 1c48e528-0420-44d2-ab42-5e5b6fd5aa8f · inbound

Unlimited OCR Works cites this paper.

Unlimited OCR Works General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:19:46.939595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T09:03:33.539615Z digest=sha256:3495ae86ee56f74ff4d1ece39d8fc4ddd267bc580f785f55b266c3410a900d5a

Observation 062bec07-ad84-4c20-9a1d-9c1bfd440594 · inbound

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods cites this paper.

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:59:57.268402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T01:06:05.381643Z digest=sha256:f3741868e362b4510939fa1220e7ec3a5c200307f37c44696893336aa96f7c68

Observation 1cf928dc-6511-42c1-beef-edc96c8a0884 · inbound

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study cites this paper.

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:04:28.313205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T07:56:37.081738Z digest=sha256:c534ddc73107907fbbebc11948b57b627ee1787957deb1411b68628b41c87e84

Observation d6f72525-8cef-4759-a7c2-9bd351202d7e · inbound

StrucTab: A Structured Optimization Framework for Table Parsing cites this paper.

StrucTab: A Structured Optimization Framework for Table Parsing General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:24:18.973209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T06:22:21.694355Z digest=sha256:055b5a28026884e5461dc6de954dcab09acca07197ced5bf822b59e1d4557d1b

Observation 847bce4b-b337-464c-9cfe-28ed85053190 · inbound

MORE: A Multilingual Document Parsing Benchmark and Evaluation cites this paper.

MORE: A Multilingual Document Parsing Benchmark and Evaluation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T05:51:08.895411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:51:08.895411Z digest=sha256:d72061d7f606edeac6f3b772f5196ebb3e7a52bb8d27ea38d75bf7ac65a3bef2

Observation 49e1fae3-e195-4641-82b5-f925d3b8703f · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 229

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:53c26b1b1d5f2df1fe860f1eac78529aa01658b7bf99e4aecef92cb029b74a16

Observation e7f01597-c3ed-4648-809f-25843272ac17 · inbound

CMDR: Contextual Multimodal Document Retrieval cites this paper.

CMDR: Contextual Multimodal Document Retrieval General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T20:35:34.408404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T20:30:59.126121Z digest=sha256:ef3e1d024b4b2486a08dfb08c9b2da265f1a01dad41e804a34cbdc1ee9054eba

Observation 78530a43-8645-4caf-a68d-4e29b682bdd7 · inbound

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI cites this paper.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 111

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:a5b8ded027f5901d52f9e1d7190e8fab488da98ee04d55455576ec8b03961b20

Observation cea1a149-72a3-40d7-8737-1d6faa99ee2c · inbound

Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study cites this paper.

Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T02:58:32.052229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:58:32.052229Z digest=sha256:20d252fe879b34528d5f91cd130316bcf0284029c9cb281ddd93cedc22c38a95

Observation ded879e8-3c3d-43d7-9fe5-eeb642491ab0 · inbound

Pixel-Space Diffusion Transformers cites this paper.

Pixel-Space Diffusion Transformers General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-01T17:35:39.974969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:35:39.974969Z digest=sha256:f0f1533f6f9fa5dd35862cce393072ecdb94a5f60567338eef764d650b7473b4

Observation c6ff5059-e7eb-402e-8a62-6df60cd5ae8c · inbound

LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR cites this paper.

LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T05:34:45.864348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:34:45.864348Z digest=sha256:3cedbe559196a68746bd88f10d16003dc5563324b4382cdee6c8d7de4e5edf29

Observation af6a8136-d713-4b21-b2a1-66bd23559e0e · inbound

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation cites this paper.

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:56.984456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:56.984456Z digest=sha256:eccd88eba6164cd11f53fbd3a2d0a43cce16893d080b81f08db8742e639a7c81

Observation 21c67e3c-020f-49c5-9493-54af6ee44773 · inbound

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression cites this paper.

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:05:39.229824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:05:39.229824Z digest=sha256:17201d8657a0a6f928d6f5faea04fe5ff765ed986db6341308316d50cf5def67