Pith. sign in

Paper Citation Record · LEDGER

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

As of 19 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 18 inbound Pith citation observations for arXiv:2412.02210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02210 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:47:10.303241Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.645575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.904748Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d312f58-a6a5-4c9c-9810-649d1a422ac7 · outbound

This paper cites an unresolved cited work.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:47:11.388706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.019976Z digest=sha256:ecf3fbb06144336fc17fc9f3fa056ea520349a222fde8081671685343e1d9f92

Observation 31aa5d48-02f9-4e3f-842d-dd6422a2d22f · outbound

This paper cites GPT-4 Technical Report.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.026585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.026585Z digest=sha256:b06a958894cc7c640e4be06704d60e91840443aa4de314ce5222ff37d51c8fbf

Observation 2e9d4d9c-600d-42f8-9dad-dcdbd68c47c7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.033823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.033823Z digest=sha256:c04e7d48907f6e03e36089ad1698d7995a2e86efe2c54a205db66cf880cf715e

Observation 1ee5e262-a344-4d7b-9077-902b4abb54ce · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.040694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.040694Z digest=sha256:556489779e542431ffc0638b1147b61a845799311364e512be9755c4262bb7a8

Observation 0ccdb98b-af15-4edf-b688-0495748d2e3c · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Onechart: Purify the chart structural extraction via one auxiliary token

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.368604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.047209Z digest=sha256:c6b299ab465bb9b3d4a77763f9fdcd6d7fd284efc1b7e4b84119ca11746bafe8

Observation 85f1cf64-7422-4120-b920-959fd9266711 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.350282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.053297Z digest=sha256:60979b7fa37c0e85b446c86dc24f003a92d11a9b31605ec31682026480c53481

Observation 1b189e01-9fe0-4e59-9b2d-6623439d76d3 · outbound

This paper cites Total-text: A com- prehensive dataset for scene text detection and recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Total-text: A com- prehensive dataset for scene text detection and recognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.330529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.060223Z digest=sha256:64780ceac7c3937f7a9d25197809b86b09850ca0c57002daef868514c6f1feb2

Observation 8700ec30-408a-4a54-928e-c9ac28f40d6e · outbound

This paper cites Icpr2018 contest on robust reading for multi- type web images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icpr2018 contest on robust reading for multi- type web images

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.308637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.066507Z digest=sha256:cf9de37e128c8ed05dd7a8b96277fb0c2181954292e0697a172bc3518da5048f

Observation 6134128f-727b-48ea-b3de-13e1898d02a4 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.073378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.073378Z digest=sha256:f354b87cf5e8e5cd35e1b23d373b15bd1f8648f82022b98d4b88f1b8d63bb6fa

Observation 00f847f6-4254-4297-9a4c-3ec2ff245a1e · outbound

This paper cites an unresolved cited work.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:47:11.290550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.079322Z digest=sha256:2b7036928781afd7df8f455e19e8a6abf98da6da3d0441a6011217a3669356d4

Observation 9a410f80-8faf-42e5-909b-6e6226005c6d · outbound

This paper cites Post-ocr parsing: building simple and robust parser via bio tagging.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Post-ocr parsing: building simple and robust parser via bio tagging

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.271230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.085071Z digest=sha256:3e51f9ffe645867c3c2f60abddfaa4c6f4ea03e29956b0e2cdc0a00ee181aa67

Observation 15702156-a23a-436d-971a-6e4dfcfcccaf · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Funsd: A dataset for form understanding in noisy scanned documents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.253184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.090339Z digest=sha256:81e0bc49bae44dc77c591de1299876bb20b475e0a1cf73ad57bc1302127e878d

Observation a83c0578-8d9f-49fc-a64c-e6378e62c762 · outbound

This paper cites Icdar 2015 competition on robust reading.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2015 competition on robust reading

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.236087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.095690Z digest=sha256:0599680801d092eb91e65ccbd53c2714a5d363a5e756214de84db489b77bdd65

Observation 23ba1a9b-30ac-4852-beab-aa6af8349b5c · outbound

This paper cites Ocr-free document understanding transformer.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Ocr-free document understanding transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.216001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.100963Z digest=sha256:9f26735caa6d0acbd0177000dd2bdadb6a3fd64d9783ce93734c912c4e1f2f82

Observation 77421711-b0f8-4c55-b65b-f8a221ac9727 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solu- tion.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Visual information extraction in the wild: practical dataset and end-to-end solu- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.196174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.106119Z digest=sha256:48d75cf00da78bdaefbb2d74bbc3e0ce3f0fe115883a742c588900a66cbf8ea9

Observation 0fa47a8c-203b-4719-a464-ed83920f3e01 · outbound

This paper cites Binary coors capable or ‘correcting deletions, insertions, and reversals.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Binary coors capable or ‘correcting deletions, insertions, and reversals

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.176016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.111356Z digest=sha256:99784e45e1eb7f6a729db66bbdd47b0d751e4d18ce36451d61dd4c1fb3ce5872

Observation 4b8f8499-f103-4abe-b669-406e7323761b · outbound

This paper cites TableBank: Table benchmark for image- based table detection and recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TableBank: Table benchmark for image- based table detection and recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.154803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.116371Z digest=sha256:4e746efdc58f26996356c6e44cb87dcaf2532d911cdfda7b618d80cdabe91b72

Observation 87624657-47b0-4cdc-9cf4-d1c3dea777f0 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.121280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.121280Z digest=sha256:cf6706a47f15defde12eb1f67c482ba28381c6b93683b531b5fa8e827cbcdd5c

Observation 19a303e8-bf54-415e-a658-ed2a8f03fdf1 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.126191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.126191Z digest=sha256:89b7780e914a0f005bead0fb73bd38dedefe3fd638fa93cdc2d6cec42c4480a6

Observation e5bd558f-d825-4799-b185-fdde556e1d39 · outbound

This paper cites Spts v2: single-point scene text spotting.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Spts v2: single-point scene text spotting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.136303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.131612Z digest=sha256:c1197492d5d341badd2543735d71ddbcdd6678c077f3d78b1df9c2b8efed182b

Observation 76f39e67-cb53-499a-8984-cc1e90f7ec3f · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.136168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.136168Z digest=sha256:487796714a51fc9100d27f82ce4e12d8d0f1f8d8e3ef6ea69f0e5b9aad24be3b

Observation 2a484f35-69bc-460f-9ce5-d3e44ed11fe5 · outbound

This paper cites Parsing table structures in the wild.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Parsing table structures in the wild

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.114773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.141463Z digest=sha256:d8164d9de2eceb205be990f54102f060b60b617a0ad922db34c700f5f47582c2

Observation fa1ea388-b6cb-47d1-8f54-40660dbb9143 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards end-to-end unified scene text detection and layout analysis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.095493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.147473Z digest=sha256:c29e9eae96a08631122d1dc7e466298f71859eff87c350372026fa2b330abc03

Observation fd6e2449-2c2f-44fb-93cc-a3bb9f6f235b · outbound

This paper cites Layoutllm: Layout instruction tuning with large language models for document understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Layoutllm: Layout instruction tuning with large language models for document understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.075697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.152557Z digest=sha256:024d1518b33baa143c4c00b1b2144fb63d54e65afc57d61ffa51f6fc07dfb601

Observation 92df3bc4-0195-4ce9-bdc4-f6412729c6ed · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy KOSMOS-2.5: A Multimodal Literate Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.158358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.158358Z digest=sha256:78c182428e8d7c6b4b2780aae46389916fc009711efaea1129a30c80d0371a0c

Observation 18dc72a7-fac3-4fa8-853d-ea813e396e4c · outbound

This paper cites Icdar 2019 crohme+ tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2019 crohme+ tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.056054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.164365Z digest=sha256:6a16ceceec61a70147027d677eed8edf164f33d34c7b005441aee409fe2a3819

Observation b149ca7d-833d-4d4d-a03b-f1f8b2fd010c · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy The iam-database: an english sentence database for offline handwriting recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.037443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.170323Z digest=sha256:42cfed72fc88e28ebc867d5cce3975792861ffb7a06e9cc8c1cede17548d5a47

Observation 8abdbd23-e6b1-4282-9347-94fe9b313cca · outbound

This paper cites Docvqa: A dataset for vqa on document images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Docvqa: A dataset for vqa on document images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.176206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.176206Z digest=sha256:c919d7c54201ec7ec464c41d754570c5565693fcb8cf6d0682cfffb9e07156c8

Observation 769ae03a-6242-4162-92e2-935b755750e9 · outbound

This paper cites Scene text recognition using higher order language priors.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Scene text recognition using higher order language priors

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.999313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.181921Z digest=sha256:f19dc791b8f285970b6ca86bb77df1c1a90c7f0ec4d38bf7706fe76ffcf4beca

Observation a766abb9-13c1-4085-8164-5ac46b60b34b · outbound

This paper cites Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.970598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.187350Z digest=sha256:e96f2365a7b2e1f02bbe5964f86b7b9791dbbbf4c3e8013b6e12651e922197af

Observation fe6323e2-5fdb-4174-a558-21e2c660c109 · outbound

This paper cites Cord: a consol- idated receipt dataset for post-ocr parsing.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Cord: a consol- idated receipt dataset for post-ocr parsing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.944130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.192956Z digest=sha256:18587b20ec0d289b9a16d474981a02d9d91f87fd1470bddd2a79cc9497a0dd1c

Observation 1b6993c0-e907-459c-9b38-e425299e7a76 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.198930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.198930Z digest=sha256:d6c5321e0489ad7cf6a33421a3d205ef51ede2bc8e2b3b6601d1f634bb10372b

Observation 0cb5d40d-e1cc-4fcc-ae6a-e02835439dff · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17).

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar2017 competition on reading chinese text in the wild (rctw-17)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.893043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.204545Z digest=sha256:9715e37d0c4335472df0166e99b33df7e859625a6b6e9267affca365ffa21e69

Observation 5f0525b2-d90d-4968-ab5d-d7a0784162fd · outbound

This paper cites Towards vqa models that can read.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards vqa models that can read

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.873328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.211029Z digest=sha256:61ca1efb38039c28a3c990e7fcd9be2f0fbfc4407a8074f59b30c530fff8bb12

Observation 1ce07450-1c4a-4083-8ba2-6ee9a0e841f8 · outbound

This paper cites Seglink++: Detecting dense and arbitrary- shaped scene text by instance-aware component grouping.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Seglink++: Detecting dense and arbitrary- shaped scene text by instance-aware component grouping

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.850688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.216475Z digest=sha256:023e9dea880f2fc48b7defd1f4501ab8f48ce1bcb547add96baf60c9f32c58cc

Observation 9378978c-3f0c-47be-9c47-877ffb7c6e69 · outbound

This paper cites Mtvqa: Benchmarking multilingual text-centric visual question answering, 2024.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Mtvqa: Benchmarking multilingual text-centric visual question answering, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.831095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.224158Z digest=sha256:1cf1e885479ff10d2d69396eda5531009d47f331b3efa05fd06db671fca79df3

Observation 22ba9c20-d406-4224-91c4-1a0f4c486bd3 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unifying vision, text, and layout for universal document processing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.810874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.230406Z digest=sha256:b166f3d30ecb2772ef7e4d799df10bd5365f7c3fe5ddf62c0decb4aba4a05098

Observation 9908f796-e438-499c-aaa8-a800c6c338f2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Gemini: A Family of Highly Capable Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.236564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.236564Z digest=sha256:acfd8279b469644a6388789e44a893beecf95a688d6704657837a01d276eb7e4

Observation 3feefc3a-554d-41bd-b8ad-2f0155313e24 · outbound

This paper cites Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T23:47:10.480277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.243109Z digest=sha256:0b06dcfe7cb0f5703e0aae8f0538750e21eb3e533f895b33cb8acce18ba81035

Observation 857d57d6-d5bb-44be-91f4-ced6d5d97b6b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.248636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.248636Z digest=sha256:8f86e5391d0361fdd7ee322eb565090a3640a28a54aa27ecd8545ef7241bd2ce

Observation f1ddf2ba-4fa9-4ea4-a4de-3b59de155505 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.254073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.254073Z digest=sha256:09c500961218bf735684cbd8e6d210ba30321065cabbb6736749574588661a0e

Observation 1a4e9fe0-786e-4466-94ef-e64971c3a63e · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.792916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.259956Z digest=sha256:31ca95b80a2609a93e0b1dd4bd6ad0afcd63d9131866cb8d16376772716361a6

Observation 13e1739b-7b3f-4fa0-a0e7-d06b2bf93856 · outbound

This paper cites Modeling entities as semantic points for visual information extraction in the wild.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Modeling entities as semantic points for visual information extraction in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.772300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.264455Z digest=sha256:acbee5719c3e20198840b4fddd3f8e3ed645075dd777406f0cd8bb52574c03df

Observation 21a08f2e-8c81-40b5-a8dc-1ed1bb00eca0 · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.753010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.270597Z digest=sha256:b3e3b5efc7c8cd90331a03d4524ec3ca18059651944eb8f12fe4eb390032b5a8

Observation 97bd656e-09d2-49df-b599-6dab9842d4cd · outbound

This paper cites Icdar 2023 competition on structured text extraction from visually-rich document images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2023 competition on structured text extraction from visually-rich document images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.735243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.275337Z digest=sha256:7c9bc54045b11346fdf256efbc5e88097cb4ff2e8e68ed54fd21ca6b3b5d7d1a

Observation f460d0ff-ae55-4d5e-9bba-2f682d51c8b9 · outbound

This paper cites Syntax-aware network for handwritten mathematical expression recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Syntax-aware network for handwritten mathematical expression recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.716552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.280683Z digest=sha256:2a7b481edff197642ef3582734fb0866d65e38cdce02cb331205ac7db78f85ec

Observation c139026f-6b05-4428-96e3-01332c86666b · outbound

This paper cites Detecting Curve Text in the Wild: New Dataset and New Solution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Detecting Curve Text in the Wild: New Dataset and New Solution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.286185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.286185Z digest=sha256:03190ff9deb0c485c20e6de555a11f1af5e2c34c3f6f98a9e1156a7005e70d7e

Observation 22305401-e979-41f9-b576-29c2070a894e · outbound

This paper cites TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.291965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.291965Z digest=sha256:b4e2d1a763b6ba53dd5fc950114373b7db92ce5615213c47924f153bb74ace2a

Observation 9018851f-08b9-4ffd-ac40-e4d3fc27a634 · outbound

This paper cites DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.297668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.297668Z digest=sha256:aff765ed2a519d5a333425402858e4eab2ffd22c2d9a3c8b2865fc485e6d8df0

Observation 9edd0b2c-ef4c-4f1d-8e9b-c4918a5ce15b · outbound

This paper cites 3.5 kg”, the true value is “3.5kg.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy 3.5 kg”, the true value is “3.5kg

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T23:47:10.696214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T23:47:10.303241Z digest=sha256:ffd1116f63a62d0b2875f0a15774b610d6f6985771487f3574488b96b0d6c9bd

Pith citing papers

Observation 44a67d08-3b66-41b7-a987-03b28e41f3aa · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.764403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:39c1cac3db07ea16782340ef10f435c66a49b0f045d65696488a813d8c80f990

Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · inbound

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? cites this paper.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.645575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.645575Z digest=sha256:b52c9356877e2ff574a4ee62a5a5985dbf32ed665387402e06f91a3ba49bddfa

Observation 736416f8-5d2a-4a1a-8b65-af190c10210b · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:21.254510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:21.254510Z digest=sha256:6a2aa13010cc27adeeebceeb065c267d8e50ba74f1fac7d57f9856d77f546915

Observation 2f7b4b42-248e-43ad-ab69-7b0e71a560c5 · inbound

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding cites this paper.

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.170204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.170204Z digest=sha256:1fb7bee37c0f846a1cb5933ba7cf78778b7727f5beabf326efb23bd6ad5e2c48

Observation c4292ea1-2568-4750-ae9c-1109fad454ab · inbound

Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation cites this paper.

Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.470496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:38.470496Z digest=sha256:8d73d4015f95c80b32ed2b1b370fdf139b7343260f9cfd278a6e2580f1c98fbf

Observation 964a464a-ceb5-4326-8e21-cf34afe55792 · inbound

Position: Reasoning After Perception Means Reasoning Without Vision cites this paper.

Position: Reasoning After Perception Means Reasoning Without Vision CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:02.315456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:02.315456Z digest=sha256:0d260dfb5ce73fd393a6de58379067ac872e4a4a886b27d0f6eae2588c2e74c2

Observation 0b2729ac-0c6d-42a0-a2a7-401f20bcb0e0 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.162023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.162023Z digest=sha256:afed9aa3011db91d13f842649e08229187d7e58173cfa62bbc601364041dc8fc

Observation 79d6a524-fbf5-4e88-97de-bc2ff4c5d8ac · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.173395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:923182b343db1cde5140b050c4adef9a4e8de5959b850cd4b6e90e1473c92284

Observation 4a312208-d36e-4512-bc46-8e0e4077c625 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.767369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:8f83cdb49ee42eecd94cc0296888f75ca0d78211621978137079f4f50e6b2e9a

Observation 8875ee5a-2409-47cc-b173-7ab2f4a3f511 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.943512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.943512Z digest=sha256:b0c806d27a7915b485143707a30eb436b32a7f6ed25a2b81ad158caa60a0c328

Observation 5740e854-0131-47fa-9469-3996b278d76e · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.078836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:003f18dbcdeb256d9941c264c810695e50a20e9158119ce8e9c4e39173b3582e

Observation 5ddc59e4-c439-4678-8284-13ebfcc68c12 · inbound

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance cites this paper.

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T22:36:09.992374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:36:09.992374Z digest=sha256:67118b1902eb8592491ba418b370ef3640efca93805b355815a4fe0a498d9d11

Observation cfe939af-8412-4dfb-ab86-9f248e99c88e · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.506437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:c5a2f5fec9fa7f38ac653c09566857836ede89977ff8633fecf01c7bdf64b810

Observation a18699c5-30c6-40c8-910e-6ffe2189aff8 · inbound

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training cites this paper.

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:18:26.614884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T01:15:26.757215Z digest=sha256:4bfe6c2410579b38788897b4e50bd31b4a4948693155404913a8ea2ff8f3e601

Observation 947b5532-e2b7-4993-a8c3-a44d4155f191 · inbound

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? cites this paper.

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.327801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T01:42:54.802658Z digest=sha256:8bcf1492e4ec6cd34a50d0172dc7f5f8ea4ea88a58d6bb5d0776bde62693bbd0

Observation 5a521e15-f43c-4ae3-b776-ccd1164a5765 · inbound

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition cites this paper.

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:53:51.525916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T18:48:45.623530Z digest=sha256:7c51235b08c70b548e9b45289e80995f49bf163aa0e93915e0a33bef42ec5244

Observation ef1ae605-bc50-4308-b92b-e56086257121 · inbound

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models cites this paper.

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:47:41.906627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T13:03:07.941337Z digest=sha256:477412a361d7e00180ea00ea16d5292ad0190de7a9bdb4ba4bce5fab04d85aa5

Observation 1aa6dd25-2830-4ce1-8020-ba4edd373b63 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 131

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.141939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.141939Z digest=sha256:65b30ba22fc233a05d33da64defae8cf25e7c5392c72062f86bc64e7a044eefc