Pith. sign in

Paper Citation Record · LEDGER

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

As of 23 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 18 inbound Pith citation observations for arXiv:2412.02210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02210 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:47:10.303241Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.645575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.904748Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d312f58-a6a5-4c9c-9810-649d1a422ac7 · outbound

This paper cites an unresolved cited work.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:47:11.388706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.019976Z digest=sha256:9293532b2156362cde10bce9730a0def361bd1b8bcc47cd9382b79140ac0dfb9

Observation 31aa5d48-02f9-4e3f-842d-dd6422a2d22f · outbound

This paper cites GPT-4 Technical Report.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.026585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.026585Z digest=sha256:edfc441d78a6fad733e0b428ad5a1e869e9cf1ba5d4161d5ca3da71c236651ff

Observation 2e9d4d9c-600d-42f8-9dad-dcdbd68c47c7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.033823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.033823Z digest=sha256:bfb1d4916ebb016eb5b7e03772bd7e8c0f0b120904432ee7bad550c15285694a

Observation 1ee5e262-a344-4d7b-9077-902b4abb54ce · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.040694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.040694Z digest=sha256:aefa321a19eea956b282875d4406453771af890d106eb3803cf227977c8d132f

Observation 0ccdb98b-af15-4edf-b688-0495748d2e3c · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Onechart: Purify the chart structural extraction via one auxiliary token

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.368604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.047209Z digest=sha256:d89e9a5ae3af1fdb0de869cc7721fd6b6d8d05fd2c0b5fd40fc775e468e9da9c

Observation 85f1cf64-7422-4120-b920-959fd9266711 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.350282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.053297Z digest=sha256:a026bca6e14423f47988e7d8b0d094a2689ba5825a60fda01ed825c353c8cda6

Observation 1b189e01-9fe0-4e59-9b2d-6623439d76d3 · outbound

This paper cites Total-text: A com- prehensive dataset for scene text detection and recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Total-text: A com- prehensive dataset for scene text detection and recognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.330529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.060223Z digest=sha256:20256d2cedfc40c1b31264ce05ef11dd76a5ef4f5228c15349e32f3aea1498ac

Observation 8700ec30-408a-4a54-928e-c9ac28f40d6e · outbound

This paper cites Icpr2018 contest on robust reading for multi- type web images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icpr2018 contest on robust reading for multi- type web images

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.308637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.066507Z digest=sha256:e52a587d359c8171283090568b960e885f335b06b44927213ef208fed4b6b2c0

Observation 6134128f-727b-48ea-b3de-13e1898d02a4 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.073378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.073378Z digest=sha256:0bcdbd95010f4a389e7052e48323bb5f2ee055ea8506d916b899c7a2bf5a8faf

Observation 00f847f6-4254-4297-9a4c-3ec2ff245a1e · outbound

This paper cites an unresolved cited work.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:47:11.290550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.079322Z digest=sha256:490c495b907ad69f74b82d48c5d147cfa829491b992f2ca77981e42a46962477

Observation 9a410f80-8faf-42e5-909b-6e6226005c6d · outbound

This paper cites Post-ocr parsing: building simple and robust parser via bio tagging.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Post-ocr parsing: building simple and robust parser via bio tagging

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.271230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.085071Z digest=sha256:1f2b019a1165effebf84c7682a39a86a6a20256fa8f2239085794449a995d3a0

Observation 15702156-a23a-436d-971a-6e4dfcfcccaf · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Funsd: A dataset for form understanding in noisy scanned documents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.253184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.090339Z digest=sha256:fc56c72824aeef2eff086d527db97178d4298186510dedc7a2b05f195f2f6713

Observation a83c0578-8d9f-49fc-a64c-e6378e62c762 · outbound

This paper cites Icdar 2015 competition on robust reading.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2015 competition on robust reading

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.236087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.095690Z digest=sha256:538962119229ddd7a77880d33c04911dbea030f5a4931c7a1d75de9334b1bbf8

Observation 23ba1a9b-30ac-4852-beab-aa6af8349b5c · outbound

This paper cites Ocr-free document understanding transformer.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Ocr-free document understanding transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.216001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.100963Z digest=sha256:f0e1ea55fe84b4ac865adc0054861aff1ad5cef121d3f115b5e206cb813af177

Observation 77421711-b0f8-4c55-b65b-f8a221ac9727 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solu- tion.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Visual information extraction in the wild: practical dataset and end-to-end solu- tion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.196174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.106119Z digest=sha256:6550f5954780ab79986e0c1197802456ac76cfe2e35017d3cea852ec28f1a446

Observation 0fa47a8c-203b-4719-a464-ed83920f3e01 · outbound

This paper cites Binary coors capable or ‘correcting deletions, insertions, and reversals.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Binary coors capable or ‘correcting deletions, insertions, and reversals

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.176016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.111356Z digest=sha256:7a5cc334a02a47d50be2c35c074a8f17ad6dd02771841ed3ecc54ef60b055406

Observation 4b8f8499-f103-4abe-b669-406e7323761b · outbound

This paper cites TableBank: Table benchmark for image- based table detection and recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TableBank: Table benchmark for image- based table detection and recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.154803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.116371Z digest=sha256:817ebb09bc14137895a81b2e149f4286935e376667d20386278a084a545bd237

Observation 87624657-47b0-4cdc-9cf4-d1c3dea777f0 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.121280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.121280Z digest=sha256:6801d0c5dbd321eb55a1487dcc39282f733faa142b55f539a16866d0674e4df6

Observation 19a303e8-bf54-415e-a658-ed2a8f03fdf1 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.126191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.126191Z digest=sha256:9f90a669531679822a697ec17b8380d2493406eebe296aa95afed7c5b91fe817

Observation e5bd558f-d825-4799-b185-fdde556e1d39 · outbound

This paper cites Spts v2: single-point scene text spotting.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Spts v2: single-point scene text spotting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.136303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.131612Z digest=sha256:7e1e2c18ab8947eca926289e04d1ab063b9c1c44bd470b245c9b2a6bebdf87da

Observation 76f39e67-cb53-499a-8984-cc1e90f7ec3f · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.136168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.136168Z digest=sha256:d59bb21227418b637f9b58edef8f521ee547e75f49689cf734d95886bd934270

Observation 2a484f35-69bc-460f-9ce5-d3e44ed11fe5 · outbound

This paper cites Parsing table structures in the wild.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Parsing table structures in the wild

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.114773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.141463Z digest=sha256:49f986fc846ec956d9336d3cc4064d94c87d1666ec6319f154eb234f35a00f4d

Observation fa1ea388-b6cb-47d1-8f54-40660dbb9143 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards end-to-end unified scene text detection and layout analysis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.095493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.147473Z digest=sha256:37298d453d6dbd7e65c3159d8e4b2c971f12972f77199c8923ce793e006e1d9c

Observation fd6e2449-2c2f-44fb-93cc-a3bb9f6f235b · outbound

This paper cites Layoutllm: Layout instruction tuning with large language models for document understanding.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Layoutllm: Layout instruction tuning with large language models for document understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.075697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.152557Z digest=sha256:710a9376500663a78d4d3e6e0609d0ed0a27757fb59d88e603eeccbbfc490c35

Observation 92df3bc4-0195-4ce9-bdc4-f6412729c6ed · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy KOSMOS-2.5: A Multimodal Literate Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.158358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.158358Z digest=sha256:e5c603cd86861baa0c2ee7ee8ed2a7e89844d9164f6cfb32668eda20e0da36cb

Observation 18dc72a7-fac3-4fa8-853d-ea813e396e4c · outbound

This paper cites Icdar 2019 crohme+ tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2019 crohme+ tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.056054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.164365Z digest=sha256:c81dcec59a6e3d8794aa53f84dff1aa3531f9bbd5ccfe9321743cf2b6f499e3a

Observation b149ca7d-833d-4d4d-a03b-f1f8b2fd010c · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy The iam-database: an english sentence database for offline handwriting recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:11.037443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.170323Z digest=sha256:d1e76ae84e71fc860d8ac1cdaeb186ad74eb98c736f7a6fcc480651310bf4cc2

Observation 8abdbd23-e6b1-4282-9347-94fe9b313cca · outbound

This paper cites Docvqa: A dataset for vqa on document images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Docvqa: A dataset for vqa on document images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.176206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.176206Z digest=sha256:6b25dc051e2f050db965866d09276b4fcf62e5f18d9a2e17adf2480b5fb39665

Observation 769ae03a-6242-4162-92e2-935b755750e9 · outbound

This paper cites Scene text recognition using higher order language priors.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Scene text recognition using higher order language priors

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.999313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.181921Z digest=sha256:a9611f6236cefafce50d8e774b37d54c7ed07de85cc97b580285385f9908a727

Observation a766abb9-13c1-4085-8164-5ac46b60b34b · outbound

This paper cites Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.970598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.187350Z digest=sha256:04ac0018a4f627914de0e1067207a98559ad3056db8328cccff71d705163e9c8

Observation fe6323e2-5fdb-4174-a558-21e2c660c109 · outbound

This paper cites Cord: a consol- idated receipt dataset for post-ocr parsing.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Cord: a consol- idated receipt dataset for post-ocr parsing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.944130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.192956Z digest=sha256:ba0e29d53dbccb671c89cdaac813b82b897ba97f1c94ba12c6d8201dd446efae

Observation 1b6993c0-e907-459c-9b38-e425299e7a76 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.198930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.198930Z digest=sha256:3eb7b2d052618640bdfdc446329dbbb7013ea65eb02734f85b60f087d23a40b5

Observation 0cb5d40d-e1cc-4fcc-ae6a-e02835439dff · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17).

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar2017 competition on reading chinese text in the wild (rctw-17)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.893043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.204545Z digest=sha256:4f010a398a50215a020ad7016406fbc39413f6b49b4718300a9ffbd439b6362c

Observation 5f0525b2-d90d-4968-ab5d-d7a0784162fd · outbound

This paper cites Towards vqa models that can read.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards vqa models that can read

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.873328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.211029Z digest=sha256:c37170660671a0b78d8c93d5884e0673d8f2961903ad7e88a44955a4c1b81061

Observation 1ce07450-1c4a-4083-8ba2-6ee9a0e841f8 · outbound

This paper cites Seglink++: Detecting dense and arbitrary- shaped scene text by instance-aware component grouping.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Seglink++: Detecting dense and arbitrary- shaped scene text by instance-aware component grouping

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.850688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.216475Z digest=sha256:7af2acb929506c65682fb716caa941e62319e9457848a9e47a35dab85490ffef

Observation 9378978c-3f0c-47be-9c47-877ffb7c6e69 · outbound

This paper cites Mtvqa: Benchmarking multilingual text-centric visual question answering, 2024.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Mtvqa: Benchmarking multilingual text-centric visual question answering, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.831095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.224158Z digest=sha256:9b217284d8bbf8e249c046f9fd894d0fb0ed22b0030f9724b10ad33cfe4f00f0

Observation 22ba9c20-d406-4224-91c4-1a0f4c486bd3 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Unifying vision, text, and layout for universal document processing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.810874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.230406Z digest=sha256:74df1dabd5507496ea8660115482cf3e467ba7a6cdc513cf918160378ce6a81d

Observation 9908f796-e438-499c-aaa8-a800c6c338f2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Gemini: A Family of Highly Capable Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.236564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.236564Z digest=sha256:30d8c162f34cb521ed98f2efd02101fb3b03aad5d2164724fc151b6d37e8cde4

Observation 3feefc3a-554d-41bd-b8ad-2f0155313e24 · outbound

This paper cites Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T23:47:10.480277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.243109Z digest=sha256:849104f14eb8bacbfb48de68b8bff5db0e008ca532026f34b31f1148fb5c2cbd

Observation 857d57d6-d5bb-44be-91f4-ced6d5d97b6b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.248636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.248636Z digest=sha256:d5fabceb72f5324642cd4909d46cc79920dfa4142f0eede3e6277b032407828c

Observation f1ddf2ba-4fa9-4ea4-a4de-3b59de155505 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.254073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.254073Z digest=sha256:0347aaae649be5383c35fb7f2072a31c9c915c76417ade8e71a5dd2c5b6bd5b7

Observation 1a4e9fe0-786e-4466-94ef-e64971c3a63e · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.792916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.259956Z digest=sha256:6d836d0d1ce15d651ab38736a611c320438f0bae474d83fc7305c8401778ff27

Observation 13e1739b-7b3f-4fa0-a0e7-d06b2bf93856 · outbound

This paper cites Modeling entities as semantic points for visual information extraction in the wild.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Modeling entities as semantic points for visual information extraction in the wild

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.772300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.264455Z digest=sha256:95210d8afa618654732cc4544547e934aa9ddebc8e4bc9964370e07f1d931d30

Observation 21a08f2e-8c81-40b5-a8dc-1ed1bb00eca0 · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.753010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.270597Z digest=sha256:4efb578ba8e6f1302417a5b6935da37fef3b5ca57a474e384d9986c94c45fcf5

Observation 97bd656e-09d2-49df-b599-6dab9842d4cd · outbound

This paper cites Icdar 2023 competition on structured text extraction from visually-rich document images.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Icdar 2023 competition on structured text extraction from visually-rich document images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.735243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.275337Z digest=sha256:1d0d5af962c4dd37074e5298389dfe44a3fb7dabb8409b252d65c1895a017918

Observation f460d0ff-ae55-4d5e-9bba-2f682d51c8b9 · outbound

This paper cites Syntax-aware network for handwritten mathematical expression recognition.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Syntax-aware network for handwritten mathematical expression recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:47:10.716552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.280683Z digest=sha256:4451844776cab2da68247768d7a5d5a495e6fe589aee3b421b343e55d6520347

Observation c139026f-6b05-4428-96e3-01332c86666b · outbound

This paper cites Detecting Curve Text in the Wild: New Dataset and New Solution.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy Detecting Curve Text in the Wild: New Dataset and New Solution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.286185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.286185Z digest=sha256:01eb918a3d229e5bae30c85d8c83c6b6468fb890aed319233c6e6e703cc134d6

Observation 22305401-e979-41f9-b576-29c2070a894e · outbound

This paper cites TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.291965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.291965Z digest=sha256:3c0014a973a4756820f9d959f9f7399b0f2e166dde94535deccc4935145e3b15

Observation 9018851f-08b9-4ffd-ac40-e4d3fc27a634 · outbound

This paper cites DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:47:10.297668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:47:10.297668Z digest=sha256:97f1a08f274ab0d671c94fbb9de5d690fb2d33e3684e6a72f4deaff819f56eee

Observation 9edd0b2c-ef4c-4f1d-8e9b-c4918a5ce15b · outbound

This paper cites 3.5 kg”, the true value is “3.5kg.

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy 3.5 kg”, the true value is “3.5kg

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T23:47:10.696214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T23:47:10.303241Z digest=sha256:babaa22eb5d978ea0068c7d924de524d6fdb65abcc41fd4ca082c03613123e4b

Pith citing papers

Observation 44a67d08-3b66-41b7-a987-03b28e41f3aa · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.764403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:73b8fc4ad1b181453ed77db65faa02b15a074bcc109d717f227bf9ab5ca693fa

Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · inbound

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? cites this paper.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.645575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.645575Z digest=sha256:8cd1148e892ebd39fa47b5c7c4e163580bc18949cf625efda342900d42588f12

Observation 736416f8-5d2a-4a1a-8b65-af190c10210b · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:21.254510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:21.254510Z digest=sha256:4bbdfe4b409b6df20b0ef9c0c27958fdd952523825e4ad1f85c971391055c8f4

Observation 2f7b4b42-248e-43ad-ab69-7b0e71a560c5 · inbound

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding cites this paper.

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.170204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.170204Z digest=sha256:5d1da966a77cd4d91c45d8bb16f6ba9368a7a501e045753d483a44315a123c3c

Observation c4292ea1-2568-4750-ae9c-1109fad454ab · inbound

Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation cites this paper.

Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:38.470496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:38.470496Z digest=sha256:9b4f64723ab1e226d106cce83caa12513db7e1bece32c88748540fa3bf5f67de

Observation 964a464a-ceb5-4326-8e21-cf34afe55792 · inbound

Position: Reasoning After Perception Means Reasoning Without Vision cites this paper.

Position: Reasoning After Perception Means Reasoning Without Vision CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:02.315456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:02.315456Z digest=sha256:db2ab5a9339156e25e682f339e7cca8dadf75dc43916bda19ae79ee36a388c8a

Observation 0b2729ac-0c6d-42a0-a2a7-401f20bcb0e0 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.162023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.162023Z digest=sha256:c91aa2ac3cb593a50aa92b4882f01666f9933461420d8a9cb4aef1419cf284e0

Observation 79d6a524-fbf5-4e88-97de-bc2ff4c5d8ac · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.173395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:fcb9b148624396c63cb917d5f694b449b4a54b23d8a6d28ca1be5ec853c8b5fc

Observation 4a312208-d36e-4512-bc46-8e0e4077c625 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.767369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:5ae2b4c0dc711a2ff0aa06123e6c875328b92a4d023d04616b903f9141659b5c

Observation 8875ee5a-2409-47cc-b173-7ab2f4a3f511 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.943512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.943512Z digest=sha256:710e78e984d53c742bae6fc8e3fa7bee5ea6b24daf1fffe8f84ee8ddd7311493

Observation 5740e854-0131-47fa-9469-3996b278d76e · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.078836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:fea5130d7e244ada69259cf75bcf11aea1b118eff5dff50ea97b7d448f28e75e

Observation 5ddc59e4-c439-4678-8284-13ebfcc68c12 · inbound

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance cites this paper.

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T22:36:09.992374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:36:09.992374Z digest=sha256:cfc9bf25c58102ec19e5d5869c1563b7fb0bc8ccc544cba177ffd97ccbcc7a54

Observation cfe939af-8412-4dfb-ab86-9f248e99c88e · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.506437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:f99ee89a9a4cc95e5deb107aa617abec716c18714bf013196e7bf70ca3389a32

Observation a18699c5-30c6-40c8-910e-6ffe2189aff8 · inbound

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training cites this paper.

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:18:26.614884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T01:15:26.757215Z digest=sha256:06afc637d609fd424e3d071fde702f21b2d6e1b09617b7c22b609397db492bba

Observation 947b5532-e2b7-4993-a8c3-a44d4155f191 · inbound

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? cites this paper.

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.327801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T01:42:54.802658Z digest=sha256:6956fb8cdf65fcd3a1a5c7a4f717461bbd6569040e490ec0adda2ef07d13c2ba

Observation 5a521e15-f43c-4ae3-b776-ccd1164a5765 · inbound

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition cites this paper.

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:53:51.525916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T18:48:45.623530Z digest=sha256:2b06ec88a78a5eb58322a9d305c5d7242982539667cf7af1c4f7bc0086401914

Observation ef1ae605-bc50-4308-b92b-e56086257121 · inbound

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models cites this paper.

Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:47:41.906627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T13:03:07.941337Z digest=sha256:08d1d79c532afded74e59c562263fc2b6508ba79194c7bd3c254948997078135

Observation 1aa6dd25-2830-4ce1-8020-ba4edd373b63 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 131

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.141939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.141939Z digest=sha256:070f5700131ba463d23bd3800047d608afdcc7192c878d3c3cc34040a3a5e044