Pith. sign in

Paper Citation Record · LEDGER

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

As of 7 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 0 inbound Pith citation observations for arXiv:2607.11562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.11562 v1

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T04:45:32.682508Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 137 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 956487cb-9bff-4207-8694-19f36768c8f5 · outbound

This paper cites Qwen3-VL Technical Report.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:80e22495fd67006645c0e46135c30ad21994f02c9ff10831649141f35dc58a3c

Observation db94fda7-7ff5-482a-8bfb-7ca78c7587a6 · outbound

This paper cites Qwen2.5-VL Technical Report.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:7e556aaa59601d8a28be9b22d1d07a8e8bdda6d208692ec8dd3335a78df622ab

Observation c8714cae-c14e-4450-8b26-9b73b2f205d1 · outbound

This paper cites BEiT: BERT pre-training of image transformers.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI BEiT: BERT pre-training of image transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:fa3f190666efb2fb627dc9e981ec46fe743b8217170f2911ec7d9a6e868276a2

Observation 9fe81157-476a-4f73-9f90-ef93006201a7 · outbound

This paper cites Scene text recognition with permuted autoregressive sequence models.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Scene text recognition with permuted autoregressive sequence models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:6a0d1cc68396231403095c40f83c06ac0c5467360e7f791243a871422e257cba

Observation f0b88761-7919-4e83-bbda-a6ca5d329c8f · outbound

This paper cites Nougat: Neu- ral optical understanding for academic documents.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Nougat: Neu- ral optical understanding for academic documents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:e25af4432f42e92066b3d193dcf1a604ad9c622c036961231146b72bea6f9352

Observation 63566e54-418f-4025-9ebd-a50ec41a9218 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Emerging properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:282d72f3b5d94694242d9f8d724e5bc6633d21c3aad8d4a202aea58b5f6018ec

Observation 5570681b-8f6e-405a-8024-858bdeaa5c01 · outbound

This paper cites Encoder-decoder with atrous separable convolution for semantic image segmentation.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Encoder-decoder with atrous separable convolution for semantic image segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:1ca3a930b6edb10fad2484dce2200f533603ada87f6367944930cb345464bc46

Observation 7d445a3c-6f95-4497-96f7-92168485031f · outbound

This paper cites Enhancing tampered text detection through frequency feature fusion and decomposition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Enhancing tampered text detection through frequency feature fusion and decomposition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:ad0e8680aaf36a4069a3372631b3f0f03e30a8fcd022ef5fe5c7a881ce4aad14

Observation c69d1d07-5d04-4bb0-82ae-e9562d2c703f · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Masked-attention mask transformer for universal image segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:4089ae1e40efe769dcaf1853ab5d548db9769d642177c3bdda9e1f30ba9e99e3

Observation 9e94e238-eb65-4307-aa30-b300977122f7 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.Advances in Neural Information Processing Systems, 34:17864–17875, 2021.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Per-pixel classification is not all you need for semantic segmentation.Advances in Neural Information Processing Systems, 34:17864–17875, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:f213117398e34f3481922d1ebbad53761a65fff248790c51ee7557b4567dead0

Observation 15c99d84-fa77-4cec-9adf-5b88bdbe8e62 · outbound

This paper cites M6doc: a large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI M6doc: a large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:0fa6f372952242f34b794155d925807be0439862c90e5b5dfd20b3f2a859c93d

Observation 0765eb9d-c493-48ba-94a6-1fe8a336c462 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Reproducible scaling laws for contrastive language-image learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:6d416b911bb7269a0a6b2b6e873516a01c06ada1347a052104ef6066c4cb6187

Observation cf0b5cba-23cb-4d6e-90b0-c82d2a309265 · outbound

This paper cites Total-text: A comprehensive dataset for scene text detection and recognition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Total-text: A comprehensive dataset for scene text detection and recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:6c3e7f2e9255b52e12776d0dab2e54c9215c9a907fecea395553ef2a5e4c9b00

Observation 4de5f807-5022-456f-a1e0-95b72d9cd0e7 · outbound

This paper cites Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:958b2b5e8f06b21649784e63f8ce512c37d8271641c815db43f84010d35bce52

Observation 8786798c-eef4-4a2e-93b1-922796038491 · outbound

This paper cites PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:7992dc47211d6074e69e7faa67df9616464e1db10b1ebb41a16966525c8cec11

Observation 5483938f-14b9-42b2-a50f-e0a1d955a917 · outbound

This paper cites Boosting document parsing efficiency and performance with coarse-to-fine visual processing.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Boosting document parsing efficiency and performance with coarse-to-fine visual processing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:a570aba3b8abd9efe1891d813daed36674b8b1da9975b198a7c78c898857eb54

Observation b8adad7e-7a38-497c-b2d9-b59a41ea8126 · outbound

This paper cites Vision grid transformer for document layout analysis.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Vision grid transformer for document layout analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:b73d35aecba4eac850fe02e69c89abec427bc5184e069283c8515630c846db08

Observation 9b558c00-1ee2-4efc-adc4-cb9060208712 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Imagenet: A large- scale hierarchical image database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:9572e5906bc0fa463ecdb3d8b04df5f76c5b5ee4f8dfbc2b5c4e462273dbc40b

Observation ea56ecce-d9ee-4c57-92ec-f3911336f702 · outbound

This paper cites Decaf: A deep convolutional activation feature for generic visual recognition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Decaf: A deep convolutional activation feature for generic visual recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:1eced912c0737565d3e8418502d4e8ab47619868714039aa6fafc382b274d99b

Observation d4f1d11a-8712-47a1-b29b-a56f85f0bd87 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI An image is worth 16x16 words: Transformers for image recognition at scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:519f49bcbdc216393f21f329ef5fac132b375cd004aaa54bb0768cc07ade12ea

Observation 6b5cdcfa-9fb6-498c-b13c-db94b6b22c5a · outbound

This paper cites Out of length text recognition with sub-string matching.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Out of length text recognition with sub-string matching

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:8936faae57ef39e8585fe9bfcd49dbe74966f5b1b7541f1fcce3b5dbc27a3d40

Observation 5e982cc3-5d97-4f1b-87a8-852587c337ed · outbound

This paper cites Context perception parallel decoder for scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Context perception parallel decoder for scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:6ba1eaf58827ca960e7c7000da80112ff5deb61807bfba1051d5f9b3c59e63a2

Observation d4194c98-d26d-4048-95c1-a77fd25bef09 · outbound

This paper cites Instruction-guided scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(4):2723–2738, 2025.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Instruction-guided scene text recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(4):2723–2738, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:b72c31b66d0ca6c4f65182c220544fe83d07477df1a36f99cb9e2e8344f9f2dd

Observation 2cf2c263-9db2-4b6b-99c9-7cacee222f4a · outbound

This paper cites Svtrv2: Ctc beats encoder-decoder models in scene text recognition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Svtrv2: Ctc beats encoder-decoder models in scene text recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:0f3203d605f6e71c155ee9b3e2d7a4555838b1d1882659d6bedc83e8593eb138

Observation 59b938e5-6407-4d77-9935-3a041565c732 · outbound

This paper cites UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:90049193164356f76012dd1f4e3bbc609de6662b2a41576f29e379a08437f4a3

Observation f145dc73-0f90-4780-8f54-b3f2ef79eeb1 · outbound

This paper cites Glm-ocr technical report.arXiv preprint arXiv:2603.10910, 2026.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Glm-ocr technical report.arXiv preprint arXiv:2603.10910, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:fd8181b38e474e75c4e2cf5c38bd38792139192de5ce2340fd8e638106f146fd

Observation fff5c9ab-7ed2-48b1-967a-b8c2f5dde944 · outbound

This paper cites Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:d50d86cad20e3deef826dcb84c1ee23d0c4fa5e327e64d630924b212b36f13a6

Observation 47b6db9a-c553-4f67-9059-d31968fd052c · outbound

This paper cites Mathwriting: A dataset for handwrit- ten mathematical expression recognition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Mathwriting: A dataset for handwrit- ten mathematical expression recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:4dcb8014efa4be899c7cb94b6c1a6cbd5aef2e9a379d4d8d08c9fad425779240

Observation 7b74a901-5379-4562-94b5-d6ad3911972e · outbound

This paper cites White, Silvia C.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI White, Silvia C

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:c02ae30433848b0197d5faf76b7f119de34fd263165ac637d6ea7af5bfb7be57

Observation 084acf49-3b43-4494-9446-2e30697605a2 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:7859e6779681ed4fcb7e81a3e8f947a573275e3497805ee44bc7f152b18ec53c

Observation af881e44-1710-45ab-88b6-af1e529bfda2 · outbound

This paper cites an unresolved cited work.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:858c63a5f600dd91de7b1bde30ac3312b471c0a50a090bf69cf827d55faeaa25

Observation 3ba7009f-7e4a-48f5-802e-b42fa9396dca · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:afececd5a5e191dd40f9a597c96a61ec6e1eb0e2852f9ac46fa29e742597f4bf

Observation 079ba129-1e5f-476b-9197-4c523e07b507 · outbound

This paper cites Speech recognition with deep recurrent neural networks.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Speech recognition with deep recurrent neural networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:f5e0c3009ca2423f0cc7dcae709c31233b4cf21c58efb259cf6f089c6abb1b18

Observation 924b39b7-45ae-4723-9350-9dbe46eeb697 · outbound

This paper cites Unimernet: A universal network for real-world mathematical expression recognition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Unimernet: A universal network for real-world mathematical expression recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:76889ecc2fb5798c037750539ef106680159c1e4c2a0e7688fe42de5dedd9d43

Observation 208a16f9-688c-4720-bdb5-cafe636234f4 · outbound

This paper cites Deep residual learning for image recognition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Deep residual learning for image recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:10176d33ae5ac056e8375b7691add280f5a6e2c039b625c1cf59796e3528e203

Observation c11203c2-00a3-4f30-a964-7f5266998a88 · outbound

This paper cites Icpr2018 contest on robust reading for multi-type web images.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Icpr2018 contest on robust reading for multi-type web images

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:c57e5e92f3f50c290d506c0727f6e71cb6eb42f2b5fa4046d25e744cbfe423e6

Observation 5444bede-46bd-447c-a75a-c5d8af7bf177 · outbound

This paper cites Radiov2.5: Improved baselines for agglomerative vision founda- tion models.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Radiov2.5: Improved baselines for agglomerative vision founda- tion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:c4a1e5eeab0ecd94bcb3ecebca0aef9615bafd697f34c2c07a1507117990aca2

Observation 412ae612-9dc8-40b0-a639-0999e1c5fa16 · outbound

This paper cites mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI mplug-docowl 1.5: Unified structure learning for ocr-free document understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:20ff035e5326a1c201d2176c91c86d42332709f0fb173211e9efc719e8318325

Observation 61516f06-91f8-4f28-844d-dc6468b132b1 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:7a9ad051d64a1b0bbd4c180ebdb830ed0e3b882fbe0268c4145811429afafd86

Observation cc7178e2-3580-4242-84b1-5f5fa48438c4 · outbound

This paper cites Revisiting scene text recognition: A data perspective.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Revisiting scene text recognition: A data perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:3c0629fa46b2c7058bef70d1c5aa80d0253bedc1f8baf4d09215808d5d8075d5

Observation 66fabf0d-d267-4115-9e41-5301b98f32d8 · outbound

This paper cites Icdar 2015 competition on robust reading.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Icdar 2015 competition on robust reading

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:4c92e859c0ac0a7651559f2b7188f20025c5b2d326fda21c9a870c382910c299

Observation 3277cd3a-9c04-4a71-8da4-7d9890a273ea · outbound

This paper cites Icdar 2013 robust reading competition.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Icdar 2013 robust reading competition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:5c88ea43f769e2d4ac9b9b61c061440bc4e9c6f9fdf218f324f00f61d286b7fc

Observation 7e7f5bc5-1c8c-4735-8a4c-e6eef1e25f3d · outbound

This paper cites Ocr-free document understanding transformer.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Ocr-free document understanding transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:b4adf31876e0da53c5a657cb25c0068a2bad1b2c172be484f02884f90b6e748f

Observation 2a2c6ce8-c79e-4450-8adf-cec8c915efe4 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:5911d45f0dd3917bead0620b6e862319f9e190d79b6c44a7ec4fa55b4f2692a3

Observation 7a008935-daa7-4bdf-b6f4-02a4e774caa9 · outbound

This paper cites Open images v5 text annotation and yet another mask text spotter.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Open images v5 text annotation and yet another mask text spotter

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:4be1f055ae3713115e641f4703b5ae56ba27f3ea0470d4bdda1cde3cef4e40f8

Observation 18557dce-9ea7-425c-9c58-b4f7c57eaf36 · outbound

This paper cites Cat-net: Compression artifact tracing network for detection and localization of image splicing.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Cat-net: Compression artifact tracing network for detection and localization of image splicing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:2528da3ce2fcc32dc7c037bb625a35817d51e79a23094c95e8a32a7eee4ab7d6

Observation 564d561b-e063-452a-822f-368e60237fd4 · outbound

This paper cites Towards better structured and less noisy web data: Oscar with register annotations.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Towards better structured and less noisy web data: Oscar with register annotations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:5ca47df20848f44623656b75c797f352d2f95e768c62c0c17d9b769a4864e089

Observation 2ded9b07-7271-4dd3-a061-49d38b8e186e · outbound

This paper cites Pix2struct: Screenshot parsing as pretraining for visual language understanding.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Pix2struct: Screenshot parsing as pretraining for visual language understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:30296c1c1c4e4eebb905ddbb0bde3576fa2a2de8474242e414c7fc0124930ed8

Observation 6b25127e-d632-4c5f-9459-162906b9b441 · outbound

This paper cites Building a test collection for complex document information processing.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Building a test collection for complex document information processing

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:e144e5cc20892f46ed6bb066469059d6ff6b1d3203e174d5ad0a2773854932bc

Observation faa64d08-d7f7-47e4-97af-c5ac7fa7ff1d · outbound

This paper cites HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:6ec16dd513d8076240e835b0d82fe998106b2830018e9bf1893caa86f17a1857

Observation 8b4844e4-6028-4cf6-a810-4a09f07cb554 · outbound

This paper cites Dit: Self-supervised pre-training for document image transformer.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Dit: Self-supervised pre-training for document image transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:f37053ae148e16a665f9a50385b1922284e09f8d4bfda7c70151be344dcaf09c

Observation a596092e-5e08-4b02-9ce0-75614f61e452 · outbound

This paper cites Trocr: Transformer-based optical character recognition with pre- trained models.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Trocr: Transformer-based optical character recognition with pre- trained models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:d6ac3945aee615a16b24fb192c8faa0815209d679948b1a118687ca89db54c45

Observation 17d8b603-2104-4512-a017-df91a66ce365 · outbound

This paper cites Openvision: A fully-open, cost- effective family of advanced vision encoders for multimodal learning.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Openvision: A fully-open, cost- effective family of advanced vision encoders for multimodal learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:9aba463ea5e5342ee51b49ba882d7ce639228f428d337c89988fcf0f3ee52815

Observation 972ce4b4-7d9b-4c2f-8d84-41cb11d90c3f · outbound

This paper cites Exploring plain vision transformer backbones for object detection.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Exploring plain vision transformer backbones for object detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:1e81fd89ec0c00f16348110ec681803710fdd73561fc55f6a3c3c5b218593f53

Observation 75c1fa4d-8313-48ff-aecb-966bc386a37e · outbound

This paper cites dots.ocr: Multilingual document layout parsing in a single vision-language model.arXiv preprint arXiv:2512.02498, 2025.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI dots.ocr: Multilingual document layout parsing in a single vision-language model.arXiv preprint arXiv:2512.02498, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:281c8bc2fe7d810c31e71200dda047acb7e18caec97d607d0c4af6beeedc9a82

Observation 1a3d7953-d21a-49ba-b273-52d579d63f9b · outbound

This paper cites Mdpbench: A benchmark for multilingual document parsing in real-world scenarios.arXiv preprint arXiv:2603.28130, 2026.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Mdpbench: A benchmark for multilingual document parsing in real-world scenarios.arXiv preprint arXiv:2603.28130, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:b9d8c7cc4f540fc918cc446eac1b01f4a79d5ca33f2d62ca8dc4e519f5817623

Observation 6227d56c-f8fb-4f99-8f2b-58e87f72a8a1 · outbound

This paper cites Monkeyocr: Document parsing with a structure-recognition-relation triplet paradigm.arXiv preprint arXiv:2506.05218, 2025.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Monkeyocr: Document parsing with a structure-recognition-relation triplet paradigm.arXiv preprint arXiv:2506.05218, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:1e9114c972a9dde512aedd7d8a4e3a9612859543af69e6911bda019bd8451ba3

Observation 0eb68467-df77-4ba9-803f-9c9d34794062 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Monkey: Image resolution and text label are important things for large multi-modal models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:063d2511b42a85d6bc21aea8361d5bfff3450701575cfbd184ddff55d4fd00a4

Observation 144f30e9-a5b7-4326-88a3-24dcb77487ae · outbound

This paper cites Real-time scene text detection with differentiable binarization.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Real-time scene text detection with differentiable binarization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:2ce454449f3c6fb25539885cb872cf40a171abcbfdb694eac5318e33a9286a61

Observation e2980ed5-5696-471b-a537-96f3857086ed · outbound

This paper cites Improved baselines with visual instruction tuning.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Improved baselines with visual instruction tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:40b997a3415269932158383f98fa58989c4e99248a1b9435e8d0b758e9fa62b8

Observation 6d9657ec-972e-42d9-b38a-3dc965cd5665 · outbound

This paper cites an unresolved cited work.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:1b36485440a858890f37bfef421b36df5279695ac9bcfd396d653c1bc21006f8

Observation c7491a30-bd47-4025-b790-957ef6961e93 · outbound

This paper cites Multi-scenario overlapping text segmen- tation with depth awareness.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Multi-scenario overlapping text segmen- tation with depth awareness

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:6e6183eaf59a87c518f44d431dd5d154b26e00f365a1bc41bedb0a31a2601b16

Observation e298b4ae-4a98-4dba-93b5-159648ec8af2 · outbound

This paper cites Openvision 2: A family of generative pretrained visual encoders for multimodal learning.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Openvision 2: A family of generative pretrained visual encoders for multimodal learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:59314689baab612c7f5000859f6e4fddaadb1587057f8ba238217213faf24487

Observation 9daeb8bc-1982-4271-b4a8-c63c5472b6c1 · outbound

This paper cites Multilingual denoising pre-training for neural machine translation.Transactions of the Association for Computational Linguistics, 8:726–742, 2020.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Multilingual denoising pre-training for neural machine translation.Transactions of the Association for Computational Linguistics, 8:726–742, 2020

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:9337cce71ba2485799c8e7093a04db40638141467a3352cb30c9c9fb3f97b135

Observation 42d42d74-75d9-4113-a0c0-3fdcdee9a47d · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Curved scene text detection via transverse and longitudinal sequence connection.Pattern Recognition, 90:337–345, 2019

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:733cf7fceafe8f039d0054bb653a283c514af1c834498fd56f833ece98477270

Observation 57fdb7dc-d7e1-4feb-a155-8d432855c425 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:6047387a840c5a63d66d14f8752811efc9a14957dfd90e1c9441bfeef77cdc87

Observation caf6b75e-cab9-4e8e-9984-09ba86f7f41b · outbound

This paper cites Textmonkey: An ocr-free large multimodal model for understanding document.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 48(5):6008–6019, 2026.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Textmonkey: An ocr-free large multimodal model for understanding document.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 48(5):6008–6019, 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:77de56797bd0cb0815afe5101ce88bb56bb514d562b956c865dd4ee714887a21

Observation fcaa1f93-3a10-4c4c-b150-b04ba9dd335c · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Swin transformer: Hierarchical vision transformer using shifted windows

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:02bda2a3b4c32146e8be8c5b5c5cceb1edc7508758f4f22a243437fedfcc44c6

Observation 5f565518-7516-45c3-bbd0-00ff66612847 · outbound

This paper cites A convnet for the 2020s.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI A convnet for the 2020s

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:aac8e90bc2c7ad6b97e41d11237c85f2da412f6844f3b50bf3ca636cc480bb38

Observation b336712c-f93f-4e92-9be6-207c7451ff63 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Towards end-to-end unified scene text detection and layout analysis

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:26892445568c7c71aceb81de55fc633e68d15363585584e1d432f5183b5d87e8

Observation a177eb79-54e5-4330-a81b-ef6bcfe04176 · outbound

This paper cites Toward real text manipulation detection: New dataset and new solution.Pattern Recognition, 157:110828, 2025.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Toward real text manipulation detection: New dataset and new solution.Pattern Recognition, 157:110828, 2025

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:c8378da1695b084e89c86f998fa1b601940cebffe92b4a537d0cc858e3b01182

Observation 027643cd-dad3-4d0c-a59c-2d95b2fab0ba · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:f8006612070f28075bca60d7e90577459feafe3109df3656d803057cc1ab64bf

Observation 726ce31e-9039-4cdc-8a26-46d5f706e084 · outbound

This paper cites Infographicvqa.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Infographicvqa

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:54f0a36e801703c0dced3e3afa9612182c4becb94e52481b7b453e3a82856890

Observation 52273c3b-b0fe-478e-8f44-b37eea290133 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Docvqa: A dataset for vqa on document images

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:9dfc8dacc4c3685f930c38cae1bbcaf4c2c6095a61e848747a5e97319c38dab3

Observation 486f3503-e9fa-4868-963f-12630bd07d44 · outbound

This paper cites Scene text recognition using higher order language priors.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Scene text recognition using higher order language priors

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:67c523e62477d55e968f73258a8159da9847c5187b12e76a2c166b82bd7cc5d3

Observation 56fee384-a9b9-465d-ae78-1fde29b31db2 · outbound

This paper cites Mineru2.5: A decoupled vision-language model for efficient high-resolution document parsing.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Mineru2.5: A decoupled vision-language model for efficient high-resolution document parsing

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:cb0e86eae64374c10ca97609f08c16ad2ff6573a377cadb3594737e3831edaa2

Observation 70d33beb-373c-4ec1-866e-3aaef0c7c233 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI DINOv2: Learning Robust Visual Features without Supervision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:ee711ad8ed8d96bbc408ade7e2aa7ed962c39c326dfd96875bc868e10da78b0d

Observation 850e8f6c-6ad3-4b18-bb0f-7bc8f9d7b94d · outbound

This paper cites Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:7b1084ae7107e6936f735dd6bcec75d24e5ae840be50554478582f4497d301dd

Observation 35a7642d-62b0-4fcc-a69d-e83b270aa870 · outbound

This paper cites Compositional semantic parsing on semi-structured tables.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Compositional semantic parsing on semi-structured tables

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:d636c9f92535b9815d01262c1f7512fc6cbef3e93f04ee8788eedd967cd3e73c

Observation 064ab8cd-112e-4251-b2a0-726ce6ab182a · outbound

This paper cites Doclaynet: A large human-annotated dataset for document-layout segmentation.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Doclaynet: A large human-annotated dataset for document-layout segmentation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:c57d85dcf1605be12c94c9a578834aae38e6f0228ae9573c965f89afb3f5cb89

Observation 4b0091e8-aa73-44c7-a41d-f62da0244813 · outbound

This paper cites Recogniz- ing text with perspective distortion in natural scenes.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Recogniz- ing text with perspective distortion in natural scenes

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:0e13b7237913ee2d358211fd301e02c79db4ff614e4e5ff58d110cfe8ac2216f

Observation a4f5d735-11b3-4965-ac60-9b86b1252f34 · outbound

This paper cites olmocr: Unlocking trillions of tokens in pdfs with vision language models.arXiv preprint arXiv:2502.18443, 2025.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI olmocr: Unlocking trillions of tokens in pdfs with vision language models.arXiv preprint arXiv:2502.18443, 2025

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:60a9b7e39d54c63e2e3b2533a34ea84ea6150023d61a376f173c233bfcd56135

Observation b3a23569-f991-49a6-900c-d324df1eeefe · outbound

This paper cites olmocr 2: Unit test rewards for document ocr.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI olmocr 2: Unit test rewards for document ocr

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:20cd9598809f7b5218825e68b13bc749758ae1aab80f443c1da99be3a814574a

Observation 2a064dbd-0d0f-4f43-9adf-068886e97a4d · outbound

This paper cites Towards robust tampered text detection in document image: New dataset and new solution.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Towards robust tampered text detection in document image: New dataset and new solution

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:270a5c3e6e78bb26d4bbed3f497e0f74cafa12e0489dfa1d0788e775487edc20

Observation 392ca933-f89d-43f3-a82b-618558fdfbad · outbound

This paper cites Learning transferable visual models from natural language supervision.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Learning transferable visual models from natural language supervision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:08a49eb1d153e75a664df9419f615b9b99044310abdb8dd9423d9addef584e12

Observation 91df2e0b-c7d5-4bfa-921d-57421cb11d10 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:184e3c8841d6763ab71b97072a2d0500bf1d816c5ed9a9393d3045992ed412c3

Observation 2733783f-b31b-4c71-a4eb-2cb09c86fb7a · outbound

This paper cites Sam 2: Segment anything in images and videos.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Sam 2: Segment anything in images and videos

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:0d4d07e223d3b147b562c2721cf5fd4ef94938b434a9e57d81e6f11ef3487646

Observation c54722b0-ffed-417f-b052-83957b64419f · outbound

This paper cites A robust arbitrary text detection system for natural scene images.Expert Systems with Applications, 41(18):8027–8048, 2014.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI A robust arbitrary text detection system for natural scene images.Expert Systems with Applications, 41(18):8027–8048, 2014

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:394e66434792316cb1b20c80d2d165f2666d43542fb217a0e3cb01376b550c6a

Observation a1fe7fa5-869c-42a3-be32-390deda8a3ba · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI U-net: Convolutional networks for biomedical image segmentation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:b8c0bf6e6d40cb4be176c87f7a9a4c52352e32667effacd5807e54db288ada99

Observation fda23e05-0668-4144-8aa4-635402b19c27 · outbound

This paper cites Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:60b33e1b8eaba620a1b1fa4cb398d1c402d97b7377d1314c2a4f9dae165a47a6

Observation 95f0fb92-a295-428f-a230-5edf53b252dd · outbound

This paper cites an unresolved cited work.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:4eaf0e25ba74ab951f14ad2fb1eda829c4d91409a167221db07b77d9928350ef

Observation 986a6732-752f-4795-84d8-61b6e2cc50cf · outbound

This paper cites DINOv3.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI DINOv3

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:5edacf92c3a4f915f13a0e9400292028a9c8e18211e9e1a2489b5dc9467fc96b

Observation a4ccb308-4618-4552-b641-99b777849004 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:c9fbef47e36f0639d704c5b96ae99e0b5b9ed3abc0cbf79ad4ea510cd4f28df4

Observation 32c0aab1-af3d-41b6-b8c9-6631daf00e51 · outbound

This paper cites Kleister: key information extraction datasets involving long documents with complex layouts.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Kleister: key information extraction datasets involving long documents with complex layouts

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:328a8bb21c5e94ab6fb1298b185cea748bc6094f593874a6d1b531bd35f7cdc5

Observation b2d65c39-6b62-4449-bdea-047b42cd2138 · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:059cbe536bedb0adc8fb5a7df7becefaf56322bda811e40a2e05083acdf1c593

Observation 1dcea6d2-649f-439f-bf23-8555ab3c4094 · outbound

This paper cites Deepform: Understand structured documents at scale.Weights & Biases report, 4, 2020.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Deepform: Understand structured documents at scale.Weights & Biases report, 4, 2020

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:de77eafa8d048423cd5fa02760010f39b9531738b7a2b5ab5999a5a615356583

Observation 2e4ede48-72c3-4f26-b347-235866852955 · outbound

This paper cites Hunyuanocr technical report.arXiv preprint arXiv:2511.19575, 2025.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Hunyuanocr technical report.arXiv preprint arXiv:2511.19575, 2025

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:bd9ed9cbe9259d6538f92e5e9963f985746a9f8fec570b45601bd944928ef24e

Observation 1908098a-9a02-4bdf-b4b0-a2e905dd33c1 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Kimi K2.5: Visual Agentic Intelligence

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:733347b884d21fb3be754eee705fedd18ff84b1b0b44154ebae331addea12484

Observation bf371730-be48-44f0-93b4-d31f8282461e · outbound

This paper cites Kwai Keye-VL Technical Report.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Kwai Keye-VL Technical Report

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:f2146bd8329f5f545d206d3ea3b404dcb3112e0745bacf4d3461e95b15be2f57

Observation 4d2e50e9-25bc-4e2e-a0bf-22e9519f3f8c · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:408b6326cffe9fd63b4ed568fef1eef3bd1423354b9278d9001cb6e940da1d63

Pith citing papers

No inbound Pith citation observations are available.