Pith. sign in

Paper Citation Record · LEDGER

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

As of 5 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 36 inbound Pith citation observations for arXiv:2509.22186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.22186 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T13:25:31.884175Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:46:40.906153Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T17:07:25.732459Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact31
  • verified fuzzy29
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0eb17dcd-cb96-42cf-bbf4-6cc66047925f · outbound

This paper cites GPT-4 Technical Report.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.932307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:f400569ed4cae20787492b34aa38eec83ebf759428f947b394cc1dacac767b39

Observation 573fdc67-6c36-4bd9-8f60-ebeefafb04cb · outbound

This paper cites Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document Understanding.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:31.943757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:38ea6fbb1d700ecce03c0120bf256b81c56b1b0be0b7deb0af4f6cb1eecfd8d3

Observation ba8d191b-7eb0-4511-8ae0-dc7bc2157634 · outbound

This paper cites Qwen2.5-VL Technical Report.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.951234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:655bbc337f56d0b8119fa3eae025360a4f863779f5532e5ce496dcc0b8f04830

Observation b6b4b551-4d46-423c-9d97-1092ce276f30 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.958813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:44b335bedc77f9084e514e796f1d6b9e1623ac98b712e396736de1289ff9d7a0

Observation 093bba68-43e4-4f8e-96c6-192b8f2bf9cc · outbound

This paper cites Ocrflux.https://github.com/chatdoc-com/OCRFlux.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Ocrflux.https://github.com/chatdoc-com/OCRFlux

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.117297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:601b06a1bc86de122a061eb9ac36b5d3bd2f4121a939604faf693f299b33ab1c

Observation 04421ea2-b829-40fd-b508-4093dac1d46d · outbound

This paper cites Ocean-OCR: Towards General OCR Application via a Vision-Language Model.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.114062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:6e2b9ad006d5e14137f89d7546a32ef0a4ceb68ae72020f91df8f046da897bdd

Observation 61d5631e-b4a9-418b-8c94-bb801dfea1f4 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.965726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:b215ebd87b6a0da176e9aa0d78fd1ccc466faa6ec32ad209ddcf610459b08c68

Observation cb70f3ae-da33-464a-b1c6-7fe38e5f921b · outbound

This paper cites PaddleOCR 3.0 Technical Report.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing PaddleOCR 3.0 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.970747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:4700a1cc9e6dc3560cd1c45338f919670e78cd462731e6155705b0d43f99b817

Observation 03980eac-ab61-4aaa-ac6c-c7734cf21115 · outbound

This paper cites Vision grid transformer for document layout analysis.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Vision grid transformer for document layout analysis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.130041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:c89a180698459f204ea37a990e85a0bb34f2094dd1afd08519b11183e38e86c4

Observation d12d51b4-6981-43e9-8205-ede07f0f853a · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.Advances in Neural Information Processing Systems, 36: 2252–2274.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.Advances in Neural Information Processing Systems, 36: 2252–2274

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.133324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:ea6d9941d771e215e4b89687e5918edf892d7f76f08293f0a5cc39ce0a6d21a5

Observation 4f1bfba0-f5f8-4ac1-a8d4-0c2ccefa47ee · outbound

This paper cites Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:31.976486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:ec7e5d12215391cf7f12f52750a784745e2ec286decd4b0a515b688e6c375dd3

Observation 52d0a6e8-8ef2-4742-ab23-e6dc1ea29c06 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:ef799a036cdec915eef89710c6ea9e02bc64ab32eff3b230ffb86c281024a3b5

Observation 4e95f10a-35a2-499a-9591-3c343fe030d8 · outbound

This paper cites Seed1.5-VL Technical Report.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.987859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:fb4957c2c69c5bde376b2aa3e3592ae6e4c89b413de555d752b83f1d2b2e4b2b

Observation 9a30f1f2-85b5-438b-b26f-f390800d78d4 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.146259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:e033618a6dde335c97ec8c8d0482d6e07ede56ab0ada50ecd932d432699bef7c

Observation f0fc35e4-6a7a-4ce7-9598-8220d09b31e4 · outbound

This paper cites Ocr-free document understanding transformer.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Ocr-free document understanding transformer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.149454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:5775da9449445f8b661aad0e49612b85e92bc833ad15225646b304a95dd5663c

Observation 4ef1a096-d4b0-416e-9151-791e8128886c · outbound

This paper cites Gon- zalez, Hao Zhang, and Ion Stoica.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Gon- zalez, Hao Zhang, and Ion Stoica

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.152409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:a06f335fb4533678e4d9d8f9e1ccce74eb8b0204c02d549c8359a63f524c5936

Observation bc1d039a-6fc1-438f-8669-2cf20ed0aabd · outbound

This paper cites arXiv preprint arXiv:2506.05218 , year=.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing arXiv preprint arXiv:2506.05218 , year=

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:25:31.993644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:8ffae4f68e252e65a49150005f83fb1632032324b61ae0732a787b6c8eebb035

Observation a57bb7ce-7adb-4547-b85c-afe3fb10108d · outbound

This paper cites Doctr: Document transformer for structured information extraction in documents.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Doctr: Document transformer for structured information extraction in documents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.158551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:2f9a4b81d3fc8fb899649fed109fe05ab6b0e6cc29a407e55d0483f74db3bdc2

Observation ec8a855e-69e8-422c-933d-891fe3663a54 · outbound

This paper cites Revolutionizing Retrieval-Augmented Generation with Enhanced PDF Structure Recognition.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Revolutionizing Retrieval-Augmented Generation with Enhanced PDF Structure Recognition

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:31.999250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:d28e6f79996e24c33c6c3937c749b3a2e5554b20ce3283963266163f1c27da8d

Observation 4fe99761-6a07-45a8-8934-3a22c03b774e · outbound

This paper cites Hrvda: High-resolution visual document assistant.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Hrvda: High-resolution visual document assistant

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.164376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:a97faf86c503167fa5283c5c69ad1d2cbe065c758b4376d7341176876de8e2d1

Observation 08762d7b-06fd-4937-88ce-1226f17b3ba2 · outbound

This paper cites PP-FormulaNet: Bridging Accuracy and Efficiency in Advanced Formula Recognition.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing PP-FormulaNet: Bridging Accuracy and Efficiency in Advanced Formula Recognition

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.004800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:4e97ea9aa3f62add508377fecb5d60e68ef716ec99551033fcb658ee81491305

Observation 58d19e35-81c9-4ed9-8fd8-5295c7f7a36d · outbound

This paper cites POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.010747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:9b146a52af14e032c6ede898214b27dbb7abbab6f03fea7b6a9665ba86e27c61

Observation b1e09e38-907f-449a-b072-d1c13ad8f18f · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.017809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:3075396300cf42f95c63417d154dfc72cb3ff70594d407def29d84a96ecfaf04

Observation 5792a4b9-125a-4cde-b6e3-c5956b76c387 · outbound

This paper cites Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.023399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:e7e23904022e9150a5da21e476e7e46bab75e17e01ddec733464101d89245caf

Observation 09fd7087-0dfa-44e7-961e-d63ffb8bf3c4 · outbound

This paper cites Optimized table tokenization for table structure recognition.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Optimized table tokenization for table structure recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.180701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:4203b931f25f9fff61af3204fbb7095ea7869be6363d51284960f41dfa3fc5fd

Observation 41e4b1c1-372c-463b-9a14-04b5ad310093 · outbound

This paper cites Nanonets-ocr-s.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Nanonets-ocr-s

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.183669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:e62bf4070cc2b3deaf1f2657c03cd021ecbcf976efc8693443ab27a055fdaeaf

Observation b0937d3c-f884-40f6-949d-b11cc6de362d · outbound

This paper cites Mathpix.https://mathpix.com/.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Mathpix.https://mathpix.com/

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.186530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:ac2c12d002c824d2bc040d3102f36842be3d6bcafdd6746d824b28b32e8d0b5f

Observation 2243c99a-0e45-49f1-8fff-ddf6d83300a6 · outbound

This paper cites SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:25:32.028685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:83e4e0dc80082e4067b1cbabfa0602068a132ed892d8ebaa7393f93cdb5b622e

Observation e7ab0eda-b5a6-4a65-9a70-0eda9239920e · outbound

This paper cites Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.034155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:1ae77a2e310d84b9fb6143596774d5b57b784bd97bad394b3b4bea57aea73fb9

Observation d91cfd00-34e8-4325-932e-e93d5f482180 · outbound

This paper cites Pdf-extract-kit.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Pdf-extract-kit

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.195460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:1e2a477c2c4c6bc2b626ee4f94c41f9730f8a22d2d304c2c6920eb6457fd5289

Observation daea5a8c-2d61-4ee5-a13f-614a0f586a9b · outbound

This paper cites Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.198535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:aadd9a06e122786cf69d682142dd15a17c4e160cd320a31b235184540ee5ae6f

Observation e8807e89-2831-4c1e-b2e3-022393b65b0c · outbound

This paper cites Marker.https://github.com/datalab-to/marker.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Marker.https://github.com/datalab-to/marker

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.202008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:d5df2f3589da56ddf858e065022b9b9aed0620c9a56de4b4a66638ff0a30c3b8

Observation 502ee242-ec00-4bfe-9ecf-fa1f03ab2649 · outbound

This paper cites Surya: A lightweight document ocr and analysis toolkit.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Surya: A lightweight document ocr and analysis toolkit

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.205165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:f180f1ba45e81173ec1d04b274c17ac92b434222eecc6f5e6c32ad72e0e5b84d

Observation 99dc6d66-502b-470d-823d-2deba20c07b5 · outbound

This paper cites Doclaynet: A large human-annotated dataset for document-layout segmentation.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Doclaynet: A large human-annotated dataset for document-layout segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.120658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:cf1254b04e1b72920096482f8d982e51b0dbc04e5fb122f29fdd220415f8af83

Observation 42e4aeed-1dbf-42ab-93de-92416fc341fd · outbound

This paper cites olmocr: Unlocking trillions of tokens in pdfs with vi- sion language models.arXiv preprint arXiv:2502.18443, 2025a.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing olmocr: Unlocking trillions of tokens in pdfs with vi- sion language models.arXiv preprint arXiv:2502.18443, 2025a

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.039105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:f9962fa1dd9ea439cbd3b148ad7e4800dafd02a7635ffc7cf5373e3b38b12972

Observation 81bd3dd6-80a7-4f3e-82db-f6e7427823d2 · outbound

This paper cites Rapid table.https://github.com/RapidAI/RapidTable.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Rapid table.https://github.com/RapidAI/RapidTable

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.126848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:336421822fab4e7122bfe2aab0e41a2222b1289c84eec1668527b1100d78289d

Observation 17d7cbcd-681d-4272-8def-c566e71e047f · outbound

This paper cites dots.ocr: Multilingual document layout parsing in a single vision-language model.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing dots.ocr: Multilingual document layout parsing in a single vision-language model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.136397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:3a5b9576025db3eb2931d59e64a8ed334a21e5ee360d81237ed11c9a50e1a509

Observation ae2b91bf-a4b2-4910-ad9e-61ba3ad00edc · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolu- tional neural network.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Real-time single image and video super-resolution using an efficient sub-pixel convolu- tional neural network

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.139697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:16cf174e0b2563c1e0dccc597f35343dd3ce8a0383013de0076be576bf0462b8

Observation a093c681-9128-4104-b413-19487e3d56fa · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.142960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:057879c2e468e6319a1c0a79ea36300aa851395f23067b3f48bb38bff3863fe4

Observation f70ad65b-3350-45c5-b78d-415902720509 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Unifying vision, text, and layout for universal document processing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.155543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:43de1d8db26432547aaa9fed680ca277813e81b5075f0a9373e86020a79ff15d

Observation 17af5ca7-490f-4b47-a1d1-61f233288be0 · outbound

This paper cites Mistral-ocr.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Mistral-ocr

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.161429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:33b11c30357d0424ed76d1e53431a4adbb7d3c1e9a6e67ff3b5f1c6ac5275716

Observation 93f37dde-91f7-4890-874a-25a952bcf41d · outbound

This paper cites Qwen2 Technical Report.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Qwen2 Technical Report

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.043633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:664ffb129ed3930170003bcec31498e3c6fc19c651543fbde5657b05e4880383

Observation d51660c9-a1b7-4d7d-8f49-009a4f899987 · outbound

This paper cites Omniparser: A unified framework for text spotting key information extraction and table recognition.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Omniparser: A unified framework for text spotting key information extraction and table recognition

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.170744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:d25400fe501c03259e2c036d42d722dca96d0f08120af0827ee37ea47593fc44

Observation e3691aa4-84fc-4532-ae02-70616618af65 · outbound

This paper cites Yolov10: Real-time end-to-end object detection.Advances in Neural Information Processing Systems, 37:107984–108011.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Yolov10: Real-time end-to-end object detection.Advances in Neural Information Processing Systems, 37:107984–108011

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.174253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:f998bbc22f0b17bf2f62ed190b0f3b4f1ba7592feee44606618752875d9f72f5

Observation 1004e3a8-1653-4f9e-a5dc-a5db11d1fcb3 · outbound

This paper cites UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.049356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:17c301aa2ce736440752d34ae49dc9a8d1fc2a1c5b76e6e3f73a14fbd6b0fb2e

Observation c7931339-2df4-41a4-abbe-cc18a883d718 · outbound

This paper cites MinerU: An Open-Source Solution for Precise Document Content Extraction.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing MinerU: An Open-Source Solution for Precise Document Content Extraction

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.054130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:e29ffc01b550e52e558fcb122094cd00a61b68ade41dafcd9f6db5a15a9f6583

Observation 8e7153a0-26d5-4d20-a281-ed8fa869d40a · outbound

This paper cites Image over text: Transforming formula recognition evaluation with character detection matching.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Image over text: Transforming formula recognition evaluation with character detection matching

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.192783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:ee239981d851e01c9ec7167447050b883428247ba39685a29704860ff450b91a

Observation 05f17076-2410-4fd1-9208-3759224807ea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.058911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:edb6434d343baec9d16024de361efad1c60378d5008ae1c8f80a8af109713f10

Observation 994d66c7-f5d7-4c88-8908-910037e8a3b0 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.063507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:a4983146f2285ffa84574b5af423661e49acaf27b3c605ed7f9909d4f3d04c8c

Observation 95ac9905-f8ce-43f7-aba8-c0db34cf4bc0 · outbound

This paper cites LayoutReader: Pre-training of Text and Layout for Reading Order Detection.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing LayoutReader: Pre-training of Text and Layout for Reading Order Detection

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.068135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:412ead2a08e8473e77f7c4c9c548ee257f4af80db8945dbc07b743d7792d4413

Observation 30b891ca-a3eb-4600-8eee-a1c01e80bee7 · outbound

This paper cites Vrdu: A benchmark for visually-rich document understanding.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Vrdu: A benchmark for visually-rich document understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.189732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:5460d80251550836034cfdc8327f592ec66fb1cc3f65d0a76195ef35ef273dd3

Observation 765ad383-e252-4db4-9a90-bec6d7788331 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:57.985910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:3dd943efb17b62cdf327ac620f49e2e73da662d6c2a4234a6458bf1ee4f213fa

Observation 5740e854-0131-47fa-9469-3996b278d76e · outbound

This paper cites CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.078836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:447a4aba6b607da0be89f31ed350d220e86335b8a5932fe796096eda1ba5502c

Observation 20685813-7daa-4a34-9c0f-8b578341645e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.083168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:f797e87533c4a8887a7a5dd9f574e51c0e8c03fd5d0921b9b52e9690a2a10783

Observation 52a366ce-7778-499b-b0a9-3e1c069b89e5 · outbound

This paper cites MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.087706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:c8de18df073d0b510990e7afbc1be00abe5739161b0a6ef748d61216deb70bce

Observation 8c64bdec-bfed-4055-8992-c5e8dc0bd9fa · outbound

This paper cites OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.092351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:b18abb61a4e3572d9dafc75cdc96b5df5462c1b72923df6802ee626d279a4f6e

Observation 07e0b3f9-3385-4af6-90fe-ea0beeff32d0 · outbound

This paper cites Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.096611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:50ee4831236d1c894f61cc67ba4f6eccaaccaa3455ceda0d1f972ddb1586ce47

Observation 3b32a8be-a118-49cc-96d6-abb1e6f9bed1 · outbound

This paper cites Retrieval-Augmented Generation for AI-Generated Content: A Survey.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Retrieval-Augmented Generation for AI-Generated Content: A Survey

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.100715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:4fcd18c8f57be72de748cd36c06bdc431a877122c9930be5ff733b7c7a073005

Observation 942e003c-d139-401d-b390-bb6c219ad5df · outbound

This paper cites DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:25:32.105068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:03cfb7c45638a5d74a97ef8d36097edbf24911d6d4ae1e14d068563bff1c7d1e

Observation c727ceae-3690-4abf-b417-d1f1dd3bfd21 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.177658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:40b6bf8c7a15316fafb0cd4dac650cbca156b776104978c405f4640fefacb6c4

Observation 15bbe352-60a6-4854-abcf-3d9247b7b49e · outbound

This paper cites Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.123813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:1cdfe975b49b430b0930ae2ea867d53151cc2084ba25bfd1c65af76ce00e38c2

Observation 7bf75f20-0a1b-4ae1-89c0-a39023d19af6 · outbound

This paper cites Image-based table recognition: data, model, and evaluation.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Image-based table recognition: data, model, and evaluation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:25:32.167584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:9eb73afbe8a71c7124ed8a558f6ff7c3c4883a3784245610eb38211474a3851f

Observation 89880a42-32d6-42e5-ae30-76fd55f6602a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:32.109409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:ed589e869c5464b56d0039d1b6507f587f71a72de1095ef1016c2a8bf15d329e

Pith citing papers

Observation 5307e98b-c8c7-4a47-bb0d-32f89cf36d49 · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:20:11.529414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:d90733a2acc5d2624332426a8e0e1bcbb6049a33fc052146ce1e447df8dd5649

Observation fa00f7a0-fc2c-4669-8566-93e8311d8c49 · inbound

Low-Resolution Editing is All You Need for High-Resolution Editing cites this paper.

Low-Resolution Editing is All You Need for High-Resolution Editing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:28:22.668743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:28:22.668743Z digest=sha256:b4d06f2a56347425af2c12105d9a911f181c57cbe03fc86a242f2f3e4f37a424

Observation 7b66e0d8-457f-4a8e-9507-8d98c1bbc67c · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.808178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.808178Z digest=sha256:a9398bd563da090aac31bd297e34c17e5a4d91117578f40e4d3321f1ca18fba7

Observation 3dcf2ae4-de54-485b-8b49-3ff70fbcc97c · inbound

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch cites this paper.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T13:12:01.889341Z digest=sha256:7837a3fcfc579425ecbb51c601ee8e7ead06fa7a1e26ba1e0ee99c597a814600

Observation 752379fa-e195-4dd4-aaa4-32406bf79e66 · inbound

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing cites this paper.

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T09:40:09.955493Z digest=sha256:eaa54bda5cf1c4cff4a898d4b3fcd05fde04e2fcb08f5edb19363f47ad0b9596

Observation c18ac8ea-f245-4968-b645-40226da52518 · inbound

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding cites this paper.

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T23:44:38.680422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:44:38.680422Z digest=sha256:f12e522a4a45e11612f1b1a0311bea0316b778f582c164d6768f79f784138909

Observation 7542957f-7bc6-487a-9840-64a2f94a9087 · inbound

DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing cites this paper.

DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:21.931070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:21.931070Z digest=sha256:ef1b4a7c946f7fb1285cc6d643e4141aafd236c6a70ec7363520f0e5fa69a577

Observation 9fb10128-ac5d-4d04-b332-6cae22609192 · inbound

HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis cites this paper.

HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T19:46:45.507654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T19:46:45.507654Z digest=sha256:1cb4f4bb3647fbe85e37cc347062dad946f494a384ea75175f047651c0dd6628

Observation 65aa4065-9a5e-4434-97c5-ff115bf9e586 · inbound

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild cites this paper.

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T18:54:56.763823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:54:56.763823Z digest=sha256:5dd43281f056f514f53acc5d953a6edd2380df21061f86ecf27e60fe52f4c590

Observation 28f09890-f0a7-4b23-bd7f-6bc39be97dee · inbound

Logics-Parsing-Omni Technical Report cites this paper.

Logics-Parsing-Omni Technical Report MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T13:37:44.189839Z digest=sha256:1c11c85306318bed260ee1728d2bc794afbe1f6861d772298dec4e70724e21e6

Observation dba10c65-cbe3-4939-b45d-da9ad6c94229 · inbound

Visual-ERM: Reward Modeling for Visual Equivalence cites this paper.

Visual-ERM: Reward Modeling for Visual Equivalence MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:19:42.002790Z digest=sha256:cd683fe5615f090a46e9264338908f5888380af00ec068bc1a78167f7c9de9f9

Observation 7e5ca173-117e-4cbd-be28-634907b1b85e · inbound

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing cites this paper.

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:25:19.782732Z digest=sha256:fa3207bc29916102de325b8afa6ad30799ced3ec8b6904461720120e22b55c87

Observation eb39dfcf-ab99-401f-b733-d8c5ea767c99 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:a0d089178e15a334d866c372f8d1fee4f5356c8ef97d4e41ccdca43b82690734

Observation 5e2c53d9-5a58-4515-a350-7428cee67402 · inbound

Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing cites this paper.

Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:07:44.073353Z digest=sha256:5dec86e1b2777b680b8fc954c4c487ea2585162a8d0cef29b38af657edfd527c

Observation 8f56fab7-0eb4-41c2-8843-e7566bff37e8 · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:fd17a144182f9c6e518f1b5d614f468ca89d693e0d9e2eb9f6e25066e1f46b48

Observation 41be4938-f2d1-4e54-b221-7dbd782dd1e7 · inbound

InstructTable: Improving Table Structure Recognition Through Instructions cites this paper.

InstructTable: Improving Table Structure Recognition Through Instructions MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:53:57.029294Z digest=sha256:068bd0c671cf1c30354f5f4dbaddcf8c88c663b8c8f81480aafe2d774cb49392

Observation 0cef38e2-9ff6-4cf5-9a26-6595a41f6fd1 · inbound

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale cites this paper.

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:58:41.377996Z digest=sha256:28419b9e2d697ab59bd5f964671ca737502bb0aa1cd0844055244fb361df549a

Observation 2733da71-3ba4-4760-9c90-0576bfba7102 · inbound

Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval cites this paper.

Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:26:44.498777Z digest=sha256:9399552e9d58972912dd520e6e77f595509528e51fef5a8caa7666e0584bb6cc

Observation 8c0534f0-97b8-4363-bd27-665910ca53f7 · inbound

The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design cites this paper.

The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T14:18:23.560561Z digest=sha256:c79e74149e3392e0e747527b94bb6c60beb2e2b3dc391300db34bb508ca459b2

Observation 22de78c7-d998-443e-bf55-9c5167dae9d5 · inbound

The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design cites this paper.

The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:40:35.576232Z digest=sha256:df7741d3790aef04598c0cf645dcab0366a94755829a7a63860dee99ac4e96c5

Observation 4f03c4df-cb9e-4ae1-92ef-4c05d49cc088 · inbound

Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval cites this paper.

Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage Retrieval MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:41:19.441055Z digest=sha256:eb92e69b9aaa134aca77615690de63ac46e611bd183295ee31f20496aaf6588a

Observation 7c18a55d-5e74-4ea7-acbf-fe93182b4c01 · inbound

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings cites this paper.

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:28:02.152600Z digest=sha256:26ac2d208e7ed2f749b8f78441d4f6990e9a036cad6873bb7ba1001149913c07

Observation cb55df40-c513-486e-9e4b-86746d7b6be5 · inbound

Information Extraction of Nested Complex Structure of Quantum Cascade Lasers via Large Language Models cites this paper.

Information Extraction of Nested Complex Structure of Quantum Cascade Lasers via Large Language Models MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:40:40.693392Z digest=sha256:8854400ce28b42a33fa92cd29d646dcab28c80fb5aae3953e7f53fdb2ed7b5ef

Observation 8a5459e4-6c2b-43ee-89df-a22adc55de57 · inbound

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence cites this paper.

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.206292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:37:36.144960Z digest=sha256:8cf3863f3c7e49f7f9553baefbcea60638e7033756e3db64c10ac18aebad2a43

Observation 29fb8c45-ea74-4d90-837b-d875804041a2 · inbound

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings cites this paper.

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:43:19.443498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:42:59.633053Z digest=sha256:0fee28940ac845b13c15ad9264151d1475ca83f19ba7ba6f4dfce03aac6ffadd

Observation 6c7d8c8c-3cab-46f4-95f8-798f47a7e05e · inbound

AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists cites this paper.

AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T03:59:32.440279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T03:58:01.585170Z digest=sha256:52478b70e513dc2f6eca966840bf5b804d369c6c97dc36c6c3a9d330e5e1b69a

Observation e8c0c729-5609-4b36-9035-54bf283284c9 · inbound

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing cites this paper.

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:01:08.740255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T05:58:04.055855Z digest=sha256:9b92a367b11ab4d55883cb2ff7403b4a875c5e3625cd524c67300ec977a32130

Observation 67fa30c4-b48b-47a9-a754-8e52982c1b40 · inbound

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing cites this paper.

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:05:48.205995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:37:33.750306Z digest=sha256:4775d752d430f6fd12a420293523035d9cff62a60fc4975ed2bb5b86e663e99e

Observation aed0d763-fb88-4e84-b5cf-0af1f25bea24 · inbound

ABot-OCR Technical Report cites this paper.

ABot-OCR Technical Report MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:33:28.142642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T13:29:17.221676Z digest=sha256:c2a4a25f0e1f2b87a78ca9afc27fbd9e1dc71dc6f1e97767dd47fbc1a733786d

Observation b12a3b30-40e9-4781-86b7-51ab969ef015 · inbound

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents cites this paper.

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:28.922465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T10:38:41.735204Z digest=sha256:54f93b285e4849a9144c97b3b83e33f64e25f9aa656e1c4bee1b8b3c1c4070c9

Observation a994715f-010f-4de8-9c72-e7859120a4ad · inbound

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training cites this paper.

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:46:28.765289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T10:39:14.444486Z digest=sha256:29757e14e936234302145a6d4f0cff6040b5a9d3fbd427e51e5e39849d9db4bc

Observation 2834b552-aa6c-49cf-907d-485eda7aab09 · inbound

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild cites this paper.

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:59:45.192479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T09:19:38.285839Z digest=sha256:c3f76e3e43f2a72ef513ba3c9375856b1438eb79bc2a3548e05b303eaf9fa172

Observation ad8863fc-922d-48e5-b117-06ef60504f21 · inbound

StrucTab: A Structured Optimization Framework for Table Parsing cites this paper.

StrucTab: A Structured Optimization Framework for Table Parsing MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:24:18.949444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T06:22:21.694355Z digest=sha256:828d503e69056e312ece74988a343f5dbfc247aa438cf53ef22d8ff90915978b

Observation 37d3c6b0-0e50-44ae-bfc2-b56fce723207 · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T17:07:25.733862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T17:02:28.089092Z digest=sha256:3a7443300f64ece968210715d0f254088e10647f1704d31b4829092163f075d8

Observation a5f84112-a861-4163-80f5-bffd6f2b7a51 · inbound

Infinity-Parser2 Technical Report cites this paper.

Infinity-Parser2 Technical Report MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:03:48.857432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:03:48.857432Z digest=sha256:adcd67bc6f367364459726074a3a4682b3e2911b39a3f4b5cb0e27998433f50c

Observation 191f36d1-a339-466d-a0fd-6d10bd7dc35b · inbound

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards cites this paper.

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:40.906153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:46:40.906153Z digest=sha256:7413d26019a3e327db043b0ee9a25bb2e6425d02f8c27ee640ac2910c4e1e396