Pith. sign in

Paper Citation Record · LEDGER

Ocean-OCR: Towards General OCR Application via a Vision-Language Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2501.15558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15558 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:25.292711Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.548550Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation acf72259-e3de-4028-b256-1aca9c9a9fb7 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.292711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:25.292711Z digest=sha256:1a16fca6e5f60b66ed7eb3636a35ece5389a6c5512dc444b50e26f45274a7a22

Observation 9111af45-1f0d-4e70-88ec-035864d55cce · inbound

Toward Reliable VLM: A Fine-Grained Benchmark and Framework for Exposure, Bias, and Inference in Korean Street Views cites this paper.

Toward Reliable VLM: A Fine-Grained Benchmark and Framework for Exposure, Bias, and Inference in Korean Street Views Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:35.856508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:08:35.856508Z digest=sha256:0ece3d5088f8eccf2e3e0639b0dec035562d26dab96e5e938a1aa98cf646ef06

Observation 3a605ce2-499d-4e75-a0b5-7159b53b774b · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:28.448857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:28.448857Z digest=sha256:84ac02cd4066be73767f473f336e4bb83f02611aac516f3032eebe35ec0f533b

Observation 06237e27-d010-4abb-9116-41f832c4d3e2 · inbound

Multi-Agent Interactive Question Generation Framework for Long Document Understanding cites this paper.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.542039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.542039Z digest=sha256:bc39916088f7f4bced628f2f6446677cdc06a67789ebcb421d05dbd333e39d3e

Observation 04421ea2-b829-40fd-b508-4093dac1d46d · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.114062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:c6c64a06377d7a4a502b393ab849bb7471817da822912f9ec8f11ea2c202e37c

Observation 158a39c1-ba1f-4bdb-bbfb-83db6cf9faef · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.723234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.723234Z digest=sha256:96c4ec34815f6d4d4a0f5ca21a4fe8eb6590d4522ac51f4ce6db6275c2063d2e

Observation a292ab14-8fd7-4629-a62c-ba80216cc5ec · inbound

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding cites this paper.

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T23:44:35.717955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:44:35.717955Z digest=sha256:26233c66c7e0f192e02dc9c409d2fd1591a50d8cd514b52ecc792d31bd9181f5

Observation 877c9eba-c007-42a2-8f49-31be1c402949 · inbound

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild cites this paper.

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:54:55.310939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:54:55.310939Z digest=sha256:c7456e5760d9f1957734ade7a00b0dd5ee5209ad679fdd5469b250e20c5df3ac

Observation e2af4052-b678-4e86-ae6e-76dc34aed837 · inbound

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models cites this paper.

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:55:19.369879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:53:18.268970Z digest=sha256:c47ea7a64623a00e84fbe33242cd0704bf389ebe40fd14e6d0d2eb0b7cfe9df1

Observation e2b5bab6-7bb0-44cc-ba69-ec0db3ca0f09 · inbound

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale cites this paper.

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:52.422492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:58:41.377996Z digest=sha256:f02b3cdcca40c2940c7f8dd637d847f80e9b55f09511acaf9f19e66a67554784

Observation 5756a774-acce-4d8e-8de8-b8e307ce04d1 · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.225427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:72336465362e796a1a1f70c505cfe44ec6e0b6024f1e6088164e11fe289176f4

Observation bb7628da-c811-4387-8c37-f289b5b38f55 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:08.438004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T19:49:28.591419Z digest=sha256:29d1dd50da7d8fdc71b4fbd5796483a702d3774406c2775e0ac6ec00010d3559

Observation 4011a6ec-8973-4c21-95f2-521124c49979 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:19.407704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:16:31.918225Z digest=sha256:3a8a4c7f321ef724628e7657cca2ad2efc06a897224c5e0cc8c9309d484c6123

Observation 5c7f47a4-8d07-4cb9-9655-a8d8bd171293 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.735949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T10:28:15.293381Z digest=sha256:a54632aa14920394a5bfb72dfffa5cc5a60d6a2f79fa26f8766d311ecddc9f40

Observation d4be3596-a0fc-45f4-b334-18dbb93cfc32 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:02:58.684066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:79f0059f854234867f2428ab1218aa8f30b0b64e6b041e6b0ac440f48d5a5705

Observation c977fcce-9763-4e75-9c73-e7bc9eece571 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:54:47.078819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T09:51:40.160096Z digest=sha256:688f4fa5d2d278d5907756efee9703cc6138d9bc358d69db38eedee0303a0da0

Observation 51067411-2d32-4029-be1b-8e4549189bce · inbound

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition cites this paper.

Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.550573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:10:03.765621Z digest=sha256:60cb01d11ff7701221a1ecf5bdd63f29064e87af69ed7095c7943123c5686777

Observation e00da1c1-55e3-452d-b664-a41a385481bd · inbound

HPD-Parsing: Hierarchical Parallel Document Parsing cites this paper.

HPD-Parsing: Hierarchical Parallel Document Parsing Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T14:16:08.139655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:16:08.139655Z digest=sha256:f1a4df8f8c3aa0e61e525d6f073b84d3ba0759002fd3011672ab89e53e31ab5a

Observation 350f618f-059a-41b7-ba55-954416cb2379 · inbound

Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models cites this paper.

Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:46:09.712264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:46:09.712264Z digest=sha256:5cec426ae25805cfe614051761ef140f0a8ac0088b52e83d8f3a952dc2f70e24