Pith. sign in

Paper Citation Record · LEDGER

COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:1601.07140.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1601.07140 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:30.447765Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:59:57.589641Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0e48bf8-6017-43ec-85aa-2b686cbdaa99 · inbound

ICDAR 2019 Competition on Scene Text Visual Question Answering cites this paper.

ICDAR 2019 Competition on Scene Text Visual Question Answering COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-25T12:16:56.500827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T12:16:41.394335Z digest=sha256:18e209b0c735724aef4a2cb082dd1b1243cf0b7ba77f207bd0890134f1aded71

Observation 3c2775f6-b66b-4c10-a044-f4cdd37bca1b · inbound

ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition -- RRC-MLT-2019 cites this paper.

ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition -- RRC-MLT-2019 COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T11:55:46.040781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T11:51:09.915069Z digest=sha256:0a2c23b0d3403ff4b8b8b37cb26fb865aafb07976b2fb745689d75d82cffb8ed

Observation a1b33149-6c26-421f-bfc4-a3ad6aac0d78 · inbound

Understanding Deep Learning Techniques for Image Segmentation cites this paper.

Understanding Deep Learning Techniques for Image Segmentation COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 203

Resolution
verified exact
local_arxiv, observed 2026-05-24T21:46:24.322680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T21:46:17.736097Z digest=sha256:7e9fa5e763c48e315ba6ec93cb99bfd9b3f900b634e27b078ea0595e4bd67e9c

Observation f2d68b78-65dd-460f-9954-608872ca8435 · inbound

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models cites this paper.

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:55:35.647328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T09:55:35.452649Z digest=sha256:c7c223d1d983dae5eb75b0e42bc97b264f97744609f973765a2087562a61d587

Observation d3f9527f-2649-430e-ad27-35004440544a · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.223609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:a9949c302b9af7294bbcaa172391e09b80e7af7dc77aec157407302ea2478fb4

Observation 891e267d-cfe6-4107-95fd-e32739877243 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:57.882154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:2db636fa43eeb6af608e485d4995181373cd5f4790c0dfc170ed12b9f3454020

Observation ceac8f92-154a-4c3a-bc42-408b61aedd69 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 233

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:15:47.131840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:83713aa90c1ad4b83f0930e3e06f5fbd91c1a1c4d330e39624229c9777dce6b5

Observation b5a23ac5-b373-4f51-9a82-df50d4e20459 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.078687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:48a1ade504acad55b6dc15e080b5850c4b4712b44cdd9487fe15a64ff22ad6c6

Observation c0d259c6-f13f-4a01-86bc-96ac9a44bcdf · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.841117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:c65439af86b850ca57e14c5dca3ef69bb83d658981e35c4bb9ecd3c928f1734c

Observation ff8a32bf-64c8-4c88-b81c-c4f14956c305 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.854550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:96d7f8f5686adb337cfe30b0096e72dfc327343320f6cca3f9b4e30f031da6d2

Observation 7e25e0db-a2aa-4f31-ad3f-0b9db62cb526 · inbound

Common Inpainted Objects In-N-Out of Context cites this paper.

Common Inpainted Objects In-N-Out of Context COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:42:15.985486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:38:44.069681Z digest=sha256:abf3f3fd3ca00c86b4819705eeca0f936dcc87f5a40a0de18bcbe15ddc9c946e

Observation a1f5b464-6f14-461f-aedd-9d6cbb42dc60 · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.447765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.447765Z digest=sha256:b336ea267c47e6bd1f9a3fe15e2a56ab06f3f8099ac32ae3e3057ce7b7c5dcbb

Observation 8dc194d2-09e1-4148-9fc2-d0f05267d9e5 · inbound

The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing cites this paper.

The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:47.198137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:47.198137Z digest=sha256:3f2a6e586264d829d565e3ebbc60723e8bdbd8b373f238a9f85a66471c5636bb

Observation 2c669783-fe6d-4e56-ba56-abe50041330b · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:13:02.871922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:dced371ab800b6717eb05ce7c181286b7350fc2f36af6efc7acd6689fe5a7fc7

Observation 7c9bffc3-6d5d-43bf-a0dd-5ee482b6adcc · inbound

Text-Aware Image Restoration with Diffusion Models cites this paper.

Text-Aware Image Restoration with Diffusion Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:41:09.684861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:41:09.684861Z digest=sha256:dc2e5b2cc430e9ffed566412c5798493269ce220143415206a3f70c4e29ecabe

Observation 0ec19cce-6c0a-4f41-be1e-fd52bd161f61 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:01.817911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:01.817911Z digest=sha256:7616711d520ec6edb9370f2c634cefc391675b62998e360f142038b5de84fc1d

Observation 97908528-2ae1-4860-b546-9b39ae56d97e · inbound

TEACH: Text Encoding as Curriculum Hints for Scene Text Recognition cites this paper.

TEACH: Text Encoding as Curriculum Hints for Scene Text Recognition COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T05:56:41.210900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:56:41.210900Z digest=sha256:2e4ecc0de0548dd4ebd2c4f10b66c78767a85449cc43a443fdf2209e6db63889

Observation 9dd17bb6-5565-460a-877d-4c48c67a0989 · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 76

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:40:42.478766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:31e83235b79a840203eb7c656039ad8bfb648fb17c2d3e25ba95409983ce27d0

Observation fce723f1-2789-4248-8839-cba58936be48 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.739428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:5964e33d6669867870a0cf981c938881c93f1de3df454e199dcd9b1be0ccb82a

Observation 71879820-a1a8-4150-a37f-9f8433f58ba0 · inbound

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models cites this paper.

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:10.153453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:44:37.891638Z digest=sha256:f76d5d73e09dff023f8a7e2c811ee59bcc9caa85cf69c28581c349623e068ee8

Observation f3710646-d8aa-44e9-8a29-42a9389dd5c0 · inbound

StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting cites this paper.

StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:33:21.527196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:30:06.831024Z digest=sha256:a286eef7a47fea597085c60d361e9cc5056629d7d26dd76ddba793c3dc2ac772

Observation fb39db39-42f0-4991-a859-0d22e6914ac3 · inbound

Object-Centric Dataset Resources for Constrained-Data Image Generation and Augmentation cites this paper.

Object-Centric Dataset Resources for Constrained-Data Image Generation and Augmentation COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:09:37.450126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T14:49:39.520109Z digest=sha256:b04fa233341e934a9bf0b364c4ad188b835469af5b7d479679048ca5ab39469d

Observation af31b272-419c-4b86-aa21-0a198f7202e9 · inbound

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications cites this paper.

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:59:57.594886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:07:54.919001Z digest=sha256:3d2ea9111f1d0428a82d69ab42d19b94d0565ad22748cb260e630b55534f3287

Observation 79b53422-c0aa-4c16-abcb-d74633c8a260 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 291

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:45:47.653390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:08ecebf2ffcf2cc03f8a78da05fb62b9e9467b5df4f464d606b94118c1767eeb

Observation 1ddba967-edfc-4ccb-a0c7-1196f832b557 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 291

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:17:23.965840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:fb49e7d7a8d9d20490b0036170cafc7fa308460d3487dd66aa448d96d22da33b

Observation 4b34ba75-39d5-4d50-bce0-1e922401e3ee · inbound

Amplifying Membership Signal Through Chained Regeneration cites this paper.

Amplifying Membership Signal Through Chained Regeneration COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:41.144448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T06:22:08.140403Z digest=sha256:3753dc77111bb2ee03431183d38458734b71cb0bcb96d1e96a100ab88f2bba96

Observation 92b91896-923d-45ec-94d7-5b4240056c52 · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:6c690b9ab39ecb0b7a1579d429e4ea7273183316ebe33ccae2b92c9f096f8fb7

Observation 37898c0a-39d9-4720-9c59-a082d52ccf1d · inbound

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI cites this paper.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:08c322acd12e8cb44309315087c7d76678ffda90ef17e35914396067a19f7c84

Observation aedb5422-ff5a-47ce-84cc-1f9660204bbb · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.237847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.237847Z digest=sha256:edf258635ede24d2d54934519a559504f7cdbfb08533c2339e9d1ebf25114a96