Pith. sign in

Paper Citation Record · LEDGER

InstructOCR: Instruction Boosting Scene Text Spotting

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2412.15523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15523 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:24:16.101773Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed2e4b1-4105-4225-b75c-1c21a627f890 · outbound

This paper cites Context Perception Parallel Decoder for Scene Text Recognition.

InstructOCR: Instruction Boosting Scene Text Spotting Context Perception Parallel Decoder for Scene Text Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.029155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.029155Z digest=sha256:48ec2a13073de9b24468ff2ee45a486396227f6ec1213c21c97c71e7878ffce1

Observation 67e2161a-9d07-4967-993d-74c303dda212 · outbound

This paper cites InstructDiffusion: A Generalist Modeling Interface for Vision Tasks.

InstructOCR: Instruction Boosting Scene Text Spotting InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.033705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.033705Z digest=sha256:dcd4028ca235503d81f8fb8f743372108e184897a4b4b7784857d34255e28c51

Observation 7cbb866a-e573-4159-aeec-6e46a2ad7c9a · outbound

This paper cites OCR-free Document Understanding Transformer.

InstructOCR: Instruction Boosting Scene Text Spotting OCR-free Document Understanding Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.048960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.048960Z digest=sha256:4aefc70bfddb8dfb167f017d67df6f56808f60057fbde31b673b05e8b0d4e0e1

Observation 7c1003f8-8de2-47c5-8774-3b8ac79bff0f · outbound

This paper cites Segment Anything.

InstructOCR: Instruction Boosting Scene Text Spotting Segment Anything

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.054357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.054357Z digest=sha256:4f1d4bf772bb3999463453709e64ee1811dd3579502c8b7a52da785ec54cd035

Observation 55638be9-ac1e-4fa9-b141-dde60e63ca80 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

InstructOCR: Instruction Boosting Scene Text Spotting Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.059498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.059498Z digest=sha256:a73fbf6c927f832e2cc41b23647d9d17da84350654a3f908ae042fe7e3e7f06d

Observation bcfc6fd0-f36c-4591-a414-ec1fbebe0002 · outbound

This paper cites SPTS v2: Single-Point Scene Text Spotting.

InstructOCR: Instruction Boosting Scene Text Spotting SPTS v2: Single-Point Scene Text Spotting

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T11:24:16.222183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:24:16.064952Z digest=sha256:e6e4669dbfd3de3c17eccf653438c93bb20032725815249160329e12693ba2fe

Observation ee1b0aa8-656e-4374-b20a-32bffae4b191 · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

InstructOCR: Instruction Boosting Scene Text Spotting KOSMOS-2.5: A Multimodal Literate Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.070213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.070213Z digest=sha256:8d9254724b29269457bf5c57eacc3fe7ed26798848b4c589bc32db4a90c1c647

Observation 08b6e3a0-caf1-43b4-82ba-31959c3f891f · outbound

This paper cites In 2017 14th IAPR international con- ference on document analysis and recognition (ICDAR), vol- ume 1, 1454–1459.

InstructOCR: Instruction Boosting Scene Text Spotting In 2017 14th IAPR international con- ference on document analysis and recognition (ICDAR), vol- ume 1, 1454–1459

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.365884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:24:16.075726Z digest=sha256:920d083327380ffecdb92549da76376dba81c1115d921ae17745bcaf82eac1cf

Observation 4dac2de8-9d73-4393-ae63-54579e5874c6 · outbound

This paper cites UPOCR: Towards Unified Pixel-Level OCR Interface.

InstructOCR: Instruction Boosting Scene Text Spotting UPOCR: Towards Unified Pixel-Level OCR Interface

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T11:24:16.185327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:24:16.088303Z digest=sha256:7d416fcf231b3f0efe975f62b8b6a197a1ce56b36f1de69a1cb19dcdbef94e00

Observation c6d0c6cb-54eb-499f-adce-c11763938a7a · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

InstructOCR: Instruction Boosting Scene Text Spotting UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.101773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.101773Z digest=sha256:b2dd33a4b2dfa0e61c367a750cb7adaf8aec4041a7f60b3f58b3dd41db31816e

Observation f113a607-6fe4-4838-bc18-46fc8ff6dbc9 · outbound

This paper cites In 12th international conference on document analysis and recognition, 1484–1493.

InstructOCR: Instruction Boosting Scene Text Spotting In 12th international conference on document analysis and recognition, 1484–1493

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.397469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:24:16.039342Z digest=sha256:ad70cd4daf8578407d98730b9db4bf3b0c5ee0b8d8d26a474a9de0d4807cb331

Observation 8f093086-639d-4626-b5c1-3082f28b7b1b · outbound

This paper cites In 13th international conference on document analysis and recognition, 1156–1160.

InstructOCR: Instruction Boosting Scene Text Spotting In 13th international conference on document analysis and recognition, 1156–1160

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.381022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:24:16.044359Z digest=sha256:d1c4ed6e0d1c574395fb92741207b967e4806d07c434ef195a5cbc6240f3f989

Observation 02c21729-2244-4df4-b7dc-09dfb6d184a5 · outbound

This paper cites In 2017 14th IAPR international conference on document anal- ysis and recognition (ICDAR), volume 1, 935–942.

InstructOCR: Instruction Boosting Scene Text Spotting In 2017 14th IAPR international conference on document anal- ysis and recognition (ICDAR), volume 1, 935–942

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:24:16.413833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:24:16.013602Z digest=sha256:7d0c67b7660dc41ceff69c668bd731a31af1a0ae3a75839aea1b79784466414f

Observation d915012a-b1e2-44be-92fa-3a29e0daee91 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

InstructOCR: Instruction Boosting Scene Text Spotting BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.017859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.017859Z digest=sha256:50c33929b0a910798f13ef6f6d3f72fbb44bd328d31ec0ddf0a1f348afcd0cc1

Observation a5d838e3-9ea4-449b-9aac-ab771b998daa · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

InstructOCR: Instruction Boosting Scene Text Spotting An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.024051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.024051Z digest=sha256:86b3332ee643fa07647746103f25dad950c90e0b0cfca75b19b2c72404d4ffec

Observation 3dca4bb9-7760-4cb3-8a55-75c53a62262a · outbound

This paper cites Pix2seq: A Language Modeling Framework for Object Detection.

InstructOCR: Instruction Boosting Scene Text Spotting Pix2seq: A Language Modeling Framework for Object Detection

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.009164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.009164Z digest=sha256:412b6634d46e1c135de60de861d6079f184fbb1cdb193253415ace9b0b4484d6

Observation 7507f045-ef53-4b51-9d07-75b3bad64f5a · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

InstructOCR: Instruction Boosting Scene Text Spotting GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.096991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.096991Z digest=sha256:c8cb1b4782fb0564eaf3ac1c613da4e3d535a14060923dbaa343e3407f9e5f5f

Observation c522d627-393f-42a0-9000-e8c0a3b9b3d5 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

InstructOCR: Instruction Boosting Scene Text Spotting Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.002128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.002128Z digest=sha256:96a4091c4731656aefad457c4c7eec70b81b0b9c4843e62b96c412fd197ae711

Observation 8c07e379-aa0b-41b6-927f-a3aafe304c2c · outbound

This paper cites OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition.

InstructOCR: Instruction Boosting Scene Text Spotting OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:24:16.092992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:24:16.092992Z digest=sha256:9aec877b9130db15feee2b8dbed880024f9e33c9775028182bd8df8f5dd30fe0

Pith citing papers

No inbound Pith citation observations are available.