Pith. sign in

REVIEW 1 cited by

UPOCR: Towards Unified Pixel-Level OCR Interface

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.02694 v2 pith:TW6GKHXR submitted 2023-12-05 cs.CV

classification cs.CV
keywords tasksupocrmodelpixel-leveltextunifiedgeneralistinterface
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InstructOCR: Instruction Boosting Scene Text Spotting

    cs.CV 2024-12 conditional novelty 6.0 of 10

    InstructOCR conditions an autoregressive scene text spotter on human-language instruction templates and reports gains on Total-Text, ICDAR2015, ICDAR2013, TextVQA, and ST-VQA.

Pith tools