Pith. sign in

REVIEW 4 cited by

Benchmarking Chinese Text Recognition: Datasets, Baselines, and an Empirical Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.15093 v2 pith:RJ4N2NW3 submitted 2021-12-30 cs.CV

classification cs.CV
keywords datasetsbaselinesrecognitiontextchineseevaluationprotocolsapplication
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The flourishing blossom of deep learning has witnessed the rapid development of text recognition in recent years. However, the existing text recognition methods are mainly proposed for English texts. As another widely-spoken language, Chinese text recognition (CTR) in all ways has extensive application markets. Based on our observations, we attribute the scarce attention on CTR to the lack of reasonable dataset construction standards, unified evaluation protocols, and results of the existing baselines. To fill this gap, we manually collect CTR datasets from publicly available competitions, projects, and papers. According to application scenarios, we divide the collected datasets into four categories including scene, web, document, and handwriting datasets. Besides, we standardize the evaluation protocols in CTR. With unified evaluation protocols, we evaluate a series of representative text recognition methods on the collected datasets to provide baselines. The experimental results indicate that the performance of baselines on CTR datasets is not as good as that on English datasets due to the characteristics of Chinese texts that are quite different from the Latin alphabet. Moreover, we observe that by introducing radical-level supervision as an auxiliary task, the performance of baselines can be further boosted. The code and datasets are made publicly available at https://github.com/FudanVI/benchmarking-chinese-text-recognition

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A document-oriented ViT family pretrained with text generation plus pixel reconstruction on 113M images transfers across recognition, detection, parsing, and understanding, setting open-source SOTA on MDPBench with a ...

  2. Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A shared transformer trained with continuous flow matching for images and discrete diffusion for text jointly restores scene text images and reads out their characters, removing the external OCR prior.

  3. Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A two-stage 'text-first, image-later' super-resolution framework restores glyph structures before enhancing the whole image, improving OCR accuracy and visual quality on a new extreme-zoom Chinese text dataset.

  4. MCCD: A Multi-Attribute Chinese Calligraphy Character Dataset Annotated with Script Styles, Dynasties, and Calligraphers

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 329,715-image Chinese calligraphy dataset with character, style (10), dynasty (15), and calligrapher (142) labels, plus single- and multi-task recognition benchmarks.

Pith tools