Pith. sign in

REVIEW 1 cited by

UnrealText: Synthesizing Realistic Scene Text Images from the Unreal World

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.10608 v6 pith:C2FKX5N3 submitted 2020-03-24 cs.CV

classification cs.CV
keywords scenetextimagesrecognitiondetectionrealisticsyntheticunrealtext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthetic data has been a critical tool for training scene text detection and recognition models. On the one hand, synthetic word images have proven to be a successful substitute for real images in training scene text recognizers. On the other hand, however, scene text detectors still heavily rely on a large amount of manually annotated real-world images, which are expensive. In this paper, we introduce UnrealText, an efficient image synthesis method that renders realistic images via a 3D graphics engine. 3D synthetic engine provides realistic appearance by rendering scene and text as a whole, and allows for better text region proposals with access to precise scene information, e.g. normal and even object meshes. The comprehensive experiments verify its effectiveness on both scene text detection and recognition. We also generate a multilingual version for future research into multilingual scene text detection and recognition. Additionally, we re-annotate scene text recognition datasets in a case-sensitive way and include punctuation marks for more comprehensive evaluations. The code and the generated datasets are released at: https://github.com/Jyouhou/UnrealText/ .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TEACH: Text Encoding as Curriculum Hints for Scene Text Recognition

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Feeding ground-truth label embeddings into an STR decoder and progressively masking them based on training loss improves accuracy on several benchmarks while leaving inference unchanged.

Pith tools