Pith. sign in

REVIEW 2 cited by

TextNet: Irregular Text Reading from Images with an End-to-End Trainable Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.09900 v1 pith:AX3VQCWX submitted 2018-12-24 cs.CV cs.MM

classification cs.CVcs.MM
keywords textfeaturesirregularrecognitionend-to-endimagesnetworkperspective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reading text from images remains challenging due to multi-orientation, perspective distortion and especially the curved nature of irregular text. Most of existing approaches attempt to solve the problem in two or multiple stages, which is considered to be the bottleneck to optimize the overall performance. To address this issue, we propose an end-to-end trainable network architecture, named TextNet, which is able to simultaneously localize and recognize irregular text from images. Specifically, we develop a scale-aware attention mechanism to learn multi-scale image features as a backbone network, sharing fully convolutional features and computation for localization and recognition. In text detection branch, we directly generate text proposals in quadrangles, covering oriented, perspective and curved text regions. To preserve text features for recognition, we introduce a perspective RoI transform layer, which can align quadrangle proposals into small feature maps. Furthermore, in order to extract effective features for recognition, we propose to encode the aligned RoI features by RNN into context information, combining spatial attention mechanism to generate text sequences. This overall pipeline is capable of handling both regular and irregular cases. Finally, text localization and recognition tasks can be jointly trained in an end-to-end fashion with designed multi-task loss. Experiments on standard benchmarks show that the proposed TextNet can achieve state-of-the-art performance, and outperform existing approaches on irregular datasets by a large margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Unconstrained End-to-End Text Spotting

    cs.CV 2019-08 conditional novelty 7.0 of 10

    A Mask R-CNN and attention-based text spotter handles curved text by masking RoI features instead of rectifying them, and uses OCR-engine labels as extra training data to set state-of-the-art results on ICDAR15 and To...

  2. Focus-Enhanced Scene Text Recognition with Deformable Convolutions

    cs.CV 2019-08 conditional novelty 4.0 of 10

    Using deformable convolutional layers in the middle of a CRNN improves recognition of irregular scene text by several accuracy points on TotalText and ICDAR 2015 without image rectification.

Pith tools