ROSA, which combines four image rotations with likelihood-ranked sampling, improves VQA accuracy on misoriented text by up to 11.7 absolute points over greedy decoding.
Improving Text Proposals for Scene Images with Fully Convolutional Networks
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Text Proposals have emerged as a class-dependent version of object proposals - efficient approaches to reduce the search space of possible text object locations in an image. Combined with strong word classifiers, text proposals currently yield top state of the art results in end-to-end scene text recognition. In this paper we propose an improvement over the original Text Proposals algorithm of Gomez and Karatzas (2016), combining it with Fully Convolutional Networks to improve the ranking of proposals. Results on the ICDAR RRC and the COCO-text datasets show superior performance over current state-of-the-art.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ROSA: Addressing text understanding challenges in photographs via ROtated SAmpling
ROSA, which combines four image rotations with likelihood-ranked sampling, improves VQA accuracy on misoriented text by up to 11.7 absolute points over greedy decoding.