A self-supervised pre-training framework that enriches text-image relations through patch permutation and block masking improves scene text recognition accuracy on 12 benchmarks.
Simmim: A simple framework for masked image modeling,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition
A self-supervised pre-training framework that enriches text-image relations through patch permutation and block masking improves scene text recognition accuracy on 12 benchmarks.