Feeding ground-truth label embeddings into an STR decoder and progressively masking them based on training loss improves accuracy on several benchmarks while leaving inference unchanged.
Experiments Setup Following standard practice in scene text recognition (STR) [19, 5], we adopt both synthetic and real-world datasets for train- ing and evaluation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TEACH: Text Encoding as Curriculum Hints for Scene Text Recognition
Feeding ground-truth label embeddings into an STR decoder and progressively masking them based on training loss improves accuracy on several benchmarks while leaving inference unchanged.