Pith. sign in

REVIEW 1 cited by

Bidirectional Scene Text Recognition with a Single Decoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.03656 v2 pith:DZCIKVAC submitted 2019-12-08 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords bidirectionalbi-stettextdecoderdecodersdecodingmethodsscene
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scene Text Recognition (STR) is the problem of recognizing the correct word or character sequence in a cropped word image. To obtain more robust output sequences, the notion of bidirectional STR has been introduced. So far, bidirectional STRs have been implemented by using two separate decoders; one for left-to-right decoding and one for right-to-left. Having two separate decoders for almost the same task with the same output space is undesirable from a computational and optimization point of view. We introduce the bidirectional Scene Text Transformer (Bi-STET), a novel bidirectional STR method with a single decoder for bidirectional text decoding. With its single decoder, Bi-STET outperforms methods that apply bidirectional decoding by using two separate decoders while also being more efficient than those methods, Furthermore, we achieve or beat state-of-the-art (SOTA) methods on all STR benchmarks with Bi-STET. Finally, we provide analyses and insights into the performance of Bi-STET.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR

    cs.CV 2025-08 conditional novelty 3.0 of 10

    Direct line-level recognition with PARSeq (trained on synthetic line images) beats word-level pipelines by 5.4% FCA and runs 4x faster on the authors' 251-page English dataset.

Pith tools