A pipeline combining Google Vision OCR, fastText language ID, masking, and neural post-correction reduces character error rate on Kwak'wala legacy books from 0.43 to 0.18 and from 0.33 to 0.15.
LAREX - A semi-automatic open-source Tool for Layout Analysis and Region Extraction on Early Printed Books
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
A semi-automatic open-source tool for layout analysis on early printed books is presented. LAREX uses a rule based connected components approach which is very fast, easily comprehensible for the user and allows an intuitive manual correction if necessary. The PageXML format is used to support integration into existing OCR workflows. Evaluations showed that LAREX provides an efficient and flexible way to segment pages of early printed books.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Developing a Mixed-Methods Pipeline for Community-Oriented Digitization of Kwak'wala Legacy Texts
A pipeline combining Google Vision OCR, fastText language ID, masking, and neural post-correction reduces character error rate on Kwak'wala legacy books from 0.43 to 0.18 and from 0.33 to 0.15.