Pith. sign in

OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Optical Character Recognition (OCR) systems often introduce errors when transcribing historical documents, leaving room for post-correction to improve text quality. This study evaluates the use of open-weight LLMs for OCR error correction in historical English and Finnish datasets. We explore various strategies, including parameter optimization, quantization, segment length effects, and text continuation methods. Our results demonstrate that while modern LLMs show promise in reducing character error rates (CER) in English, a practically useful performance for Finnish was not reached. Our findings highlight the potential and limitations of LLMs in scaling OCR post-correction for large historical corpora.

fields

cs.DL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Comparing OCR Pipelines for Folkloristic Text Digitization

cs.DL · 2025-07-25 · conditional · novelty 5.0

The best OCR pipeline for Slovene folklore depends on document type: olmOCR for clean typewritten pages, Tesseract plus LLM post-processing for complex layouts and degraded newspapers.

citing papers explorer

Showing 1 of 1 citing paper.

  • Comparing OCR Pipelines for Folkloristic Text Digitization cs.DL · 2025-07-25 · conditional · none · ref 8 · internal anchor

    The best OCR pipeline for Slovene folklore depends on document type: olmOCR for clean typewritten pages, Tesseract plus LLM post-processing for complex layouts and degraded newspapers.