Pith. sign in

REVIEW 1 cited by

A Conglomerate of Multiple OCR Table Detection and Extraction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.08591 v1 pith:YDMPG7BZ submitted 2020-10-16 cs.IR cs.AI

classification cs.IRcs.AI
keywords tablesmultiplealgorithmdocumentsextractingimagetextappropriate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Information representation as tables are compact and concise method that eases searching, indexing, and storage requirements. Extracting and cloning tables from parsable documents is easier and widely used, however industry still faces challenge in detecting and extracting tables from OCR documents or images. This paper proposes an algorithm that detects and extracts multiple tables from OCR document. The algorithm uses a combination of image processing techniques, text recognition and procedural coding to identify distinct tables in same image and map the text to appropriate corresponding cell in dataframe which can be stored as Comma-separated values, Database, Excel and multiple other usable formats.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting the Past: Estimating Historical Appraisals with OCR and Machine Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A hand-annotated dataset of 1933 Hamilton County property appraisals is extracted with template-aligned OCR, and a random forest trained on contemporary features estimates those historical values with about 17% MAPE.

Pith tools