REVIEW 2 cited by
Exploring OCR Capabilities of GPT-4V(ision) : A Quantitative and In-depth Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents a comprehensive evaluation of the Optical Character Recognition (OCR) capabilities of the recently released GPT-4V(ision), a Large Multimodal Model (LMM). We assess the model's performance across a range of OCR tasks, including scene text recognition, handwritten text recognition, handwritten mathematical expression recognition, table structure recognition, and information extraction from visually-rich document. The evaluation reveals that GPT-4V performs well in recognizing and understanding Latin contents, but struggles with multilingual scenarios and complex tasks. Specifically, it showed limitations when dealing with non-Latin languages and complex tasks such as handwriting mathematical expression recognition, table structure recognition, and end-to-end semantic entity recognition and pair extraction from document image. Based on these observations, we affirm the necessity and continued research value of specialized OCR models. In general, despite its versatility in handling diverse OCR tasks, GPT-4V does not outperform existing state-of-the-art OCR models. How to fully utilize pre-trained general-purpose LMMs such as GPT-4V for OCR downstream tasks remains an open problem. The study offers a critical reference for future research in OCR with LMMs. Evaluation pipeline and results are available at https://github.com/SCUT-DLVCLab/GPT-4V_OCR.
Forward citations
Cited by 2 Pith papers
-
Exploring Light-Weight Object Recognition for Real-Time Document Detection
Adapting IWPOD-Net to document detection gives a 1.8M-parameter rectifier that is faster than YOLO11, RTMDet, and Jdeskew while keeping OCR quality competitive on a synthetic ID dataset.
-
Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models
Classification-guided dynamic prompts improve zero-shot visual information extraction from 16 certificate types, reaching 86.43 F1 without LVLM fine-tuning on a private bidding dataset.
Discussion (0). Continue with ORCID to comment.