Pith. sign in

REVIEW 2 cited by

Representing Online Handwriting for Recognition in Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15307 v1 pith:5AOAF5FJ submitted 2024-02-23 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords handwritingonlinerecognitionvlmsimagerepresentationappliedapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The adoption of tablets with touchscreens and styluses is increasing, and a key feature is converting handwriting to text, enabling search, indexing, and AI assistance. Meanwhile, vision-language models (VLMs) are now the go-to solution for image understanding, thanks to both their state-of-the-art performance across a variety of tasks and the simplicity of a unified approach to training, fine-tuning, and inference. While VLMs obtain high performance on image-based tasks, they perform poorly on handwriting recognition when applied naively, i.e., by rendering handwriting as an image and performing optical character recognition (OCR). In this paper, we study online handwriting recognition with VLMs, going beyond naive OCR. We propose a novel tokenized representation of digital ink (online handwriting) that includes both a time-ordered sequence of strokes as text, and as image. We show that this representation yields results comparable to or better than state-of-the-art online handwriting recognizers. Wide applicability is shown through results with two different VLM families, on multiple public datasets. Our approach can be applied to off-the-shelf VLMs, does not require any changes in their architecture, and can be used in both fine-tuning and parameter-efficient tuning. We perform a detailed ablation study to identify the key elements of the proposed representation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions

    cs.CV 2025-01 reject novelty 5.0 of 10

    The paper introduces a face-attribute hallucination benchmark and MIAVLM, a model that fuses multiview generated images and negative-instruction training, reporting higher balanced accuracy than zero-shot baselines.

  2. Early evidence of how LLMs outperform traditional systems on OCR/HTR tasks for historical records

    cs.CV 2025-01 conditional novelty 4.0 of 10

    Using 20 scanned pages of 1921 Belgian handwritten tables, the authors report that GPT-4o and Claude Sonnet 3.5 transcribe more accurately than EasyOCR, Keras, Pytesseract, and TrOCR, with two-shot prompting giving th...

Pith tools