Pith. sign in

REVIEW 4 cited by

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.08387 v3 pith:QBQYKSKP submitted 2022-04-18 cs.CL cs.CV

classification cs.CLcs.CV
keywords documentlayoutlmv3imagetextmultimodalpre-trainedpre-trainingtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised pre-training techniques have achieved remarkable progress in Document AI. Most multimodal pre-trained models use a masked language modeling objective to learn bidirectional representations on the text modality, but they differ in pre-training objectives for the image modality. This discrepancy adds difficulty to multimodal representation learning. In this paper, we propose \textbf{LayoutLMv3} to pre-train multimodal Transformers for Document AI with unified text and image masking. Additionally, LayoutLMv3 is pre-trained with a word-patch alignment objective to learn cross-modal alignment by predicting whether the corresponding image patch of a text word is masked. The simple unified architecture and training objectives make LayoutLMv3 a general-purpose pre-trained model for both text-centric and image-centric Document AI tasks. Experimental results show that LayoutLMv3 achieves state-of-the-art performance not only in text-centric tasks, including form understanding, receipt understanding, and document visual question answering, but also in image-centric tasks such as document image classification and document layout analysis. The code and models are publicly available at \url{https://aka.ms/layoutlmv3}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 31 citations worldwide. Full citation record

  1. FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Fine-tuning small language models on synthetic financial errors yields high detection and editing scores, but the evaluation is limited to synthetic data from the same pipeline.

  2. Interfaze: The Future of AI is built on Task-Specific Small Models

    cs.AI 2026-02 reject novelty 4.0 of 10

    Interfaze-Beta uses small specialist models and tools to build a compact context that a general-purpose LLM answers from, reporting competitive benchmark scores without reproducible evidence.

  3. Vector embedding of multi-modal texts: a tool for discovery?

    cs.IR 2025-09 conditional novelty 4.0 of 10

    Using ColPali embeddings of 3,600 textbook page images, cosine similarity beats dot product, Euclidean, and Manhattan distances on top-5 retrieval, but only reaches 0.51 precision@5 without a text-only baseline.

  4. Template-Based Schema Matching of Multi-Layout Tenancy Schedules:A Comparative Study of a Template-Based Hybrid Matcher and the ALITE Full Disjunction Model

    cs.DB 2025-07 conditional novelty 4.0 of 10

    A template-based hybrid schema matcher aligns multi-layout tenancy schedules to a fixed target schema and reports an F1 of 0.881, but the score is obtained by grid search on the evaluation ground truth.

Pith tools