Pith. sign in

REVIEW 2 cited by

Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.09550 v3 pith:5V6EQOHD submitted 2021-02-18 cs.CL cs.LG

classification cs.CLcs.LG
keywords informationlayoutdocumentslanguagemodelnaturaltransformerunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the challenging problem of Natural Language Comprehension beyond plain-text documents by introducing the TILT neural network architecture which simultaneously learns layout information, visual features, and textual semantics. Contrary to previous approaches, we rely on a decoder capable of unifying a variety of problems involving natural language. The layout is represented as an attention bias and complemented with contextualized visual information, while the core of our model is a pretrained encoder-decoder Transformer. Our novel approach achieves state-of-the-art results in extracting information from documents and answering questions which demand layout understanding (DocVQA, CORD, SROIE). At the same time, we simplify the process by employing an end-to-end model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DRISHTIKON: Visual Grounding at Multiple Granularities in Documents

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A combined OCR, LLM, and fuzzy-matching pipeline locates answer spans in document images at block, line, word, and point granularity, with line-level grounding F1 of 69.10 on a new 70-document benchmark.

  2. Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale

    cs.CL 2025-07 reject novelty 4.0 of 10

    Spatial ModernBERT is a token-classification model that adds layout coordinates to ModernBERT to extract tables and key-value fields, with benchmark scores that fall short of the claimed state of the art.

Pith tools