REVIEW 2 cited by
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We address the challenging problem of Natural Language Comprehension beyond plain-text documents by introducing the TILT neural network architecture which simultaneously learns layout information, visual features, and textual semantics. Contrary to previous approaches, we rely on a decoder capable of unifying a variety of problems involving natural language. The layout is represented as an attention bias and complemented with contextualized visual information, while the core of our model is a pretrained encoder-decoder Transformer. Our novel approach achieves state-of-the-art results in extracting information from documents and answering questions which demand layout understanding (DocVQA, CORD, SROIE). At the same time, we simplify the process by employing an end-to-end model.
Forward citations
Cited by 2 Pith papers
-
DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
A combined OCR, LLM, and fuzzy-matching pipeline locates answer spans in document images at block, line, word, and point granularity, with line-level grounding F1 of 69.10 on a new 70-document benchmark.
-
Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
Spatial ModernBERT is a token-classification model that adds layout coordinates to ModernBERT to extract tables and key-value fields, with benchmark scores that fall short of the claimed state of the art.
Discussion (0). Sign in to comment.