A 7B LLaVA model fine-tuned on Claude 3 Sonnet's synthetic labels matches Sonnet's extraction accuracy on sharp expense receipts while cutting cost by 85% and increasing speed 5x.
Wukong-reader: Multi-modal pre-training for fine-grained visual document understanding, 2022
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Information Extraction from Heterogeneous Documents without Ground Truth Labels using Synthetic Label Generation and Knowledge Distillation
A 7B LLaVA model fine-tuned on Claude 3 Sonnet's synthetic labels matches Sonnet's extraction accuracy on sharp expense receipts while cutting cost by 85% and increasing speed 5x.