A weighted combination of text, image, caption, and OCR embeddings improves a custom retrieval score by 57.3% over text-only retrieval on a private 19-question enterprise HR document set.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Evaluating VisualRAG: Quantifying Cross-Modal Performance in Enterprise Document Understanding
A weighted combination of text, image, caption, and OCR embeddings improves a custom retrieval score by 57.3% over text-only retrieval on a private 19-question enterprise HR document set.