A new benchmark converts text-only QA datasets into text images and shows that vision-language models degrade sharply on long visually presented contexts.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models
A new benchmark converts text-only QA datasets into text images and shows that vision-language models degrade sharply on long visually presented contexts.