Open vision-language models fail to select spatial demonstratives based on object distance in a human-like manner across four languages.
ArXiv , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
High OCR accuracy on standard metrics does not guarantee strong downstream RAG performance because structural and semantic errors cause retrieval and generation failures on challenging industrial documents.
citing papers explorer
-
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
Open vision-language models fail to select spatial demonstratives based on object distance in a human-like manner across four languages.
-
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
High OCR accuracy on standard metrics does not guarantee strong downstream RAG performance because structural and semantic errors cause retrieval and generation failures on challenging industrial documents.