A new benchmark shows frontier image-editing models can sometimes follow instructions rendered inside an image and answer on the same canvas, with the best model, Nano Banana Pro, reaching 48.7% strict proxy accuracy versus 96.1% for humans.
Are VLMs Really Blind
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Vision Language Models excel in handling a wide range of complex tasks, including Optical Character Recognition (OCR), Visual Question Answering (VQA), and advanced geometric reasoning. However, these models fail to perform well on low-level basic visual tasks which are especially easy for humans. Our goal in this work was to determine if these models are truly "blind" to geometric reasoning or if there are ways to enhance their capabilities in this area. Our work presents a novel automatic pipeline designed to extract key information from images in response to specific questions. Instead of just relying on direct VQA, we use question-derived keywords to create a caption that highlights important details in the image related to the question. This caption is then used by a language model to provide a precise answer to the question without requiring external fine-tuning.
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Image-Space Rule Discovery
A new benchmark shows frontier image-editing models can sometimes follow instructions rendered inside an image and answer on the same canvas, with the best model, Nano Banana Pro, reaching 48.7% strict proxy accuracy versus 96.1% for humans.