A new multimodal benchmark shows that vision-language models are much worse at combining object counting with spatial-relation reasoning than at either task alone.
Geobench-vlm: Benchmarking vision-language models for geospatial tasks, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
A new multimodal benchmark shows that vision-language models are much worse at combining object counting with spatial-relation reasoning than at either task alone.