On the CopsRef dataset, referring expression comprehension reveals that vision-language models fail at dynamic spatial relations, multi-relation expressions, and negations, while task-specific models handle geometric relations better.
A.2 Other Models In our analysis, we also experimented with InstructBLIP[Dai et al., 2023] and OpenFlamingo [Awadalla et al., 2023] mod- els
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploring Spatial Language Grounding Through Referring Expressions
On the CopsRef dataset, referring expression comprehension reveals that vision-language models fail at dynamic spatial relations, multi-relation expressions, and negations, while task-specific models handle geometric relations better.