On the CopsRef dataset, referring expression comprehension reveals that vision-language models fail at dynamic spatial relations, multi-relation expressions, and negations, while task-specific models handle geometric relations better.
Eg:- The large poster that is leaning against the wall
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploring Spatial Language Grounding Through Referring Expressions
On the CopsRef dataset, referring expression comprehension reveals that vision-language models fail at dynamic spatial relations, multi-relation expressions, and negations, while task-specific models handle geometric relations better.