Seven different agent personas made a vision-language model describe the same COCO image with under 10% lexical overlap, which the paper interprets as >90% context-dependent affordance computation.
A survey on efficient vision-language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Context-Dependent Affordance Computation in Vision-Language Models
Seven different agent personas made a vision-language model describe the same COCO image with under 10% lexical overlap, which the paper interprets as >90% context-dependent affordance computation.