Vision-language models show some human-like patterns in visual search effort (flat for features, rising for conjunctions) but diverge on target-present vs absent slopes and enumeration accuracy when reasoning tokens proxy reaction time.
Frankland, Thomas L
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
VLMs fail at exact visual path following primarily at self-intersections, where performance drops sharply after the first crossing on a controlled polyline benchmark.
citing papers explorer
-
Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms
Vision-language models show some human-like patterns in visual search effort (flat for features, rising for conjunctions) but diverge on target-present vs absent slopes and enumeration accuracy when reasoning tokens proxy reaction time.
-
TraversalBench: Challenging Paths to Follow for Vision Language Models
VLMs fail at exact visual path following primarily at self-intersections, where performance drops sharply after the first crossing on a controlled polyline benchmark.