A new VLM benchmark with interlocking jigsaw pieces shows frontier and fine-tuned vision-language models solve 4x4 puzzles but collapse to near random on 8x8 and larger grids.
CVPR , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
A new VLM benchmark with interlocking jigsaw pieces shows frontier and fine-tuned vision-language models solve 4x4 puzzles but collapse to near random on 8x8 and larger grids.