A new maze-navigation benchmark claims LLM spatial reasoning is language-dependent, with O3 exceptional and other models failing by looping, but the looping result is an artifact of the termination rule.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
A new maze-navigation benchmark claims LLM spatial reasoning is language-dependent, with O3 exceptional and other models failing by looping, but the looping result is an artifact of the termination rule.