A new diagnostic benchmark decomposes LLM spatial navigation into three cognitive scales and shows that cross-scale aggregation, not single-level deficits, causes failure beyond small mazes.
arXiv:2502.16690 (2025).https://arxiv
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
MentalMap benchmark identifies a universal L3 reasoning cliff in LLMs' text-based spatial reasoning that persists across languages, scales, and prompting, and is replicated in human evaluations.
Reasoning-tuned LLMs reliably complete navigation in partial-observability gridworlds but take longer paths than oracle optima, with few-shot prompting reducing invalid moves and action priors like UP/RIGHT causing loops.
citing papers explorer
-
Lost in Aggregation: A Multi-Scale Diagnostic Benchmark for LLM Spatial Navigation
A new diagnostic benchmark decomposes LLM spatial navigation into three cognitive scales and shows that cross-scale aggregation, not single-level deficits, causes failure beyond small mazes.
-
Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning
MentalMap benchmark identifies a universal L3 reasoning cliff in LLMs' text-based spatial reasoning that persists across languages, scales, and prompting, and is replicated in human evaluations.
-
LLMs for Text-Based Exploration and Navigation Under Partial Observability
Reasoning-tuned LLMs reliably complete navigation in partial-observability gridworlds but take longer paths than oracle optima, with few-shot prompting reducing invalid moves and action priors like UP/RIGHT causing loops.