MAPs is a new amusement-park simulator benchmark on which frontier LLM agents score 7–15% of human performance, exposing persistent gaps in long-horizon planning, active learning, spatial reasoning, and handling stochasticity.
Are LLM s capable of data-based statistical and causal reasoning? benchmarking advanced quantitative reasoning with data
1 Pith paper cite this work, alongside 23 external citations. Polarity classification is still indexing.
1
Pith paper citing it
23
external citations · OpenAlex
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
MAPs is a new amusement-park simulator benchmark on which frontier LLM agents score 7–15% of human performance, exposing persistent gaps in long-horizon planning, active learning, spatial reasoning, and handling stochasticity.