Pith. sign in

VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Escape rooms present a unique cognitive challenge that demands exploration-driven planning: with the sole instruction to 'escape the room', players must actively search their environment, collecting information, and finding solutions through repeated trial and error. Motivated by this, we introduce VisEscape, a benchmark of 20 virtual escape rooms specifically designed to evaluate AI models under these challenging conditions, where success depends not only on solving isolated puzzles but also on iteratively constructing and refining spatial-temporal knowledge of a dynamically changing environment. On VisEscape, we observe that even state-of-the-art multi-modal models generally fail to escape the rooms, showing considerable variation in their progress and problem-solving approaches. We find that integrating memory management and reasoning contributes to efficient exploration and enables successive hypothesis formulation and testing, thereby leading to significant improvements in dynamic and exploration-driven environments

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

TextAtari: 100K Frames Game Playing with Language Agents

cs.CL · 2025-06-04 · conditional · novelty 5.0

TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.

citing papers explorer

Showing 1 of 1 citing paper.

  • TextAtari: 100K Frames Game Playing with Language Agents cs.CL · 2025-06-04 · conditional · none · ref 55 · internal anchor

    TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.