Pith. sign in

REVIEW 2 cited by

Planning in a recurrent neural network that plays Sokoban

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15421 v3 pith:ICJ2MJKW submitted 2024-07-22 cs.LG cs.AI

classification cs.LGcs.AI
keywords neuralplanningbehaviornetworksokobanmodelplanrecurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Planning is essential for solving complex tasks, yet the internal mechanisms underlying planning in neural networks remain poorly understood. Building on prior work, we analyze a recurrent neural network (RNN) trained on Sokoban, a challenging puzzle requiring sequential, irreversible decisions. We find that the RNN has a causal plan representation which predicts its future actions about 50 steps in advance. The quality and length of the represented plan increases over the first few steps. We uncover a surprising behavior: the RNN "paces" in cycles to give itself extra computation at the start of a level, and show that this behavior is incentivized by training. Leveraging these insights, we extend the trained RNN to significantly larger, out-of-distribution Sokoban puzzles, demonstrating robust representations beyond the training regime. We open-source our model and code, and believe the neural network's interesting behavior makes it an excellent model organism to deepen our understanding of learned planning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Latent Programming Horizons in Coding Agents

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Linear probes on coding-agent residual streams decode current program properties (AUC up to 0.83) and predict future edit outcomes up to 25 steps in advance.

  2. Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Relational hidden states anchored to environment states are what let a model-free RL agent plan, and a free-slot control without that anchoring shows no planning signatures.

Pith tools