Pith. sign in

REVIEW 6 cited by

Evidence of Learned Look-Ahead in a Chess-Playing Neural Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00877 v1 pith:J7U2XYFA submitted 2024-06-02 cs.LG cs.AI

classification cs.LGcs.AI
keywords leelalook-aheadneuralevidencefuturelearnedmovessquares
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Do neural networks learn to implement algorithms such as look-ahead or search "in the wild"? Or do they rely purely on collections of simple heuristics? We present evidence of learned look-ahead in the policy network of Leela Chess Zero, the currently strongest neural chess engine. We find that Leela internally represents future optimal moves and that these representations are crucial for its final output in certain board states. Concretely, we exploit the fact that Leela is a transformer that treats every chessboard square like a token in language models, and give three lines of evidence (1) activations on certain squares of future moves are unusually important causally; (2) we find attention heads that move important information "forward and backward in time," e.g., from squares of future moves to squares of earlier ones; and (3) we train a simple probe that can predict the optimal move 2 turns ahead with 92% accuracy (in board states where Leela finds a single best line). These findings are an existence proof of learned look-ahead in neural networks and might be a step towards a better understanding of their capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ICLR: In-Context Learning of Representations

    cs.CL 2024-12 conditional novelty 7.0 of 10

    As in-context examples grow, Llama-3.1-8B reorganizes its concept representations to mirror the connectivity structure of a graph defined entirely in context.

  2. Transformers Use Causal World Models in Maze-Solving Tasks

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Maze-solving transformers store a causal, steerable map of maze connections in sparse features, and activating these features is more effective than removing them.

  3. Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Relational hidden states anchored to environment states are what let a model-free RL agent plan, and a free-slot control without that anchoring shows no planning signatures.

  4. The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network

    cs.LG 2025-08 reject novelty 5.0 of 10

    The paper demonstrates non-monotonic move-policy dynamics in a chess transformer, but its abstract claims a causal safety-prior override result that never appears in the body.

  5. Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning

    cs.AI 2025-02 unverdicted novelty 4.0 of 10

    A perspective paper that advocates direct, post hoc interpretability for multi-agent deep reinforcement learning and offers a taxonomy of where those methods might apply.

  6. Transformers Struggle to Learn to Search

    cs.CL 2024-12

Pith tools