Pith. sign in

REVIEW 5 cited by

Thinking Fast and Slow with Deep Learning and Tree Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1705.08439 v4 pith:MVHWZ6DN submitted 2017-05-23 cs.AI cs.LG

classification cs.AIcs.LG
keywords searchnetworkneuralplanstreedeeplearningplanning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sequential decision making problems, such as structured prediction, robotic control, and game playing, require a combination of planning policies and generalisation of those plans. In this paper, we present Expert Iteration (ExIt), a novel reinforcement learning algorithm which decomposes the problem into separate planning and generalisation tasks. Planning new policies is performed by tree search, while a deep neural network generalises those plans. Subsequently, tree search is improved by using the neural network policy to guide search, increasing the strength of new plans. In contrast, standard deep Reinforcement Learning algorithms rely on a neural network not only to generalise plans, but to discover them too. We show that ExIt outperforms REINFORCE for training a neural network to play the board game Hex, and our final tree search agent, trained tabula rasa, defeats MoHex 1.0, the most recent Olympiad Champion player to be publicly released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Partial Label Learning for Automated Theorem Proving

    cs.LO 2025-07 conditional novelty 6.0 of 10

    Using partial label learning losses, especially Libra and meritocratic losses, improves the plCoP theorem prover's solved-problem count by roughly 14 to 28 percent over the MCTS-imitation baseline.

  2. Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening

    cs.LG 2025-06 conditional novelty 6.0 of 10

    GRPO's rank bias reinforces likely answers and neglects rare correct proofs; an unlikeliness reward that down-weights likely correct samples improves pass@N in formal theorem proving.

  3. Inference Scaling Reshapes AI Governance

    cs.CY 2025-02 conditional novelty 6.0 of 10

    If frontier AI progress shifts from pre-training compute to inference-time compute, AI governance must be rebuilt around deployment-time capabilities and transparency, with different implications depending on whether ...

  4. AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Vanilla AlphaZero fails to preserve oracle-optimal trajectories in Connect Four and Chomp; AZAL with oracle policy supervision substantially raises oracle consistency, reaching perfect full-game play on Chomp 10×11.

  5. CogniPlay: a work-in-progress Human-like model for General Game Playing

    cs.AI 2025-07 unverdicted novelty 4.0 of 10

    A position paper proposing CogniPlay, a dual-process architecture for human-like general game playing, with no implementation or evaluation yet.

Pith tools