Pith. sign in

REVIEW 5 cited by

Giraffe: Using Deep Reinforcement Learning to Play Chess

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1509.01549 v2 pith:CNCYNLEY submitted 2015-09-04 cs.AI cs.LGcs.NE

classification cs.AIcs.LGcs.NE
keywords chessgiraffelearningevaluationhand-craftedfunctionsknowledgemachine
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This report presents Giraffe, a chess engine that uses self-play to discover all its domain-specific knowledge, with minimal hand-crafted knowledge given by the programmer. Unlike previous attempts using machine learning only to perform parameter-tuning on hand-crafted evaluation functions, Giraffe's learning system also performs automatic feature extraction and pattern recognition. The trained evaluation function performs comparably to the evaluation functions of state-of-the-art chess engines - all of which containing thousands of lines of carefully hand-crafted pattern recognizers, tuned over many years by both computer chess experts and human chess masters. Giraffe is the most successful attempt thus far at using end-to-end machine learning to play chess.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A transformer encoder trained with supervised contrastive learning on Stockfish win probabilities, combined with an advantage-axis cosine score and 6-ply beam search, reaches an estimated Elo of 2593.

  2. Reinforcement Learning for Game-Theoretic Resource Allocation on Graphs

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DQN and PPO policies, trained with a graph-aware action mask called an action-displacement adjacency matrix, beat random and greedy baselines on multi-step Colonel Blotto games on small graphs.

  3. DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization

    cs.LG 2025-02 conditional novelty 5.0 of 10

    SPO fine-tunes an LLM to rewrite drug molecules into analogs that score higher on docking, drug-likeness, solubility, and synthesizability, beating baselines on two protein targets.

  4. Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Graph-based RL agents can solve logic puzzles larger than anything seen in training, with graph structure, reward design, and recurrence each changing how far extrapolation goes.

  5. Exploring the Performance of Deep Residual Networks in Crazyhouse Chess

    cs.LG 2019-08 conditional novelty 4.0 of 10

    A residual-network policy/value engine trained on Stockfish self-play reached a 2278 Lichess rating in Crazyhouse, though the evidence is anecdotal and baseline comparisons are missing.

Pith tools