REVIEW 5 cited by
Giraffe: Using Deep Reinforcement Learning to Play Chess
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This report presents Giraffe, a chess engine that uses self-play to discover all its domain-specific knowledge, with minimal hand-crafted knowledge given by the programmer. Unlike previous attempts using machine learning only to perform parameter-tuning on hand-crafted evaluation functions, Giraffe's learning system also performs automatic feature extraction and pattern recognition. The trained evaluation function performs comparably to the evaluation functions of state-of-the-art chess engines - all of which containing thousands of lines of carefully hand-crafted pattern recognizers, tuned over many years by both computer chess experts and human chess masters. Giraffe is the most successful attempt thus far at using end-to-end machine learning to play chess.
Forward citations
Cited by 5 Pith papers
-
Learning to Plan via Supervised Contrastive Learning and Strategic Interpolation: A Chess Case Study
A transformer encoder trained with supervised contrastive learning on Stockfish win probabilities, combined with an advantage-axis cosine score and 6-ply beam search, reaches an estimated Elo of 2593.
-
Reinforcement Learning for Game-Theoretic Resource Allocation on Graphs
DQN and PPO policies, trained with a graph-aware action mask called an action-displacement adjacency matrix, beat random and greedy baselines on multi-step Colonel Blotto games on small graphs.
-
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization
SPO fine-tunes an LLM to rewrite drug molecules into analogs that score higher on docking, drug-likeness, solubility, and synthesizability, beating baselines on two protein targets.
-
Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks
Graph-based RL agents can solve logic puzzles larger than anything seen in training, with graph structure, reward design, and recurrence each changing how far extrapolation goes.
-
Exploring the Performance of Deep Residual Networks in Crazyhouse Chess
A residual-network policy/value engine trained on Stockfish self-play reached a 2278 Lichess rating in Crazyhouse, though the evidence is anecdotal and baseline comparisons are missing.
Discussion (0). Continue with ORCID to comment.