REVIEW 3 cited by
Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language models have shown unprecedented capabilities, sparking debate over the source of their performance. Is it merely the outcome of learning syntactic patterns and surface level statistics, or do they extract semantics and a world model from the text? Prior work by Li et al. investigated this by training a GPT model on synthetic, randomly generated Othello games and found that the model learned an internal representation of the board state. We extend this work into the more complex domain of chess, training on real games and investigating our model's internal representations using linear probes and contrastive activations. The model is given no a priori knowledge of the game and is solely trained on next character prediction, yet we find evidence of internal representations of board state. We validate these internal representations by using them to make interventions on the model's activations and edit its internal board state. Unlike Li et al's prior synthetic dataset approach, our analysis finds that the model also learns to estimate latent variables like player skill to better predict the next character. We derive a player skill vector and add it to the model, improving the model's win rate by up to 2.6 times.
Forward citations
Cited by 3 Pith papers
-
One mechanism for many mental spaces: a shared router over a value slot in language models
A subspace trained to control one mental-space builder also controls others, indicating a shared router/slot mechanism across counterfactual, belief, fictional, and temporal spaces in LMs.
-
Three-Body Alignment: Aligning Chess Agent with Human Reasoning through Reranked Rationale
Reranking retrieved grandmaster rationales by FEN similarity raises a chess LLM's semantic alignment with grandmaster explanations from 0.61 to 0.73 cosine similarity, while reducing tactical quality.
-
Linear Spatial World Models Emerge in Large Language Models
Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.
Discussion (0). Continue with ORCID to comment.