REVIEW 9 cited by
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language models show a surprising range of capabilities, but the source of their apparent competence is unclear. Do these networks just memorize a collection of surface statistics, or do they rely on internal representations of the process that generates the sequences they see? We investigate this question by applying a variant of the GPT model to the task of predicting legal moves in a simple board game, Othello. Although the network has no a priori knowledge of the game or its rules, we uncover evidence of an emergent nonlinear internal representation of the board state. Interventional experiments indicate this representation can be used to control the output of the network and create "latent saliency maps" that can help explain predictions in human terms.
Forward citations
Cited by 9 Pith papers
-
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
LLMs' source-attribution ability is not fixed: it flips with conversational memory structure, and corrective feedback can invert judgments or sever confidence from accuracy.
-
When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary
High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.
-
Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study
In information-matched tiny transformers, zero-shot compositional binding fails for every route, while few-shot efficiency is governed by input-pathway sharing and code readability.
-
Transformers converge to invariant algorithmic cores
Trained transformers contain low-dimensional causal subspaces — algorithmic cores — that recur across runs and scales and can be extracted, characterized, and steered.
-
What Does it Mean for a Neural Network to Learn a "World Model"?
Defines a world model as a simple commutative-diagram factorization through an intermediate representation, with conditions that the model be learned and emergent rather than inherited from input or output.
-
The Limits of Predicting Agents from Behaviour
Observed behavior only weakly constrains an intentional agent's choices under distribution shift, and its perceived fairness and harm cannot be identified from behavior alone.
-
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
A survey showing that common systematic generalization benchmarks measure behavioural systematicity, not the representational systematicity that Fodor and Pylyshyn's challenge requires, and mapping them onto Hadley's ...
-
Linear Spatial World Models Emerge in Large Language Models
Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.
-
Learning Implicit Causal World Models from Multi-Agent Demonstrations
Random action noise improves multi-agent world-model OOD accuracy, but the paper's own common-cause analysis shows the causal graph contributes little at matched data.
Discussion (0). Sign in to comment.