Pith. sign in

REVIEW 9 cited by

Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.13382 v5 pith:RFNBILVC submitted 2022-10-24 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords boardemergentgameinternalmodelnetworkrepresentationrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models show a surprising range of capabilities, but the source of their apparent competence is unclear. Do these networks just memorize a collection of surface statistics, or do they rely on internal representations of the process that generates the sequences they see? We investigate this question by applying a variant of the GPT model to the task of predicting legal moves in a simple board game, Othello. Although the network has no a priori knowledge of the game or its rules, we uncover evidence of an emergent nonlinear internal representation of the board state. Interventional experiments indicate this representation can be used to control the output of the network and create "latent saliency maps" that can help explain predictions in human terms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

    cs.AI 2026-07 conditional novelty 7.0 of 10

    LLMs' source-attribution ability is not fixed: it flips with conversational memory structure, and corrective feedback can invert judgments or sever confidence from accuracy.

  2. When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary

    cs.LG 2026-07 conditional novelty 7.0 of 10

    High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.

  3. Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study

    cs.LG 2026-07 accept novelty 6.0 of 10

    In information-matched tiny transformers, zero-shot compositional binding fails for every route, while few-shot efficiency is governed by input-pathway sharing and code readability.

  4. Transformers converge to invariant algorithmic cores

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Trained transformers contain low-dimensional causal subspaces — algorithmic cores — that recur across runs and scales and can be extracted, characterized, and steered.

  5. What Does it Mean for a Neural Network to Learn a "World Model"?

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Defines a world model as a simple commutative-diagram factorization through an intermediate representation, with conditions that the model be learned and emergent rather than inherited from input or output.

  6. The Limits of Predicting Agents from Behaviour

    cs.AI 2025-06 accept novelty 6.0 of 10

    Observed behavior only weakly constrains an intentional agent's choices under distribution shift, and its perceived fairness and harm cannot be identified from behavior alone.

  7. Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A survey showing that common systematic generalization benchmarks measure behavioural systematicity, not the representational systematicity that Fodor and Pylyshyn's challenge requires, and mapping them onto Hadley's ...

  8. Linear Spatial World Models Emerge in Large Language Models

    cs.AI 2025-06 reject novelty 5.0 of 10

    Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.

  9. Learning Implicit Causal World Models from Multi-Agent Demonstrations

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Random action noise improves multi-agent world-model OOD accuracy, but the paper's own common-cause analysis shows the causal graph contributes little at matched data.

Pith tools