Pith. sign in

REVIEW 14 cited by

Emergent Linear Representations in World Models of Self-Supervised Sequence Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.00941 v2 pith:IZKNQXEC submitted 2023-09-02 cs.LG

classification cs.LG
keywords modelslinearmodelrepresentationsboardcolourinternalsequence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023). In this work, we provide evidence of a closely related linear representation of the board. In particular, we show that probing for "my colour" vs. "opponent's colour" may be a simple yet powerful way to interpret the model's internal state. This precise understanding of the internal representations allows us to control the model's behaviour with simple vector arithmetic. Linear representations enable significant interpretability progress, which we demonstrate with further exploration of how the world model is computed.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How are linear representations learned? Exact solutions to the dynamics of abstraction

    cs.LG 2026-07 conditional novelty 8.0 of 10

    Exact solutions show abstraction is set by input/target geometry, rises with depth, peaks under small init, and is attenuated by nonlinearities—improving LLM probes via GELU ablation.

  2. Context Is King: How In-Context Specification Shapes the Geometry of Concepts

    cs.LG 2026-07 accept novelty 7.5 of 10

    In capable Gemma and Qwen models, declarative in-context rules set the relational geometry and topology type that the model represents and causally uses, overriding strong pretrained priors.

  3. When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary

    cs.LG 2026-07 conditional novelty 7.0 of 10

    High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.

  4. A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A single linear direction in an observer model's residual stream detects contextual hallucinations, transfers across models and datasets, and causally steers generation hallucination rates.

  5. Convergent Linear Representations of Emergent Misalignment

    cs.LG 2025-06 conditional novelty 7.0 of 10

    A single activation direction extracted from one misaligned fine-tune can ablate emergent misalignment across models trained on different datasets with different LoRA setups.

  6. Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

    cs.LG 2026-07 conditional novelty 6.5 of 10

    LLM-as-judge scoring biases concentrate in low-dimensional, type-specific activation subspaces that support bidirectional causal steering and cross-domain failure prediction.

  7. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  8. Geometry-Guided Constraint Learning for LLM Safety Classification

    cs.AI 2026-06 conditional novelty 6.0 of 10

    Sparse-autoencoder features reduce the number of safety constraints needed to two for most BeaverTails categories, and a three-phase-trained cone constraint modestly beats a flat polytope on in-distribution Qwen3.5-9B...

  9. Distinct Computations Emerge From Compositional Curricula in In-Context Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    When transformer models see easy component examples before a harder combined math problem in one prompt, they solve unseen versions of the combined problem and store intermediate steps internally, unlike models traine...

  10. Model Organisms for Emergent Misalignment

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Emergent misalignment can be induced in small models via a single rank-1 LoRA adapter, and its onset coincides with a phase transition in the adapter's weight direction.

  11. Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task

    cs.LG 2025-10 conditional novelty 5.0 of 10

    In an RL-trained transformer solving adjacent-swap sorting, larger embedding dimensions improve the monotonicity of an internal order-encoding in attention weights and the match to a largest-adjacent-difference swap r...

  12. The Geometry of Harmfulness in LLMs through Subconcept Probing

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Fifty-five harmfulness subconcept directions in Llama-3.1-8B-Instruct form a nearly rank-1 subspace, and steering along the dominant direction cuts jailbreak success but costs accuracy and fails on Qwen.

  13. Large Language Models and Emergence: A Complex Systems Perspective

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A perspective paper arguing that LLM emergence claims are incomplete without evidence of internal coarse-grained representations, and that LLMs have not shown emergent intelligence.

  14. Towards Atoms of Large Language Models

    cs.CL 2025-09 reject novelty 4.0 of 10

    The authors define 'atoms' as sparse, near-orthogonal directions in LLM representations under a data-adaptive inner product, and show threshold-activated sparse autoencoders can recover them with about 99.9% reconstru...

Pith tools