Jacobian-lens analysis of two looped transformers finds workspace-like representations survive weight tying, but reads and writes are bounded by supervised loop checkpoints or a ~2-recurrence transport window.
Verbalizable Representations Form a Global Workspace in Language Models
8 Pith papers cite this work. Polarity classification is still indexing.
abstract
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.
years
2026 8representative citing papers
A Jacobian-lens audit predicts model-level relearning recovery in LLM unlearning but cannot pick which facts return and backfires when used as a training penalty.
Language is best understood as a shared codebook compressor; for multimodal models, it belongs at the boundary (input/output) and not as an internal representation, a design principle supported by experiments on cue integration and modality neglect.
Attention gathering, not a gating mechanism, brings a task-relevant latent variable into a form the model can report on at the queried position.
A single low-rank, sparsely trained lens family decodes residual, attention, and MLP states in models up to 70B parameters, revealing that visible and causally effective locations for a behavior can differ.
Logic pre-pretraining, training a small LM on next-step formal derivations before natural language, accelerates skill acquisition and improves pruning robustness at a 100B-token scale.
A projection of LLM internal states, trained on some runs, predicts repetition on held-out runs and can be steered to change repetition; the headline entropy maximum is a reparameterization of an occupancy split.
In two tiny engineered models, a verified internal state was used to bias word choices in generated text, and a detector recovered that state from the text even when the final answer was unchanged.
citing papers explorer
-
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
Jacobian-lens analysis of two looped transformers finds workspace-like representations survive weight tying, but reads and writes are bounded by supervised loop checkpoints or a ~2-recurrence transport window.
-
Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning
A Jacobian-lens audit predicts model-level relearning recovery in LLM unlearning but cannot pick which facts return and backfires when used as a training penalty.
-
Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition
Language is best understood as a shared codebook compressor; for multimodal models, it belongs at the boundary (input/output) and not as an internal representation, a design principle supported by experiments on cue integration and modality neglect.
-
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Attention gathering, not a gating mechanism, brings a task-relevant latent variable into a form the model can report on at the queried position.
-
Interpreting Language Model Hidden States at Scale
A single low-rank, sparsely trained lens family decodes residual, attention, and MLP states in models up to 70B parameters, revealing that visible and causally effective locations for a behavior can differ.
-
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
Logic pre-pretraining, training a small LM on next-step formal derivations before natural language, accelerates skill acquisition and improves pruning robustness at a 100B-token scale.
-
Temperature-driven inversion and nonlinear dynamics in ChatGPT-like AIs
A projection of LLM internal states, trained on some runs, predicts repetition on held-out runs and can be steered to change repetition; the headline entropy maximum is a reparameterization of an occupancy split.
-
Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text
In two tiny engineered models, a verified internal state was used to bias word choices in generated text, and a detector recovered that state from the text even when the final answer was unchanged.