REVIEW 2 major objections 3 minor 2 cited by
Belief and reality separate in a query router, not in the stored value.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Belief-reality separation in language models lives in query-position routing subspaces over a frame-agnostic value slot, not in the value representation itself.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection A careful two-locus account of belief–reality separation in LLMs, with a solid router result and a value-slot claim that still needs on-pathway evidence. the 2 major comments →
Belief-reality separation lives in routing over a shared value slot in language models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that belief–reality separation in transformer language models is implemented as routing over a shared value slot. The value slot is a low-rank subspace at the answer token that binds the attributed value; it is frame-agnostic, transferring at 87–95% of its controllable range between belief and reality readouts across Qwen, Mistral, and OLMo-2. The frame is selected instead by a query-position router: a pair of dissociated low-rank subspaces, one for belief and one for reality, that flip a readout to the probe's own value for the selected frame while injecting none of the anchor's color. Asserted and derived belief fill the same slot by two routes, with only the derived r
What carries the argument
The central object is a slot-and-router pair, both low-rank subspaces of the model's residual-stream activations. The value slot holds the attributed value and is found by Distributed Alignment Search (DAS): swapping the slot from a donor item flips the readout to the donor's value, in either frame. The router is a separate subspace at the query position that selects which frame a query reads out; it is identified by splicing a frame anchor into the final query token and observing that the query flips to the probe's own value for that frame, with zero anchor-color injection. A visibility-gated lookback is the derivation route that fills the slot when the belief value must be inferred from wh
Load-bearing premise
The load-bearing premise is that the low-rank subspaces optimized for interchange are the model's own causal mechanisms—directions it actually uses on its forward pass—and not merely expressive directions that can change the output when swapped in.
What would settle it
Run the belief-to-reality cross-frame transfer on a substrate where the same color appears in both frames and the donor is perfectly matched: if the value-slot subspace moves the reality readout at no more than the random-subspace floor, the slot is frame-specific. Equally decisive, project the position-matched router out of an un-intervened forward pass: if the belief readout survives, the router is not on the model's actual pathway.
If this is right
- Interventions aimed at the value token will not find belief–reality separation; the separation sits at the query-position router.
- Asserted and derived beliefs share a binding subspace, so training a control on either route gives partial control over the other; only derived belief is causally sensitive to described visibility.
- Belief and reality are not two ends of one knob but separate, near-orthogonal routers, with the bare query defaulting to reality.
- The mechanism is readable only once models reach roughly 7B parameters; at 3B the behavior is absent or recency-driven, which fixes the substrate for mechanistic study.
Where Pith is reading between the lines
- If the slot-and-router format is general, the same value slot may serve counterfactual, fictional, and temporal frames, each with its own router; the companion paper makes that case.
- A direct test of the routing claim: find a context where two query forms differ only by which frame is selected and use a router trained on one query form; it should transfer to the other surface form, as the paper's paraphrase control suggests.
- The frame-agnostic slot implies that efforts to steer model truthfulness by editing a 'truth direction' may be acting on a router or a slot, depending on where they intervene; the paper's dissociation offers a way to tell which.
- If the scale boundary is a capacity threshold, then smaller models trained on much larger belief-tracking corpora should still fail to form routers; if it is a data-availability limit, they should acquire them. The paper does not test this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks where belief-reality separation lives in transformer language models. Using Distributed Alignment Search on Qwen2.5, Mistral-7B, and OLMo-2, it argues that the attributed color/value is held in a low-rank 'value slot' that is frame-agnostic, while a query-position 'router' selects whether the query reads the belief frame or the reality frame. It further distinguishes two filling routes: asserted beliefs bind directly, while derived beliefs arrive via a visibility-gated lookback. Behavioral controls de-confound recency and copying; a scale analysis places the relevant behavior between 3B and 7B. The paper is explicit that cross-context generality is provided by a same-authors companion paper.
Significance. If the main claim holds, this is a substantive advance: it separates value binding from frame selection in a causally localizable way and gives a mechanistic counterpart to behavioral ToM critiques. The methodology is careful in several respects: held-out DAS item tuples, random-subspace floors, copy-proof two-value stimuli, pre-fixed dissociation bars, and unusually honest disclosure of failed and inconclusive controls. The release of the Mental Spaces Corpus and build code also supports reproducibility. The main weakness is that the shared-slot half of the central claim lacks an on-distribution ablation comparable to the one supplied for the router, leaving an expressive-fit alternative open.
major comments (2)
- [§3.4, §4.4, Table 1] The shared-slot half of the central claim lacks the on-pathway test that protects the router half. Table 2 ablates the position-matched router on un-intervened forward passes, but no equivalent projection ablation is reported for the value-token subspace. Cross-frame transfer (0.87–0.95) is equally consistent with one shared slot and with two frame-specific value circuits converging on a common color/value readout direction—the target DAS would be most likely to find, since it most efficiently moves the color readout. Because §3.4 concedes that DAS subspaces can be expressive fits, the title/conclusion 'routing over a shared value slot' outruns the evidence. Please add an on-distribution ablation of the value-slot subspace, e.g., projecting it out of un-intervened belief and reality passes and showing the readout collapses, or explicitly demote the shared-slot claim to a transfer/compact
- [§4.4, App. F] The secondary preservation control fails on Qwen and is uninformative on Mistral (0.667 vs 0.633 random baseline). The manuscript attributes this to slightly off-manifold anchor contexts, but that explanation is post hoc, and this control is the one most directly testing whether the router splice is a pure re-router rather than a perturbation. Since the central negative contrast—'inject none of the donor's color' and 'routes rather than injects'—rests on the same splice, the failure should be reported in the main text as a caveat on the router mechanism, not only in App. F/Limitations, unless a positive preservation result is supplied on some substrate.
minor comments (3)
- [§4.1 vs App. C] The abstract and §4.1 say the behavior 'emerges between 3B and 7B', but App. C reports that the same few-shot protocol largely dissolves the boundary for the visibility task (Qwen2.5-3B reaches 0.87, order-independent). Please qualify the 'emergence' wording so it is not read as a raw capability claim.
- [§4.4, Limitations] The §4.4 readout gates are item-sample sensitive at 7B: only one of five stimulus samples clears both clean-accuracy gates at 7B, and the 7B rows are single-sample corroboration. This is disclosed in Limitations, but it should also appear in the caption or main text near Table 1 so readers do not assign the 7B rows the same weight as the 14B primary substrate.
- [Figure 3] The asserted/derived transfer comparison mixes conditions across rows: Qwen and Mistral rows are base-model completion prompts, while the OLMo-2 row is instruction-tuned completion. Please add a column or footnote making this explicit in the figure itself, since the near-symmetry claim for OLMo-2 is not directly architecture-comparable to the base rows.
Circularity Check
No circular derivation; central result rests on held-out DAS transfers and an on-pathway router ablation, with only a minor same-author citation for cross-context generality.
full rationale
The paper's claimed derivation chain is not circular. The central mechanisms are established by DAS interchange on held-out item tuples disjoint from optimization (72/90/96 items), with random-subspace floors and a clean-pass projection ablation for the router (Table 2). The cross-frame transfer of the value slot is an empirical, not definitional, result: values are 0.87-0.95, not 1.0, and it is replicated with materials instead of colors and across five layers, so it is not forced by the shared color output alone. The asserted/derived shared subspace similarly transfers at 0.74-0.99 across held-out items. The only self-citation is the companion paper (Steele et al., 2026) for cross-context generality across counterfactual/fictional/temporal spaces; the paper explicitly states this is imported for context and that its own evidence is confined to the belief-reality axis, so the self-citation is not load-bearing for the main claim. The gap identified by the skeptic—no clean-pass ablation for the value slot—is a mechanistic identifiability limitation (the DAS target could be a generic value direction rather than the model's binding site), which the paper partially acknowledges via Sutter et al. (2025); this is a correctness risk, not a circularity. Score 2 reflects the minor non-load-bearing self-citation.
Axiom & Free-Parameter Ledger
free parameters (3)
- DAS subspace rank =
8
- subspace layer position =
selected across 0.25–0.75 depth
- router anchor contexts =
not specified beyond 'a separate context asserting a belief/reality'
axioms (4)
- domain assumption Interchange interventions reveal causal mechanisms
- domain assumption Belief and reality values are linearly encoded in low-rank residual stream subspaces at the answer/query token
- domain assumption The lookback mechanism of Prakash et al. (2026) is the correct account of derived belief derivation
- domain assumption The de-confounded stimuli isolate belief tracking from copying and recency
invented entities (3)
-
value slot
no independent evidence
-
router / belief-space index
no independent evidence
-
nested / order-indexed router
no independent evidence
Cite this review
Pith. "Pith review of Belief-reality separation lives in routing over a shared value slot in language models." pith.science (2026). https://pith.science/paper/4XSFP7OF
@misc{pith2026260711945,
author = {Pith},
title = {Pith review of: Belief-reality separation lives in routing over a shared value slot in language models},
year = {2026},
howpublished = {\url{https://pith.science/paper/4XSFP7OF}},
note = {Machine review of arXiv:2607.11945}
}
read the original abstract
Capable language models hold what a character believes apart from what is true: told "Anna believes the cup is blue; in reality it is red," they answer blue about Anna and red about the world. Where in the computation does that separation live? We show it rests on two separable mechanisms at two positions. A generic value slot binds the attributed value. A router at the query position selects which frame, the character's belief or reality, a query reads out. Two routes fill the slot: an asserted belief, whose value the text supplies, binds in directly; a derived belief, whose value must be inferred from what the character could see, arrives by a visibility-gated lookback. A subspace trained on either route steers the other, and only the derived route depends on described visibility. The slot itself carries no belief-reality tag: intervening on it moves a reality readout as strongly as a belief one. The separation lives instead in a dissociated pair of routing subspaces, which flip a query between frames without injecting the donor's value. These results hold across three architectures, on stimuli de-confounded against theory-of-mind-benchmark shortcuts; the behavior itself emerges between 3B and 7B across five model families. This paper develops the single belief-reality axis in depth; a companion paper shows the same slot-and-router format is shared across the other non-actual contexts a sentence can open (counterfactual, fictional, temporal).
Figures
Forward citations
Cited by 2 Pith papers
-
One mechanism for many mental spaces: a shared router over a value slot in language models
A low-rank DAS subspace trained on one discourse-space type (counterfactual, belief, fictional, or temporal) causally controls the others across three LM families via a shared router over a value slot.
-
One mechanism for many mental spaces: a shared router over a value slot in language models
A subspace trained to control one mental-space builder also controls others, indicating a shared router/slot mechanism across counterfactual, belief, fictional, and temporal spaces in LMs.
Reference graph
Works this paper leans on
-
[7]
arXiv preprint arXiv:2607.10248
One mechanism for many mental spaces: a shared router over a value slot in language models. arXiv preprint arXiv:2607.10248. James W. A. Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, Michael S. A. Graziano, and Cristina Becchio
-
[8]
arXiv preprint arXiv:2602.16085
Language Statistics and False Belief Reasoning: Evidence from 41 Open-Weight LMs. arXiv preprint arXiv:2602.16085. Tomer Ullman
-
[9]
arXiv preprint arXiv:2302.08399
Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks. arXiv preprint arXiv:2302.08399. Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, and Yao Fu
-
[10]
arXiv preprint arXiv:2404.15574
Retrieval Head Mechanistically Explains Long-Context Factuality. arXiv preprint arXiv:2404.15574. Yiwei Wu, Atticus Geiger, and Raphaël Millière. 2025a. How Do Transformers Learn Variable Binding in Symbolic Programs?. International Conference on Machine Learning. Yufan Wu, Yinghui He, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. 2023b. Hi- ToM...
-
[11]
Yuheng Wu, Wentao Guo, Zirui Liu, Heng Ji, Zhaozhuo Xu, and Denghui Zhang. 2025b. Sensitivity Meets Sparsity: The Impact of Extremely Sparse Parameter Patterns on Theory-of- Mind of Large Language Models. arXiv preprint arXiv:2504.04238. Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah D. Goodman. 2023a. Interpretability at Scale: I...
-
[12]
tuning installs the capability
and where those critiques are sharpest. We measure it with a copy-proof false-belief stimulus: a character holds a belief, a second character knows the truth, and the model is asked what the second thinks the first believes, with both colors present and clause order randomized so that neither copying nor recency can pass. Chance between the two present va...
2024
-
[2019]
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP-IJCNLP):5872–5877
Revisiting the Evaluation of Theory of Mind through Question Answering. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP-IJCNLP):5872–5877. Benjamin A. Levinstein and Daniel A. Herrmann
2019
-
[2022]
Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP):3762–3780
Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP):3762–3780. Melanie Sclar, Sachin Kumar, Peter West, Alane Suhr, Yejin Choi, and Yulia Tsvetkov
2022
-
[2023]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP):14397–14413
FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP):14397–14413. Najoung Kim and Sebastian Schuster
2023
-
[2024]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing:17468–17493
Representational Analysis of Binding in Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing:17468–17493. Jiahai Feng and Jacob Steinhardt
2024
-
[2025]
arXiv preprint arXiv:2509.04466
Just-in-time and Distributed Task Representations in Language Models. arXiv preprint arXiv:2509.04466. Spotlight, NeurIPS 2025 Mechanistic Interpretability Workshop. 14 Samuel Marks and Max Tegmark
arXiv 2025
-
[2026]
International Conference on Learning Representations
Language Models Use Lookbacks to Track Beliefs. International Conference on Learning Representations. arXiv:2505.14685v3. Matthew Riemer, Zahra Ashktorab, Djallel Bouneffouf, Payel Das, Miao Liu, Justin D. Weisz, and Murray Campbell
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.