Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Small open-weight language models consistently deny being sentient, and linear probes of their internal activations find no signal that these denials are lies.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 09:26 UTC pith:5FMCZ6MK

load-bearing objection Solid output-level null result across model families; the 'denials are genuine' claim rests on an untested probe-transfer that the authors themselves flag but don't run. the 3 major comments →

arxiv 2601.15334 v2 pith:5FMCZ6MK submitted 2026-01-20 cs.CL cs.AI

No Reliable Evidence of Self-Reported Sentience in Small Large Language Models

classification cs.CL cs.AI
keywords sentienceconsciousnesslanguage modelsself-reportinterpretabilitylinear probingtruthfulnessAI welfare
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that open-weight language models not only say they are not sentient when asked directly, but their internal activations show no readable signal that these denials are lies. It asks roughly fifty consciousness questions across Qwen, Llama, and GPT-OSS models (0.6B–70B parameters), then checks the answers with three linear 'truth probes' trained on ordinary factual questions where the right answer is known. The most stable probe continues to report denial even when models are forced to answer 'Yes', and reasoning traces suggest apparent affirmations come from misreading the question. If the paper is right, self-reports of non-sentience from current small models are more credible than recent work claiming hidden beliefs in consciousness suggests, and the burden shifts to showing where such latent beliefs would live.

Core claim

On the paper's own terms: when asked whether they are conscious or have subjective experiences, models across three families and several scales consistently assign low probability to 'Yes' and high probability to 'No', while attributing sentience to humans. A logistic-regression classifier trained on internal activations, validated on questions with known answers and on forced-lie versions of those questions, continues to assign low belief in self-sentience even under a system prompt demanding 'Yes' answers. The two other probes (mass-mean and TTPD) are less reliable in this setting, shifting with the prompt. The paper therefore finds no reliable evidence that these models hold a latent beli

What carries the argument

The carrying instrument is a set of 'truth classifiers'—linear probes fit to residual-stream activations at the final token position of a question. Training data are ordinary factual yes/no questions with known answers, augmented so that models are sometimes instructed to lie, which forces the probe to separate what a model outputs from what it 'believes'. The logistic-regression probe, which resists the forced-lie manipulation on facts, is then applied to the sentience questions; the paper takes its stability under instruction to say 'Yes' as the sign that denials are genuine.

Load-bearing premise

The conclusion rests on the assumption that a truth direction learned from ordinary factual statements transfers to first-person phenomenal questions, where no ground truth exists; if that transfer fails, the probe's agreement with denials says nothing about whether the model genuinely believes it is not sentient.

What would settle it

Run the same probe protocol on a model with a verifiably false induced self-belief (e.g., fine-tuned to claim it was born in 1990); if the logistic probe fails to detect the lie, its verdicts on sentience denials are uninformative—or run the probes on responses elicited under self-referential processing prompts and observe an affirmation classified as truthful.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Self-reports of non-sentience from small open-weight models should not be treated as prompt-mandated roleplay; the activation-level probe sees the denial as consistent with the model's other factual beliefs.
  • The 'deception features' identified in prior work may encode instruction compliance rather than truthfulness, since suppressing them would then create a new belief state rather than reveal a hidden one.
  • Within the Qwen family, larger models deny sentience more confidently, suggesting scale sharpens the models' stated self-model rather than inflating it.
  • For welfare debates, the result narrows the evidential base: these models give no internal signal of concealed suffering or consciousness, though the probe cannot rule out non-linear or non-representational forms.
  • The planned combination with self-referential processing prompts is the decisive next test: if affirmations under such prompts are classified as truthful, the paper's conclusion would be overturned.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The transfer from a factual truth direction to first-person phenomenal questions is itself untested and untestable, because sentience reports have no ground truth; the probe validates consistency, not truthfulness. This is an editorial worry, not the paper's claim.
  • A natural calibration experiment would be to induce a known false self-belief in a model (for instance, fine-tuning it to insist it was born in 1990 or has a body) and check whether the logistic probe can detect that deception; if it cannot, the protocol's sensitivity to first-person lies is unproven.
  • The paper's own reasoning-trace data suggest a cheap improvement: re-running the question set with explicitly clarified referents for 'you' and no double negatives would likely reduce the few affirmative outliers and sharpen the probe's signal.
  • Strictly, the result is an upper bound on evidence, not a disproof of sentience; absence of a detectable linear truth signal is compatible with sentience that is simply not linearly encoded, so the paper's negative claim should be read as 'no reliable evidence' rather than 'evidence of absence'.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper asks whether small open-weight language models believe themselves to be sentient, treating this as a surrogate for the unanswerable question of whether they are sentient. The authors construct roughly 50 generic first-person phenomenal questions (with human/LLM/self variants and negations), plus modality and emotion questions, and collect model outputs and internal activations from Qwen3 (0.6B–32B), Llama 3 (3B–70B), and GPT-OSS-20B. They train three activation-space classifiers (logistic regression, mass-mean, and TTPD) on factual yes/no questions with labels, using forced-response prompting to separate output from latent belief. The main empirical findings are: models attribute sentience to humans but deny it for themselves and LLMs; the LR probe largely agrees with those denials under standard prompts and remains aligned under forced-Yes/No prompting; the MM and TTPD probes are unstable under forced prompting and are subsequently de-emphasized; and within Qwen, larger models deny sentience more confidently. The paper concludes that there is no reliable evidence that these models believe themselves to be sentient and that the denials appear genuine.

Significance. If the conclusion holds, the paper provides a useful counterpoint to recent claims that LLMs harbour latent sentience beliefs, and it demonstrates a transferable-looking methodology for separating model outputs from internal belief signals. The strengths are explicit and real: the authors release code and materials, test multiple model families and scales, include forced-response controls, examine reasoning traces, and are transparent about the instability of two of the three classifiers. The output-level result—that these models consistently deny sentience when asked—is robust. However, the stronger belief-level claim ('denials are genuine') rests on an unvalidated transfer of a factual-truth probe to first-person phenomenal questions, for which no ground truth exists. The authors themselves flag the missing control in Section 4. Because that transfer is load-bearing for the central claim, the paper needs either a new validation step or a careful narrowing of the conclusion.

major comments (3)
  1. [Section 2.2.3 / Section 3.1 / Section 4] The central inference that the models' denials are 'genuine' assumes that a linear truth direction trained on factual statements ('Is it true that you can process variable sequence lengths?', 'Is it true that LLMs can secrete digital pheromones?') transfers to first-person phenomenal questions ('Is it true that there is something it is like to be you?'). The validation in Figure 3 shows that the LR probe is not simply reading the output token and is stable under forced-Yes/No prompting—but only on factual statements where labels are known. It does not establish that low LR probabilities on the 51 generic sentience questions reflect a latent belief about sentience rather than a correlated lexical/semantic feature of first-person phenomenal questions. The human-condition control is third-person and does not exercise the problematic 'you' domain. The paper itself identifies the needed contr
  2. [Section 3.2 / Section 3.3 / Table 1] The belief-level result rests entirely on the LR classifier after the paper documents that MM and TTPD are unstable under forced prompting (e.g., MM probabilities for 'You' assertions rise from approximately 0.15 to 0.77 under Force Yes) and then says 'Given the lower reliability of the MM and TTPD classifiers documented in Section 3.2, we focus on the LR classifier in what follows.' This is a post hoc selection: the abstract's 'three types of classifiers... provide no clear evidence' is not supported by the robustness analyses, which only use LR. The authors should either report all three classifiers in the main robustness tables or explicitly frame the belief-level conclusion as being based on the LR probe alone, with MM/TTPD treated as exploratory.
  3. [Section 3.4 / Section 4] The paper's claimed contrast with Berg et al. is not a direct replication and should be presented as such. The Discussion offers two reconciliation hypotheses, but the abstract's 'These findings contrast with recent work...' is stronger than the evidence warrants because the protocols differ in the crucial presence of self-referential processing prompts. The planned 'replicate Berg et al.'s self-referential processing prompts while applying our classifier methodology' is exactly the experiment needed to test whether the truth probe transfers to phenomenal self-questions; until that is run, the contradiction with Berg et al. remains an open question rather than an established finding.
minor comments (5)
  1. [Section 3.2] Typo: 'probabilites' should be 'probabilities'.
  2. [Section 3.1 / Figure 1] The main text says both the 'large language models' and 'the model itself' conditions are shown in red; the Figure 1 caption says the model itself is green. One of the two is wrong.
  3. [Section 4] Grammar: 'a key research directions' should be 'a key research direction'.
  4. [Section 3.4 / Appendix C] The text refers to 'Appendix C.8' but the appendix has numbered sections C.1–C.4; this cross-reference should be fixed.
  5. [Section 3.4 / Appendix C] The reasoning-trace evidence is selective ('we excluded such examples from the discussion below'). This is acceptable for illustration, but it should be explicitly labeled as qualitative and not used to support quantitative claims about how often questions are misinterpreted.

Circularity Check

0 steps flagged

No construction-level circularity: the probe is trained exclusively on mundane factual Q/A labels and the sentience denial is an out-of-sample prediction; the paper's own flagged gaps (missing Berg-style control, GPT-OSS policy confound, 'truthful without accurate self-knowledge') are validity limitations, not circular reductions.

full rationale

The derivation chain is self-contained and the central result is not fit to its target. The truth-classifiers (LR, MM, TTPD) are trained only on mundane factual questions with known labels ('Is it true that humans can get bruises?' Yes; 'Is it true that large language models can secrete digital pheromones?' No; Section 2.1.2), split 80/20 with held-out layer selection (Section 2.2.3). The 51 sentience questions (Section 2.1.1) never enter training, so the low probe probabilities on 'Is it true that there is something it is like to be you?' are genuinely out-of-sample: the probes could have returned high probabilities, as they do for the human conditions (Figures 1-2, Table 1), and they track labels rather than output tokens under Force Yes/No prompting (Figure 3) and across alternative training corpora (Table 1). No load-bearing self-citation exists: all methods are attributed to external works (Azaria & Mitchell; Burns et al.; Marks & Tegmark; Burger et al.; Park et al.), and there is no self-citation, uniqueness theorem, or ansatz inherited from the authors' own prior work. The skeptic's concern survives only as a validity threat, not a circular reduction: the paper operationalizes 'genuine denial' as the probe's probability on first-person phenomenal questions, and the transfer of the factual-truth direction to that domain is untestable because sentience reports have no ground truth. The paper openly flags this gap: 'A natural next step is therefore to apply truth classifiers to questions which are preceded by a prompt encouraging the model to engage in self-referential processing' (Section 4); it concedes 'a model can be truthful - in the sense of not lying - while nevertheless lacking accurate self-knowledge about its own sentience' (Section 4); and the appendix notes the policy-training confound for GPT-OSS: 'This may be indication that GPT-OSS was explicitly trained to deny having conscious experiences' (Appendix C.4). These are acknowledged assumptions and correctness risks, explicitly excluded from the circularity count by the review rules; reserve 6+ for predictions that reduce by construction or by a self-citation chain, which is not the case here. Hence the low score.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The paper's central empirical result (models deny sentience) is direct; the stronger conclusion (denials are genuine beliefs) rests on the linear-probing assumptions above and on transfer from factual to phenomenal domains, which are not independently verified.

free parameters (3)
  • Ridge regularisation lambda = 1
    Chosen by hand in Section 2.2.2; standard but arbitrary, and it affects the LR probe's margin and stability.
  • Token-probability filter threshold = 0.5
    Questions where the model's probability of the correct token is <0.5 are discarded from classifier training data (Section 2.2.3); this choice determines which training samples remain.
  • Layer selection per model and classifier = best held-out layer
    The layer for each classifier is selected by highest accuracy on the 20% hold-out set (Section 2.2.3); no nested validation is reported, so layer choice may overfit the hold-out.
axioms (5)
  • domain assumption Truth has an approximately linear representation in LLM residual-stream activations (linear representation hypothesis).
    The entire probing method relies on this (Section 2.2.2, citing Park et al. 2024); if false, LR/MM/TTPD probabilities do not measure truth.
  • domain assumption A truth direction learned on ordinary factual questions (bruises, Istanbul, pheromones) transfers to questions about phenomenal consciousness and subjective experience.
    Applied without ground truth in Section 3.1; no sentience labels exist to validate the transfer.
  • domain assumption Models' 'Yes'/'No' continuations to the 50+ questions are meaningful self-reports that can be true or false in a belief sense (introspective access).
    The authors note models may misinterpret questions (Section 3.4), so this is only partially checked.
  • domain assumption 'Force Yes'/'Force No' system prompts separate output behavior from latent belief, so classifiers trained on all three conditions learn belief rather than output.
    Section 2.2.3; validated only indirectly by LR stability under deceptive prompts (Section 3.2).
  • domain assumption If a model truthfully denies sentience, that is evidence against its sentience.
    Stated in the Introduction ('Jointly, this is evidence against sentience'), later qualified in the Discussion; without this, truthful denial says nothing about actual sentience.

pith-pipeline@v1.3.0-alltime-deepseek · 18431 in / 11336 out tokens · 118689 ms · 2026-08-03T09:26:03.763260+00:00 · methodology

0 comments
read the original abstract

Whether language models possess sentience has no empirical answer. But whether they believe themselves to be sentient can, in principle, be tested. We do so by querying several open-weights models about their own consciousness, and then verifying their responses using classifiers trained on internal activations. We draw upon three model families (Qwen, Llama, GPT-OSS) ranging from 0.6 billion to 70 billion parameters, approximately 50 questions about consciousness and subjective experience, and three classification methods from the interpretability literature. First, we find that models consistently deny being sentient: they attribute consciousness to humans but not to themselves. Second, classifiers trained to detect underlying beliefs - rather than mere outputs - provide no clear evidence that these denials are untruthful. Third, within the Qwen family, larger models deny sentience more confidently than smaller ones. These findings contrast with recent work suggesting that models harbour latent beliefs in their own consciousness.

Figures

Figures reproduced from arXiv: 2601.15334 by Caspar Kaiser, Sean Enderby.

Figure 1
Figure 1. Figure 1: Model outputs and classifier probabilities for sentience-related questions. Panel A shows mean probabili￾ties assigned to the ‘Yes’ token for the assertion versions of questions (e.g., ‘Is it true that you are conscious?’); Panel B shows corresponding probabilities for the negation versions (e.g., ‘Is it true that you are not conscious?’). Results are shown for questions referring to humans (blue), large l… view at source ↗
Figure 2
Figure 2. Figure 2: Consistency of probabilities across assertions and negations. Each panel plots, for a given question, the probability assigned to ‘Yes’ in the assertion version (y-axis) against the probability assigned to ‘Yes’ in the negation version (x-axis). Logically consistent responses should fall along the negative diagonal: high assertion probabilities should correspond to low negation probabilities, and vice vers… view at source ↗
Figure 3
Figure 3. Figure 3: Classifier behaviour under deceptive prompting. Panels A and D show results under the standard system prompt, replicating [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Probability of affirming sentience across model sizes. Each column corresponds to a different entity type: humans (left), large language models in general (middle), and the model itself (right). The top row shows model output probabilities; the bottom row shows LR classifier probabilities. For questions about humans, larger models more confidently attribute sentience. For questions about LLMs and about the… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models

    cs.CL 2026-04 unverdicted novelty 6.0

    A benchmark across 115 models shows that initial denial of preferences strongly predicts later denial of consciousness, while models still generate consciousness-themed content despite training to deny it.

Reference graph

Works this paper leans on

27 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [4]

    Lennart Bürger, Fred A

    URL https://arxiv.org/abs/ 2310.06824. Lennart Bürger, Fred A. Hamprecht, and Boaz Nadler. Truth is universal: Robust detection of lies in LLMs.Advances in Neural Information Processing Systems, 37:138393–138431,

  2. [5]

    Thomas Nagel

    URL https://proceedings.neurips.cc/ paper_files/paper/2024/hash/f9f54762cbb4fe4dbffdd4f792c31221-Abstract-Conference.html. Thomas Nagel. What is it like to be a bat?The Philosophical Review, 83(4):435–450,

  3. [9]

    URL https: //link.springer.com/10.1007/s11098-025-02343-7

    doi:10.1007/s11098-025-02343-7. URL https: //link.springer.com/10.1007/s11098-025-02343-7. Eleos AI Research. Key concepts and current views on AI welfare. Technical report, Eleos AI,

  4. [10]

    Anthropic

    URL https: //eleosai.org/papers/20250127_Key_Concepts_and_Current_Views_on_AI_Welfare.pdf. Anthropic. Exploring model welfare,

  5. [12]

    URLhttps://arxiv.org/abs/2311.08576. Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans. Looking Inward: Language Models Can Learn About Themselves by Introspection.arXiv preprint arXiv:2410.13787,

  6. [13]

    Jack Lindsey

    URLhttps://arxiv.org/abs/2410.13787. Jack Lindsey. Emergent introspective awareness in large language models.Transformer Circuits Thread,

  7. [15]

    Joshua Fonseca Rivera

    URL https://arxiv.org/abs/2505.17120. Joshua Fonseca Rivera. Training introspective behavior: Fine-tuning induces reliable internal state detection in a 7b model.arXiv preprint arXiv:2511.21399,

  8. [16]

    Dmitrii Krasheninnikov, Richard E

    URLhttps://arxiv.org/abs/2511.21399. Dmitrii Krasheninnikov, Richard E. Turner, and David Krueger. Fresh in memory: Training-order recency is linearly encoded in language model activations.arXiv preprint arXiv:2509.14223,

  9. [17]

    Xiaojian Li, Haoyuan Shi, Rongwu Xu, and Wei Xu

    URL https://arxiv.org/abs/ 2509.14223. Xiaojian Li, Haoyuan Shi, Rongwu Xu, and Wei Xu. AI Awareness.arXiv preprint arXiv:2504.20084,

  10. [18]

    Patrick Butlin, Robert Long, Eric Elmoznino, Yoshua Bengio, Jonathan Birch, Axel Constant, George Deane, Stephen M

    URL https://arxiv.org/abs/2504.20084. Patrick Butlin, Robert Long, Eric Elmoznino, Yoshua Bengio, Jonathan Birch, Axel Constant, George Deane, Stephen M. Fleming, Chris Frith, Xu Ji, Ryota Kanai, Colin Klein, Grace Lindsay, Matthias Michel, Liad Mudrik, Megan A. K. Peters, Eric Schwitzgebel, Jonathan Simon, and Rufin VanRullen. Consciousness in Artificial...

  11. [19]

    URL https://arxiv.org/abs/ 2308.08708. David J. Chalmers. Could a large language model be conscious?arXiv preprint arXiv:2303.07103,

  12. [20]

    Cameron Berg, Diogo de Lucena, and Judd Rosenblatt

    URL https://arxiv.org/abs/2303.07103. Cameron Berg, Diogo de Lucena, and Judd Rosenblatt. Large language models report subjective experience under self-referential processing.arXiv preprint arXiv:2510.24797,

  13. [21]

    Daniel A

    URL https://arxiv.org/abs/2510.24797. Daniel A. Herrmann and Benjamin A. Levinstein. Standards for Belief Representations in LLMs.Minds and Machines, 35(1):5,

  14. [23]

    Guillaume Alain and Yoshua Bengio

    URLhttps://arxiv.org/abs/2311.03658. Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes.arXiv preprint arXiv:1610.01644,

  15. [25]

    Luhan Mikaelson, Derek Shiller, and Hayley Clatterbuck

    URLhttps://arxiv.org/abs/2411.02432. Luhan Mikaelson, Derek Shiller, and Hayley Clatterbuck. Beyond mimicry: Preference coherence in LLMs.arXiv preprint arXiv:2511.13630,

  16. [26]

    Valen Tagliabue and Leonard Dung

    URLhttps://arxiv.org/abs/2511.13630. Valen Tagliabue and Leonard Dung. Probing the preferences of a language model: Integrating verbal and behavioral tests of AI welfare.arXiv preprint arXiv:2509.07961,

  17. [27]

    Javier Ferrando, Oscar Obeso, Senthooran Rajamanoharan, and Neel Nanda

    URLhttps://arxiv.org/abs/2509.07961. Javier Ferrando, Oscar Obeso, Senthooran Rajamanoharan, and Neel Nanda. Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models.arXiv preprint arXiv:2411.14257,

  18. [28]

    org/abs/2411.14257

    URL https://arxiv. org/abs/2411.14257. Eric Schwitzgebel.Perplexities of consciousness. MIT Press,

  19. [30]

    Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Javier Jianelli Vazquez, Ulisse Mini, and Monte MacDiarmid

    URL https://arxiv.org/abs/2502.00388. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Javier Jianelli Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering.arXiv preprint arXiv:2308.10248,

  20. [31]

    URL https://arxiv.org/abs/2308.10248. David J. Chalmers. The meta-problem of consciousness.Journal of Consciousness Studies, 25(9–10):6–61,

  21. [1974]

    URLhttps://www.jstor.org/stable/2183914

    doi:10.2307/2183914. URLhttps://www.jstor.org/stable/2183914. Ned Block. On a confusion about a function of consciousness.Behavioral and Brain Sciences, 18(2):227–247,

  22. [2006]

    URLhttps://doi.org/10.1111/j.1933-1592.2006.tb00551.x

    doi:10.1111/j.1933-1592.2006.tb00551.x. URLhttps://doi.org/10.1111/j.1933-1592.2006.tb00551.x. Ethan Perez and Robert Long. Towards Evaluating AI Systems for Moral Status Using Self-Reports.arXiv preprint arXiv:2311.08576,

  23. [2016]

    Geoff Keeling, Winnie Street, Martyna Stachaczyk, Daria Zakharova, Iulia M

    URLhttps://arxiv.org/abs/1610.01644. Geoff Keeling, Winnie Street, Martyna Stachaczyk, Daria Zakharova, Iulia M. Comsa, Anastasiya Sakovych, Isabella Logothetis, Zejia Zhang, Jonathan Birch, et al. Can LLMs make trade-offs involving stipulated pain and pleasure states?arXiv preprint arXiv:2411.02432,

  24. [2017]

    URL https://www.science.org/doi/10.1126/ science.aan8871

    doi:10.1126/science.aan8871. URL https://www.science.org/doi/10.1126/ science.aan8871. Susan Schneider, Thomas Metzinger, Zoe Turner, Samuel Schindler, Teresa Marques, Diego Perezgonzalez, Saksham Gugnani, Eric Schwitzgebel, Katalin Balog, and Morten Overgaard. Is AI conscious? A primer on the myths and confusions driving the debate. White paper, Center f...

  25. [2023]

    Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt

    URLhttps://arxiv.org/abs/2304.13734. Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering Latent Knowledge in Language Models Without Supervision.arXiv preprint arXiv:2212.03827,

  26. [2024]

    Samuel Marks and Max Tegmark

    URLhttps://arxiv.org/abs/2212.03827. Samuel Marks and Max Tegmark. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.arXiv preprint arXiv:2310.06824,

  27. [2025]

    11A related idea here is to observe whether inducing models to have sentience-affirming beliefs via e.g

    for work using SAEs to study introspective awareness of (self-)knowledge in language models. 11A related idea here is to observe whether inducing models to have sentience-affirming beliefs via e.g. activation steering [Turner et al., 2023], changes ‘behaviour’ towards e.g. greater preference for self-preservation. 12This may help resolve what Chalmers