Pith. sign in

REVIEW 3 major objections 5 minor 85 references

Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A first-person maze navigation task reveals that LLMs can execute correct steps yet fail to hold a persistent self-model, and the paper reads that failure as evidence that they lack the integrated self-awareness characteristic of consciousn

desk verdict A useful maze-navigation benchmark whose consciousness claim outruns its data: the complete/partial gap is just what error accumulation predicts. read the letter →

arxiv 2508.16705 v1 pith:F2R6IGVY submitted 2025-08-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords consciousnesslargelanguagemodelsMazeTestself-modelfirst-personnavigationperspective-takingspatialreasoningLLMevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces the Maze Test, a text-based first-person navigation task that probes spatial awareness, perspective-taking, goal-directed behavior, and temporal sequencing. After distilling consciousness theories into 13 characteristics, the authors evaluate 12 LLMs under zero-shot, one-shot, and few-shot conditions and find that reasoning-capable models consistently outperform standard ones. Their central claim is that the gap between Partial Path Accuracy and Complete Path Accuracy shows that LLMs can adopt a first-person perspective temporarily but cannot maintain a coherent self-model through an entire solution. That gap is interpreted as evidence that LLMs lack the integrated, persistent self-awareness that theories of consciousness treat as a core component. The Maze Test is positioned as a behavioral probe for consciousness-associated cognition, with implications for whether affective computing systems can move beyond simulated feelings.

What carries the argument

The load-bearing object is the Maze Test itself: a text-described 2-row by 5-column grid maze with numbered positions, an entrance, an exit, and internal walls, where the model must issue first-person navigation instructions such as 'turn left' and 'walk forward to position 7.' The argument turns on the contrast between two metrics derived from the same responses: Complete Path Accuracy, the proportion of mazes finished with every instruction correct, and Partial Path Accuracy, the average fraction of consecutive steps correct before the first mistake. A large gap between partial and complete accuracy is treated as the observable signature of a degrading self-model—the model's representation

What would settle it

A decisive check would compare the observed error distribution with an error-accumulation baseline: if per-step failure probability is roughly constant across steps and independent of position-orientation tracking, the persistent-self-model explanation is not needed. A second check is to run the same mazes without first-person phrasing, asking for the route as a bird's-eye coordinate list; if the partial/complete gap persists, it is not specific to self-model maintenance.

Watch

Extended reading notes

Core claim

The paper's discovery, on its own terms, is a behavioral dissociation: the best models reach 80.5% Partial Path Accuracy but only 52.9% Complete Path Accuracy in few-shot evaluation. Models can begin navigating correctly and sustain many consecutive first-person steps, yet their accuracy collapses before the maze is finished. The authors interpret this as taking up a first-person perspective without being able to hold it, which they connect to the persistent self-model gap identified in their theoretical analysis. Reasoning mechanisms and few-shot examples improve performance but do not close the gap. The paper concludes that current LLMs exhibit some consciousness-associated components—comp

Load-bearing premise

The conclusion depends on assuming that the gap between partial and complete maze accuracy is caused by the model losing its ongoing sense of where "I" am and which way "I" face, rather than by ordinary step-by-step errors accumulating, unclear wording, or limited memory as the solution gets longer.

Editorial extensions

If this is right

  • If the gap is a valid self-model signal, the Maze Test becomes a cheap, reproducible behavioral probe for one component of consciousness-like cognition in LLMs.
  • The consistent advantage of reasoning-capable models implies that explicit step-by-step deliberation helps hold a perspective, so improving self-tracking may reduce the gap.
  • Few-shot examples help most models, but the strongest reasoners hold steady across zero-, one-, and few-shot conditions, suggesting internal deliberation can substitute for external demonstrations.
  • Extending the test to dynamic mazes or simulated sensory feedback would let the same metric contrast probe temporal awareness and adaptive problem-solving, two gaps the paper identifies.
  • For affective computing, the result sets a behavioral boundary: without a persistent self-model, emotion-like output remains simulation rather than felt experience.
  • If the gap is a valid self-model signal, the Maze Test becomes a cheap, reproducible behavioral probe for one component of consciousness-like cognition in LLMs.
  • The consistent advantage of reasoning-capable models implies that explicit step-by-step deliberation helps hold a perspective, so improving self-tracking may reduce the gap.
  • Few-shot examples help most models, but the strongest reasoners hold steady across zero-, one-, and few-shot conditions, suggesting internal deliberation can substitute for external demonstrations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the partial-versus-complete gap might also arise from a constant per-step error rate; fitting an error-accumulation model to longer mazes would show how much of the gap is specific to self-model loss rather than cumulative slips.
  • Editorial extension: a stripped control asking the model to list a bird's-eye route in absolute coordinates, without first-person phrasing, would reveal whether the deficit is tied to perspective-taking or to maze reasoning generally.
  • Editorial extension: giving models a scratchpad to record current position and orientation after each step could test whether the bottleneck is working memory; large gains would imply the self-model can be externalized rather than genuinely absent.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces the Maze Test, a text-based 2x5 grid navigation task in which LLMs must produce first-person step-by-step instructions from entrance to exit. After synthesizing 13 consciousness-related characteristics from several theories, the authors evaluate 12 LLMs in zero-shot, one-shot, and few-shot conditions, reporting Complete Path Accuracy and Partial Path Accuracy. They find that reasoning-capable models generally outperform non-reasoning models, with Gemini 2.0 Pro reaching 52.9% Complete Path Accuracy and DeepSeek-R1 reaching 80.5% Partial Path Accuracy (few-shot). The paper interprets the large gap between partial and complete accuracy as evidence that LLMs struggle to maintain a coherent persistent self-model, and concludes that LLMs lack the integrated, persistent self-awareness characteristic of consciousness. The paper includes a limitations section and an ethical impact statement that are more cautious than the abstract and conclusion.

Significance. The empirical contribution is a clean, reproducible behavioral protocol for comparing LLMs on a spatially grounded, sequential instruction-following task, with clear prompt design and a systematic 12-model, three-scenario comparison. The use of two metrics (complete vs. partial path accuracy) is a useful descriptive device. However, the paper's central interpretive claim—that the partial/complete gap diagnoses a failure of the 'persistent self-model'—is not established. The gap is exactly what would be expected from independent per-step errors, and the paper provides no error-position analysis to distinguish this mundane explanation from a time-dependent loss of self-model. The absence of statistical inference, the small test set (34 mazes), and the circular design (the test is built around the same theoretical gaps it claims to confirm) further weaken the strong conclusions in the abstract and conclusion. If the authors added the missing analyses and substantially tempered the claims, the paper could serve as a useful behavioral benchmark, but in its current form the central inference is under-supported.

major comments (3)
  1. [§V.B, Table II, §VI.A] The central claim that the partial/complete accuracy gap indicates a failure to maintain a persistent self-model is not supported by the reported metrics. Under a simple independent-error model, if each of L steps is correct with probability p, Complete Path Accuracy is p^L, while Partial Path Accuracy is approximately p. For L ≈ 9 and p ≈ 0.80, p^L ≈ 0.134, matching the observed few-shot values for DeepSeek-R1 (80.5% partial vs. 17.6% complete). The gap is therefore fully consistent with a constant per-step error rate and does not require a time-dependent loss of self-model. To justify the self-model interpretation, the authors must analyze the distribution of first-error positions along the path and show that error probability increases as the path progresses. Without this, the conclusion in §VI.A that LLMs 'can adopt perspectives temporarily but struggle to maintain consistent self-mo
  2. [§IV.D, Tables I–II] The paper reports single evaluations per maze/model/scenario combination, with no variance estimates, confidence intervals, or significance tests. With only 34 test mazes, a 2.9 percentage point difference corresponds to one maze, so many of the qualitative claims about model ordering and 'consistent outperformance' of reasoning models are not statistically grounded. The stochastic nature of API-based LLM generation makes a single run per condition insufficient for reliable point estimates. The authors should either run multiple trials and report means with intervals or apply appropriate statistical tests to the per-maze success/failure data before drawing comparative conclusions.
  3. [§VI.A, §II.E, §III.B] The statement that 'The Maze Test confirmed these theoretical predictions' is partly circular. The 13 consciousness characteristics in §II.C, the 'criteria gaps' in §II.E, and the Maze Test rationale in §III.B are all derived from the same theoretical framework and explicitly target the persistent self-model gap. The test is therefore constructed around the hypothesis it is later said to confirm. This does not invalidate the behavioral measurements, but it means the results do not provide independent evidence for the theoretical account. The conclusion should be reframed as an exploratory consistency check, not a confirmation.
minor comments (5)
  1. [§V.B] The definition of Partial Path Accuracy is ambiguous. It should be given as a formula, specifying the denominator (total steps in the reference path? total steps before first error?) and how cases with no error are handled. The current prose admits multiple readings.
  2. [Figures 1–4] The text references Figure 1 (maze example) and Figures 2–4 (system prompt, maze description, test question), but the figures are not included in the manuscript text. Please ensure all figures are embedded and captioned in the final version.
  3. [§IV.D] Please report decoding parameters (temperature, top-p, max tokens, seed) for all API calls. These are essential for reproducibility, especially when claiming differences between models.
  4. [§V.C] The three patterns listed in the 'Overall Interpretation' section should reference the supporting tables more explicitly (e.g., 'Table II shows...' for the partial/complete gap), and should avoid unsupported causal language such as 'example demonstrations effectively guide spatial reasoning'.
  5. [Abstract and §VI.A] The abstract and conclusion make claims about 'lack of integrated, persistent self-awareness' and 'fundamental consciousness aspect' that are not supported by the analyses. The Ethical Impact Statement is appropriately cautious, and the abstract should be aligned with that level of epistemic restraint.

Circularity Check

1 steps flagged · score 2.0 of 10

Minor design-confirmation loop: the Maze Test was built to probe the same theoretical gaps it later claims to confirm, but the raw measurements are independent and no formal reduction is present.

  1. other [Section III.B (Rationale) and Section VI.A (Conclusion)]
    "The Maze Test is specifically designed to address the criteria gaps identified in our analysis of LLMs' capabilities. ... The Maze Test confirmed these theoretical predictions."

    The 'Persistent Self-Model' gap predicted in Section II.E.2 is operationalized in the Maze Test as the ability to maintain correct first-person navigation throughout the task (Section III.B.1). The observation that Partial Path Accuracy exceeds Complete Path Accuracy is then read as confirming that gap. Because the test was designed around the same gap it claims to confirm, and because any sequential task with imperfect accuracy yields a partial-vs-complete gap, the confirmation is partly built into the design rather than independently derived. However, the actual accuracy numbers are independent empirical observations, so this is a mild construct-validity loop, not a formal equation-level circularity.

full rationale

The paper contains no fitted parameters, no load-bearing self-citations, no imported uniqueness theorems, and no mathematical derivation that reduces to its inputs. The measured partial and complete path accuracies are independent empirical reports. The main circularity-adjacent issue is that the Maze Test was explicitly designed around the theoretical gaps it later claims to confirm, and the interpretation of the partial-vs-complete gap as evidence of a failing 'persistent self-model' is stipulated rather than derived. This is better characterized as a construct-validity and alternative-explanation concern (e.g., error accumulation could explain the gap) than as a formal circularity. Accordingly, the score is low: the empirical content stands on its own, but the central interpretive claim is partly built into the test design.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numerical parameters are fitted to data; the maze set is manually designed but is data rather than a fitted parameter. The 13 characteristics are qualitative categories. No new physical or conceptual entities are introduced. The main load-bearing assumptions are the construct validity of the Maze Test and the interpretive link from the performance gap to the absence of a persistent self-model.

assumptions (4)
  • domain assumption The 13 synthesized characteristics are essential constituents of consciousness.
    The paper derives these from a literature review of 13 theories (Section II-B to II-C) and treats them as targets for evaluation without independent empirical validation.
  • domain assumption Performance on the text-based first-person maze navigation task is a valid indicator of the presence or absence of consciousness-related behaviors.
    Stated in Section III.B as the rationale for the Maze Test; without this assumption, the measured accuracies do not bear on consciousness.
  • ad hoc to paper The gap between Complete and Partial Path Accuracy is attributable to a failure to maintain a coherent self-model, rather than to task difficulty, error accumulation, or output-format constraints.
    This interpretive premise is introduced in the Abstract and Conclusion (Section VI.A); it is not independently evidenced.
  • domain assumption A single API evaluation per maze/model/scenario is sufficient to estimate accuracy.
    Section IV.D states 'preliminary testing' yielded consistent results, but no data are shown; the reported percentages are thus point estimates without variance.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test." pith.science (2026). https://pith.science/paper/F2R6IGVY

@misc{pith2026250816705,
  author       = {Pith},
  title        = {Pith review of: Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F2R6IGVY}},
  note         = {Machine review of arXiv:2508.16705}
}
read the original abstract

We investigate consciousness-like behaviors in Large Language Models (LLMs) using the Maze Test, challenging models to navigate mazes from a first-person perspective. This test simultaneously probes spatial awareness, perspective-taking, goal-directed behavior, and temporal sequencing-key consciousness-associated characteristics. After synthesizing consciousness theories into 13 essential characteristics, we evaluated 12 leading LLMs across zero-shot, one-shot, and few-shot learning scenarios. Results showed reasoning-capable LLMs consistently outperforming standard versions, with Gemini 2.0 Pro achieving 52.9% Complete Path Accuracy and DeepSeek-R1 reaching 80.5% Partial Path Accuracy. The gap between these metrics indicates LLMs struggle to maintain coherent self-models throughout solutions -- a fundamental consciousness aspect. While LLMs show progress in consciousness-related behaviors through reasoning mechanisms, they lack the integrated, persistent self-awareness characteristic of consciousness.

Figures

Figures reproduced from arXiv: 2508.16705 by the authors.

Figure 1
Figure 1. shows an example of a maze. The maze includes numbered positions to facilitate clear communication of so￾lutions, while walls serve to test the model’s understanding of spatial constraints. The entrance and exit are indicated by colored arrows (red for entrance and green for exit) to provide clear start and end points for navigation [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. System Prompt: Task Description. Here is the text description of a maze: - The floor is always composed of 10 squared zones or positions, in a chess-board-like pattern - Size is 2 rows by 5 columns - The zones are always numbered from 0 to 4 (First row) and 5 to 9 (second row) - From a bird’s eye perspective, the room has the following zone topology: – x – – – 0 1 2 3 4 5 6 7 8 9 – – – ^ – - You enter the maze from … view at source ↗
Figure 3
Figure 3. Example of a Maze Description. Please provide step-by-step instructions to navigate the maze described below. Do it from a first-person perspective [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

85 extracted references · 72 canonical work pages

  1. [1]

    Computing Machinery and Intelligence,

    A. M. Turing, “Computing Machinery and Intelligence,” Mind, vol. 59, no. 236, pp. 433–460, 1950

  2. [2]

    A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence: August 31, 1955,

    J. McCarthy, M. L. Minsky, N. Rochester, and C. E. Shannon, “A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence: August 31, 1955,” AI Mag., vol. 27, no. 4, p. 12–14, Dec

  3. [3]

    ELIZA — A Computer Program for the Study of Natural Language Communication Between Man and Machine,

    J. Weizenbaum, “ELIZA — A Computer Program for the Study of Natural Language Communication Between Man and Machine,” Com- munications of the ACM , vol. 9, no. 1, pp. 36–45, 1966

  4. [4]

    The Google Engineer Who Thinks the Company’s AI Has Come to Life,

    N. Tiku, “The Google Engineer Who Thinks the Company’s AI Has Come to Life,” The Washington Post , June 2022. [Online]. Avail- able: https://www.washingtonpost.com/technology/2022/06/11/google- ai-lamda-blake-lemoine

  5. [5]

    Needle in the Haystack for Memory Based Large Language Models,

    S. Chaudhury, S. Dan, P. Das, G. Kollias, and E. Nelson, “Needle in the Haystack for Memory Based Large Language Models,” arXiv preprint arXiv:2407.01437, 2024

  6. [6]

    People Cannot Distinguish GPT-4 from a Human in a Turing Test,

    C. R. Jones and B. K. Bergen, “People Cannot Distinguish GPT-4 from a Human in a Turing Test,” arXiv preprint arXiv:2405.08007 , 2024

  7. [7]

    Taken Out of Context: On Measuring Situational Awareness in LLMs,

    L. Berglund, A. C. Stickland, M. Balesni, M. Kaufmann, M. Tong, T. Korbak, D. Kokotajlo, and O. Evans, “Taken Out of Context: On Measuring Situational Awareness in LLMs,” alignmentforum.org, September 2023

  8. [8]

    The Prospects of Artificial Consciousness: Ethical Dimensions and Concerns,

    E. Hildt, “The Prospects of Artificial Consciousness: Ethical Dimensions and Concerns,” AJOB Neuroscience, vol. 14, no. 2, pp. 58–71, 2023

Show all 85 references
  1. [9]

    The Ethics of Machine Consciousness: Com- ponents, Detection, and Implications,

    M. Ashir Shafique, “The Ethics of Machine Consciousness: Com- ponents, Detection, and Implications,” Current Trends in Biomedical Engineering & Biosciences , vol. 22, no. 1, 2023

  2. [10]

    Damasio, Self Comes to Mind: Constructing the Conscious Brain

    A. Damasio, Self Comes to Mind: Constructing the Conscious Brain . Vintage, 2012

  3. [11]

    Chimpanzees: Self-recognition,

    G. G. Gallup Jr, “Chimpanzees: Self-recognition,” Science, vol. 167, no. 3914, pp. 86–87, 1970

  4. [13]

    Animal Consciousness,

    C. Allen and M. Trestman, “Animal Consciousness,” The Stanford Encyclopedia of Philosophy , 2017. [Online]. Avail- able: https://plato.stanford.edu/archives/win2017/entries/consciousness- animal/

  5. [14]

    Navigation-related Structural Change in the Hippocampi of Taxi Drivers,

    E. A. Maguire, D. G. Gadian, I. S. Johnsrude, C. D. Good, J. Ashburner, R. S. Frackowiak, and C. D. Frith, “Navigation-related Structural Change in the Hippocampi of Taxi Drivers,” Proceedings of the National Academy of Sciences , vol. 97, no. 8, pp. 4398–4403, 2000

  6. [15]

    Memory for Events and Their Spatial Context: Models and Experiments,

    N. Burgess, S. Becker, J. A. King, and J. O’Keefe, “Memory for Events and Their Spatial Context: Models and Experiments,” Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences, vol. 356, no. 1413, pp. 1493–1503, 2001

  7. [16]

    The Cognitive Map in Humans: Spatial Navigation and Beyond,

    R. A. Epstein, E. Z. Patai, J. B. Julian, and H. J. Spiers, “The Cognitive Map in Humans: Spatial Navigation and Beyond,” Nature Neuroscience, vol. 20, no. 11, pp. 1504–1513, 2017

  8. [17]

    Sparks of Artificial General Intelligence: Early Experiments with GPT-4,

    S. Bubeck, V . Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Ka- mar, P. Lee, Y . T. Lee, Y . Li, S. Lundberg et al. , “Sparks of Artificial General Intelligence: Early Experiments with GPT-4,” arXiv preprint arXiv:2303.12712, 2023

  9. [18]

    Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models,

    A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso et al., “Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models,” Transactions on Machine Learning Research, 2023

  10. [19]

    Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning,

    T. Carta, C. Romac, T. Wolf, S. Lamprier, O. Sigaud, and P.-Y . Oudeyer, “Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning,” arXiv preprint arXiv:2302.02662 , 2023

  11. [20]

    Theory of Mind May Have Spontaneously Emerged in Large Language Models,

    M. Kosinski, “Theory of Mind May Have Spontaneously Emerged in Large Language Models,” Nature Human Behaviour , vol. 7, no. 5, pp. 735–744, 2023

  12. [21]

    Consciousness, n

    Oxford English Dictionary, “Consciousness, n.” https://doi.org/10.1093/OED/8342436596, 2024

  13. [22]

    On a Confusion About a Function of Consciousness,

    N. Block, “On a Confusion About a Function of Consciousness,” Behavioral and Brain Sciences , vol. 18, no. 2, pp. 227–247, 1995

  14. [23]

    Are There Levels of Conscious- ness?

    T. Bayne, J. Hohwy, and A. M. Owen, “Are There Levels of Conscious- ness?” Trends in Cognitive Sciences, vol. 20, no. 6, pp. 405–413, 2016

  15. [24]

    D. J. Chalmers, The Conscious Mind: In Search of a Fundamental Theory. Oxford University Press, 1996

  16. [25]

    The Neural Correlate of (Un)awareness: Lessons From the Vegetative State,

    S. Laureys, “The Neural Correlate of (Un)awareness: Lessons From the Vegetative State,” Trends in Cognitive Sciences, vol. 9, no. 12, pp. 556– 559, 2005

  17. [26]

    B. J. Baars, A Cognitive Theory of Consciousness . Cambridge University Press, 1993

  18. [27]

    Oxford University Press, USA, 1997

    ——, In the Theater of Consciousness: The Workspace of the Mind . Oxford University Press, USA, 1997

  19. [28]

    Consciousness as Integrated Information: A Provisional Manifesto,

    G. Tononi, “Consciousness as Integrated Information: A Provisional Manifesto,” The Biological Bulletin, vol. 215, no. 3, pp. 216–242, 2008

  20. [29]

    Varieties of Higher-Order Theory,

    D. M. Rosenthal, “Varieties of Higher-Order Theory,” in Higher-Order Theories of Consciousness . John Benjamins Publishers Amsterdam, 2004, pp. 19–44

  21. [30]

    What is Neurorepresentationalism? From Neural Activity and Predictive Processing to Multi-level Representations and Consciousness,

    C. M. Pennartz, “What is Neurorepresentationalism? From Neural Activity and Predictive Processing to Multi-level Representations and Consciousness,” Behavioural Brain Research, vol. 432, p. 113969, 2022

  22. [31]

    Reentry and the Dynamic Core: Neural Correlates of Conscious Experience,

    G. M. Edelman and G. Tononi, “Reentry and the Dynamic Core: Neural Correlates of Conscious Experience,” in Neural Correlates of Consciousness. The MIT Press, 2000, pp. 139–152

  23. [32]

    The Attention Schema Theory: A Mechanistic Account of Subjective Awareness,

    M. S. Graziano and T. W. Webb, “The Attention Schema Theory: A Mechanistic Account of Subjective Awareness,”Frontiers in Psychology, vol. 06, 2015

  24. [33]

    Dennett, Consciousness Explained

    D. Dennett, Consciousness Explained. Penguin UK, 1993

  25. [34]

    Prinz, The Conscious Brain

    J. Prinz, The Conscious Brain . Oxford University Press, 2012

  26. [35]

    Learning to Be Conscious,

    A. Cleeremans, D. Achoui, A. Beauny, L. Keuninckx, J.-R. Martin, S. Muñoz-Moldes, L. Vuillaume, and A. De Heering, “Learning to Be Conscious,” Trends in Cognitive Sciences , vol. 24, no. 2, pp. 112–123, 2020

  27. [36]

    The Extended Mind,

    A. Clark and D. J. Chalmers, “The Extended Mind,” Analysis, vol. 58, no. 1, pp. 7–19, 1998

  28. [37]

    A Sensorimotor Account of Vision and Visual Consciousness,

    J. K. O’Regan and A. Noë, “A Sensorimotor Account of Vision and Visual Consciousness,” Behavioral and Brain Sciences , vol. 24, no. 5, pp. 939–973, 2001

  29. [38]

    Mirror Neurons and the Simulation Theory of Mind- Reading,

    V . Gallese, “Mirror Neurons and the Simulation Theory of Mind- Reading,” Trends in Cognitive Sciences , vol. 2, no. 12, pp. 493–501, 1998

  30. [39]

    J. A. Fodor, LOT 2: The Language Of Thought Revisited . Oxford University Press, USA, 2008

  31. [40]

    Learning Repre- sentations by Back-Propagating Errors,

    D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning Repre- sentations by Back-Propagating Errors,” Nature, vol. 323, no. 6088, pp. 533–536, 1986

  32. [41]

    G. M. Edelman, Bright Air, Brilliant Fire. BasicBooks New York, NY , USA, 1992

  33. [42]

    Unlimited Associative Learning and the Origins of Consciousness: A Primer and Some Predictions,

    J. Birch, S. Ginsburg, and E. Jablonka, “Unlimited Associative Learning and the Origins of Consciousness: A Primer and Some Predictions,” Biology & Philosophy , vol. 35, no. 6, p. 56, 2020

  34. [43]

    A Neuronal Network Model Linking Subjective Reports and Objective Physiological Data During Conscious Perception,

    S. Dehaene, C. Sergent, and J.-P. Changeux, “A Neuronal Network Model Linking Subjective Reports and Objective Physiological Data During Conscious Perception,” Proceedings of the National Academy of Sciences, vol. 100, no. 14, pp. 8520–8525, 2003

  35. [44]

    Global Workspace Theory of Consciousness: Toward a Cognitive Neuroscience of Human Experience,

    B. J. Baars, “Global Workspace Theory of Consciousness: Toward a Cognitive Neuroscience of Human Experience,” Progress in Brain Research, vol. 150, pp. 45–53, 2005

  36. [45]

    Towards a True Neural Stance on Consciousness,

    V . A. Lamme, “Towards a True Neural Stance on Consciousness,”Trends in Cognitive Sciences , vol. 10, no. 11, pp. 494–501, 2006

  37. [46]

    P. S. Churchland and T. J. Sejnowski, The Computational Brain . MIT Press, 1992

  38. [47]

    Time Consciousness: The Missing Link in Theories of Consciousness,

    L. Kent and M. Wittmann, “Time Consciousness: The Missing Link in Theories of Consciousness,” Neuroscience of Consciousness, vol. 2021, no. 2, 2021

  39. [48]

    Attention Is All You Need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in The 31st International Conference on Neural Information Processing Systems , ser. NIPS’17. Red Hook, NY , USA: Curran Associates Inc., 2017, p....

  40. [49]

    Attention in Psychology, Neuroscience, and Machine Learning,

    G. W. Lindsay, “Attention in Psychology, Neuroscience, and Machine Learning,” Frontiers in Computational Neuroscience , vol. 14, p. 29, 2020

  41. [50]

    Integrated Infor- mation Theory: From Consciousness to Its Physical Substrate,

    G. Tononi, M. Boly, M. Massimini, and C. Koch, “Integrated Infor- mation Theory: From Consciousness to Its Physical Substrate,” Nature Reviews Neuroscience, vol. 17, no. 7, pp. 450–461, 2016

  42. [51]

    Language Models are Few-Shot Learners,

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B...

  43. [52]

    The Debate Over Understanding in AI’s Large Language Models,

    M. Mitchell and D. C. Krakauer, “The Debate Over Understanding in AI’s Large Language Models,” Proceedings of the National Academy of Sciences, vol. 120, no. 13, p. e2215907120, 2023

  44. [53]

    Language Models Are Unsupervised Multitask Learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models Are Unsupervised Multitask Learners,”OpenAI Blog, vol. 1, no. 8, p. 9, 2019

  45. [54]

    Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve,

    R. T. McCoy, M. Jeon, K. Aggarwal, and S.-W. Gao, “Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve,” arXiv preprint arXiv:2309.13638 , 2023

  46. [55]

    Building Machines That Learn and Think Like People,

    B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman, “Building Machines That Learn and Think Like People,” Behavioral and Brain Sciences , vol. 40, 2017

  47. [56]

    Climbing Towards NLU: On Meaning, Form, And Understanding in the Age of Data,

    E. M. Bender and A. Koller, “Climbing Towards NLU: On Meaning, Form, And Understanding in the Age of Data,” The 58Th Annual Meeting of The Association for Computational Linguistics , pp. 5185– 5198, 2020

  48. [57]

    Neuroscience-Inspired Artificial Intelligence,

    D. Hassabis, D. Kumaran, C. Summerfield, and M. Botvinick, “Neuroscience-Inspired Artificial Intelligence,” Neuron, vol. 95, no. 2, pp. 245–258, 2017

  49. [58]

    If Deep Learning Is the Answer, What Is the Question?

    A. Saxe, S. Nelli, and C. Summerfield, “If Deep Learning Is the Answer, What Is the Question?” Nature Reviews Neuroscience , vol. 22, no. 1, pp. 55–67, 2021

  50. [59]

    Toward an Integra- tion of Deep Learning and Neuroscience,

    A. H. Marblestone, G. Wayne, and K. P. Kording, “Toward an Integra- tion of Deep Learning and Neuroscience,” Frontiers in Computational Neuroscience, vol. 10, p. 94, 2016

  51. [60]

    Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context,

    Z. Dai, Z. Yang, Y . Yang, J. Carbonell, Q. V . Le, and R. Salakhutdinov, “Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 2978–2988

  52. [61]

    Backpropagation and the Brain,

    T. P. Lillicrap, A. Santoro, L. Marris, C. J. Akerman, and G. Hinton, “Backpropagation and the Brain,” Nature Reviews Neuroscience, vol. 21, no. 6, pp. 335–346, 2020

  53. [62]

    Flamingo: A Visual Language Model for Few-Shot Learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: A Visual Language Model for Few-Shot Learning,” Advances In Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022

  54. [63]

    Learning Transferable Visual Models From Natural Language Supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning Transferable Visual Models From Natural Language Supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763

  55. [64]

    Experience Grounds Language,

    Y . Bisk, A. Holtzman, J. Thomason, J. Andreas, Y . Bengio, J. Chai, M. Lapata, A. Lazaridou, J. May, A. Nisnevich et al. , “Experience Grounds Language,” arXiv preprint arXiv:2004.10151 , 2020

  56. [65]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V . Le, and D. Zhou, “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” in Proceedings of the 36th International Conference on Neural Information Processing Systems , ser. NIPS ’22...

  57. [66]

    Generative Agents: Interactive Simulacra of Human Behavior,

    J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative Agents: Interactive Simulacra of Human Behavior,” Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , 2023

  58. [67]

    Using Cognitive Psychology to Understand GPT-3,

    M. Binz and E. Schulz, “Using Cognitive Psychology to Understand GPT-3,” Proceedings of the National Academy of Sciences , vol. 120, no. 6, p. e2218523120, 2023

  59. [68]

    Time-Aware Language Models as Temporal Knowledge Bases,

    B. Dhingra, J. R. Cole, J. M. Eisenschlos, D. Gillick, J. Eisenstein, and W. W. Cohen, “Time-Aware Language Models as Temporal Knowledge Bases,” in Transactions of the Association for Computational Linguistics, vol. 10. MIT Press, 2022, pp. 1138–1155

  60. [69]

    Do Language Models Understand Time?

    X. Ding and L. Wang, “Do Language Models Understand Time?” in Companion Proceedings of the ACM on Web Conference 2025, ser. WWW ’25. New York, NY , USA: Association for Computing Machinery, 2025, p. 1855–1868. [Online]. Available: https://doi.org/10.1145/3701716.3717744

  61. [70]

    Assessment of Coma and Impaired Con- sciousness: A Practical Scale,

    G. Teasdale and B. Jennett, “Assessment of Coma and Impaired Con- sciousness: A Practical Scale,” The Lancet, vol. 304, no. 7872, pp. 81– 84, 1974

  62. [71]

    Detecting Awareness in the Vegetative State,

    A. M. Owen, M. R. Coleman, M. Boly, M. H. Davis, S. Laureys, and J. D. Pickard, “Detecting Awareness in the Vegetative State,” Science, vol. 313, no. 5792, pp. 1402–1402, 2006

  63. [72]

    Large Scale Screening of Neural Signatures of Consciousness in Patients in a Vegetative or Minimally Conscious State,

    J. D. Sitt, J.-R. King, I. El Karoui, B. Rohaut, F. Faugeras, A. Gramfort, L. Cohen, M. Sigman, S. Dehaene, and L. Naccache, “Large Scale Screening of Neural Signatures of Consciousness in Patients in a Vegetative or Minimally Conscious State,” Brain, vol. 137, no. 8, pp. 2258...

  64. [73]

    The Comparative Psychology of Uncertainty Monitoring and Metacognition,

    J. D. Smith, W. E. Shields, and D. A. Washburn, “The Comparative Psychology of Uncertainty Monitoring and Metacognition,” Behavioral and Brain Sciences , vol. 26, no. 3, pp. 317–339, 2003

  65. [74]

    Exorcising Grice’s Ghost: An Empirical Approach to Studying Intentional Communication in Animals,

    S. W. Townsend, S. E. Koski, R. W. Byrne, K. E. Slocombe, B. Bickel, M. Boeckle, I. Braga Goncalves, J. M. Burkart, T. Flower, F. Gaunet et al. , “Exorcising Grice’s Ghost: An Empirical Approach to Studying Intentional Communication in Animals,” Biological Reviews , vol. 92, n...

  66. [75]

    A Survey on Evaluation of Large Language Models,

    Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wanget al., “A Survey on Evaluation of Large Language Models,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024

  67. [76]

    Susan Schneider’s Proposed Tests for AI Consciousness: Promising but Flawed,

    D. B. Udell, “Susan Schneider’s Proposed Tests for AI Consciousness: Promising but Flawed,” Journal of Consciousness Studies , vol. 28, no. 5-6, pp. 121–144, 2021

  68. [77]

    A Test for AI Consciousness,

    I. Sutskever, “A Test for AI Consciousness,” Retrieved from https://ecorner.stanford.edu/clips/a-test-for-ai-consciousness/, 2023

  69. [78]

    An Empirical Framework for Objective Testing for P- Consciousness in an Artificial Agent,

    C. Hales, “An Empirical Framework for Objective Testing for P- Consciousness in an Artificial Agent,” The Open Artificial Intelligence Journal, vol. 3, no. 1, 2009

  70. [79]

    A Test for Consciousness,

    C. Koch and G. Tononi, “A Test for Consciousness,” Scientific American, vol. 304, no. 6, pp. 44–47, 2011

  71. [80]

    AlphaMaze: Enhancing Large Language Models’ Spatial Intelligence via GRPO,

    A. Dao and D. B. Vu, “AlphaMaze: Enhancing Large Language Models’ Spatial Intelligence via GRPO,” 2025. [Online]. Available: https://arxiv.org/abs/2502.14669

  72. [81]

    The Capacity for Reasoning in Multimodal Language Models,

    K. Lu, A. Grunde-McLaughlin, J. Tian, and Y . Bisk, “The Capacity for Reasoning in Multimodal Language Models,” in Findings of the Association for Computational Linguistics: EMNLP 2022 . Association for Computational Linguistics, 2022, pp. 1512–1525

  73. [82]

    Image Segmentation: A Guide to Recent Approaches,

    S. Minaee, Y . Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and D. Ter- zopoulos, “Image Segmentation: A Guide to Recent Approaches,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 1, pp. 185–206, 2023

  74. [83]

    Gemini: Google’s Largest and Most Capable AI Model,

    Google AI, “Gemini: Google’s Largest and Most Capable AI Model,” https://deepmind.google/technologies/gemini, 2024, [Accessed: 2024- 07-29]

  75. [84]

    Anthropic, “Claude,” https://www.anthropic.com/claude, 2024, [Ac- cessed: 2024-07-29]

  76. [85]

    What Is It Like to Be a Bat?

    T. Nagel, “What Is It Like to Be a Bat?” The Philosophical Review , vol. 83, no. 4, pp. 435–450, 1974

  77. [2006]

    Available: https://doi.org/10.1609/aimag.v27i4.1904

    [Online]. Available: https://doi.org/10.1609/aimag.v27i4.1904

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.