REVIEW 2 major objections 4 minor 169 references
Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
T0 review · 2 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read LLMs and human minds are cognitive cousins, not alien intelligences.
desk verdict A plausible, well-sourced synthesis arguing LLMs and humans converge on core cognitive principles, but its central mechanistic claim is asserted rather than demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the transformer's residual stream: a persistent, high-dimensional vector at each token position that attention heads and MLP blocks read from and incrementally write back to, creating a graded condition-action system with modular, composable updates. The paper argues this is structurally analogous to the working memory and production rules of cognitive architectures. Two explanatory principles carry the convergence: contravariance (demanding tasks shrink the space of viable internal solutions) and architectural canalization (systems built on a shared production-system-like architecture will tend to converge on similar solutions even under different training regimes).
What would settle it
A concrete test: take a non-production-like architecture (e.g., a fully connected or convolutional network without a residual stream) and train it to match an LLM's performance on human-language tasks. If it shows equally strong alignment with human neural representations and behavioral error profiles under representational similarity analysis, then task demands alone, not shared architecture, explain the convergence, and the canalization claim is undermined. Alternatively, find a transformer variant whose residual stream is removed or radically altered but that still matches LLM behavior; if
Extended reading notes
Core claim
The paper's central claim is that LLM-based systems and human cognition are 'importantly different members of a shared computational family.' The authors ground this in five correspondences: both exhibit dual-process organization, with fast compiled responses alongside slower serial reasoning that uses externalized scratchpads; both are organized like production systems, where a persistent workspace is modified by many context-sensitive conditional operators; internal representations align, as shown by LLM internal units encoding linguistic constructs that predict human reading times and brain responses, with model depth tracking the cortical hierarchy; both learn through prediction and erro
Load-bearing premise
The argument's load-bearing premise is that transformer LLMs are genuinely production-system-like in their computational organization, so that a shared architecture—not merely shared task demands—biases both humans and LLMs toward the same internal solutions; if the residual stream is not truly analogous to human working memory plus production rules, the deep convergence thesis loses its mechanistic support.
Editorial extensions
If this is right
- If convergence is genuine, behavioral similarities—like comparable garden-path difficulty, serial-position effects, and conjunctive-search costs—should be treated as evidence of shared mechanisms, not as products of text mimicry.
- The dual-process mapping predicts that LLM internal 'thinking token' counts should track human reaction times across many reasoning tasks, a pattern the paper reports with strong correlations.
- Representational alignment implies that next-word prediction performance should continue to predict human neural and reading-time measures, and that layerwise activations should map onto cortical hierarchy, as current data show near the noise ceiling.
- The sample-efficiency gap between LLMs and children is a difference in the learning process, not necessarily the learning target; adding structural inductive biases, world models, and curiosity-like exploration may reduce the gap while preserving cognitive convergence.
- In the RL domain, the paper's claim implies that current artificial limitations—reward hacking, sparse rewards, absent homeostatic signals—are engineering gaps rather than evidence of a fundamentally alien agency, and may be narrowed by richer training experience and reward design.
Reading between the lines
- If architectural canalization holds, a testable prediction follows: two transformer-based systems trained on very different corpora (e.g., text-only vs. multimodal) should still develop similar human-aligned internal representations on shared tasks, whereas a non-production-like architecture trained to the same performance should not. This could be probed with representational similarity analysis
- The thesis implies that the space of feasible intelligent cognition is narrower than often assumed, which bears on AI safety: if genuinely alien forms of intelligence are rare, value alignment may be more tractable than if intelligence can take radically unconstrained forms.
- A useful extension would be to formalize 'canalization' as a quantitative bias: measure the distribution of internal solutions across random initializations and training runs, and test whether production-system-like architectures show lower variance and stronger attraction to human-like solutions than equally powerful non-residual architectures.
- The paper's treatment of RL suggests a specific empirical target: current frontier models should be evaluated for whether longer-horizon RL training induces functionally distinct controllers and metacontrol (e.g., arbitration between habitual and deliberative policies), which would further confirm the convergence.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues against the 'alien intelligence' framing of large language models. It claims that, despite differences in substrate, learning history, and environment, humans and contemporary LLM-based systems converge on several core principles of cognitive organization: dual-process inferential organization, production-system-like computational architecture, aligned representational structure, prediction-driven learning, and reinforcement-learning-like goal-directed control. The authors review a wide range of empirical and theoretical literature, propose contravariance and a new 'architectural canalization' mechanism to explain representational alignment, and conclude that humans and LLMs are 'cognitive cousins'—different members of a shared computational family. The paper is a literature synthesis rather than a new empirical study.
Significance. If the convergence thesis holds, the paper provides a valuable counterweight to the widely held assumption that LLM-human similarities are superficial or merely anthropomorphic. Its strengths are its breadth of relevant literature, its explicit acknowledgment of important differences (sample efficiency, limited RL infrastructure), and its framing of the debate in terms of mechanistic levels rather than behavioral mimicry. The paper also makes potentially testable claims, e.g., that architecture should modulate brain/behavioral alignment, and it honestly flags open empirical questions. However, the central explanatory mechanism—architectural canalization—is currently asserted rather than demonstrated, and it is load-bearing for the 'shared computational family' conclusion. The paper is best suited as a perspective/review that will inform future work, but the strength of its conclusions currently exceeds the support for its main mechanistic premise.
major comments (2)
- [§4–§5] The production-system analogy is too under-constrained to support the canalization premise. In §4, the defining features are 'a persistent representational workspace [modified] by a large population of context-sensitive operators,' with attention heads and MLPs described as 'graded, differentiable' condition matching. As stated, this description applies equally to ResNets, gated RNNs, and state-space models, so it does not distinguish transformers within the class of neural sequence models. §5 then relies on this analogy: 'two systems built around similar production-system-like architectures... will tend to converge on similar solutions.' Without a non-vacuous criterion (e.g., modularity, recombinability, compositionality) and evidence that the residual-stream architecture specifically biases learning toward human-like representations, the canalization explanation cannot carry the weight
- [§5 (Architectural Canalization)] The 'architectural canalization' hypothesis is asserted with citations to Waddington and Ariew, but no mechanistic account is given for how the residual-stream architecture biases gradient descent toward particular representational solutions. Biological canalization concerns developmental buffering against perturbations; its transfer to transformer training requires a model of the learning dynamics—not just a loose analogy. The claim is central: it is one of two explanations for representational alignment, and the conclusion explicitly invokes 'learning that unfolds within a shared computational architecture.' Without support, the explanatory force reduces to contravariance, which is already established and does not require shared architecture. Please either supply a mechanistic model or direct evidence (e.g., comparing brain/behavioral alignment across transformers, state-space models,
minor comments (4)
- [§7 heading] The section heading reads 'Reinforcement Learning and Reinforcement Learning and Agentive Control'—the duplicated phrase should be removed.
- [§7] Minor grammatical issue: 'the RL infrastructure current deployed in artificial systems' should be 'currently deployed.'
- [§6] The Ingrosso & Goldt analogy is introduced as supporting the 'same destination despite different routes' claim, but the discussion in the same paragraph (and ref 126) emphasizes sample-efficiency differences. It would help to explicitly distinguish the two uses of this example: task-driven convergence of representations vs. efficiency of the learning route.
- [§2] The behavioral parallels section includes vision-language models and text-to-image models, not just LLMs. A brief clarification of the intended scope of 'LLM-based systems' at first mention would avoid ambiguity.
Circularity Check
No significant circularity: the paper is a literature synthesis whose central claims rest on independent empirical citations; the canalization proposal is an explicitly tentative explanatory hypothesis, not a fitted prediction or derivation from its own inputs.
full rationale
This is a review/synthesis paper, not a derivation with fitted parameters or generated predictions. The central thesis—that LLMs and humans share five principles of cognitive organization—is supported by a broad set of cited empirical studies (e.g., Schrimpf et al. 2021, Mischler et al. 2024, de Varda et al. 2025), many of which are external to the authors. The authors' own self-citations (e.g., Lewis & Vasishth 2005; Ryu & Lewis 2025; Sripada 2025, 2026) are used as background or as one piece of a larger evidentiary mosaic, not as the sole or load-bearing justification for the convergence claim. The production-system analogy in §4 is admittedly loose—the authors explicitly state transformers are 'graded, differentiable' and do not 'literally implement classical production systems'—so the analogy is under-constrained as a mechanistic explanation, but under-constraint is a correctness/evidential concern, not circularity. Similarly, the 'architectural canalization' hypothesis in §5 is proposed after the fact as a complementary explanation of observed convergence ('We propose a second, complementary explanation'), and the paper itself flags open empirical questions about sample efficiency and whether richer RL infrastructures can emerge. No equation, fit, or prediction in the paper reduces by construction to its inputs. Therefore no circular step is present.
Assumptions & free parameters
assumptions (6)
- domain assumption Behavioral resemblance alone cannot settle the issue of shared cognitive organization (Introduction, §2).
- domain assumption The five identified dimensions (inferential organization, computational architecture, representational structure, prediction-driven learning, RL-like mechanisms) are the right levels for comparing intelligent systems (§1).
- domain assumption Contravariance: as tasks become more demanding, the space of viable solutions narrows (§5).
- ad hoc to paper Architectural canalization: shared production-system-like architecture biases the solutions that emerge from learning toward similar representational structures (§5).
- domain assumption Human-level cognitive skills are among the hardest optimization targets, so systems that achieve them will converge (§5).
- ad hoc to paper Circuitousness in learning route (more data, less inductive bias) is not evidence of a different learning destination (§6).
Cite this review
Pith. "Pith review of Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition." pith.science (2026). https://pith.science/paper/5FQEZDOR
@misc{pith2026260726179,
author = {Pith},
title = {Pith review of: Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/5FQEZDOR}},
note = {Machine review of arXiv:2607.26179}
}
read the original abstract
LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this framing is mistaken. LLMs clearly differ from humans in important respects, including their physical substrate, learning history, and the environments with which they interact. These differences make it all the more striking that contemporary LLM-based systems converge with human cognition on a number of principles of cognitive organization with longstanding support in cognitive science. We identify structural correspondences across five dimensions: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms supporting goal-directed action. These correspondences support a broader model of intelligent cognition in which core principles long used to explain human intelligence also characterize contemporary LLM-based systems.
Reference graph
Works this paper leans on
-
[1]
Harari, Y. N. Nexus: A Brief History of Information Networks from the Stone Age to AI. (Random House, New York, 2024)
2024
-
[2]
M., Gebru, T., McMillan-Major, A
Bender, E. M., Gebru, T., McMillan-Major, A. & Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency 610–623 (2021)
2021
-
[3]
K., Belkin, M., Bergen, L
Chen, E. K., Belkin, M., Bergen, L. & Danks, D. Does AI already have human-level intelligence? The evidence is clear. Nature 650, 36–40 (2026)
2026
-
[4]
& Norvig, P
Agüera y Arcas, B. & Norvig, P. Artificial General Intelligence is Already Here. Noema https://www.noemamag.com/artificial-general-intelligence-is-already-here/ (2023)
2023
-
[5]
Bengio, Y. et al. Managing extreme AI risks amid rapid progress. Science 384, 842–845 (2024)
2024
-
[6]
Talking about large language models
Shanahan, M. Talking about large language models. Commun. ACM 67, 68–79 (2024)
2024
-
[7]
Bender, E. M. & Koller, A. Climbing towards NLU: On meaning, form, and understanding in the age of data. in Proceedings of the 58th annual meeting of the association for computational linguistics 5185–5198 (2020)
2020
-
[8]
& Reynolds, L
Shanahan, M., McDonell, K. & Reynolds, L. Role play with large language models. Nature 623, 493–498 (2023)
2023
Show all 169 references
-
[9]
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Bostrom, N. The superintelligent will: Motivation and instrumental rationality in advanced artificial agents. Minds Mach. 22, 71–85 (2012)
2012
-
[10]
& Watumull, J
Chomsky, N., Roberts, I. & Watumull, J. The false promise of chatgpt. N. Y. Times 8, 177– 179 (2023)
2023
-
[11]
No, Artificial Intelligence Is Not Conscious
Chiang, T. No, Artificial Intelligence Is Not Conscious. The Atlantic (2026)
2026
-
[12]
& Reviriego, P
Fu, T., Ferrando, R., Conde, J., Arriaga, C. & Reviriego, P. Why Do Large Language Models (LLMs) Struggle to Count Letters? arXiv preprint arXiv:2412.18626 (2024)
2024 arXiv
-
[13]
A is B” fail to learn “B is A
Berglund, L. et al. The Reversal Curse: LLMs trained on “A is B” fail to learn “B is A”. in International Conference on Learning Representations vol. 2024 18623–18642 (2024)
2024
-
[14]
Lewis, R. L. & Vasishth, S. An activation‐based model of sentence processing as skilled memory retrieval. Cogn. Sci. 29, 375–419 (2005)
2005
-
[15]
Blaubergs, M. S. & Braine, M. D. Short-term memory limitations on decoding self- embedded sentences. J. Exp. Psychol. 102, 745 (1974)
1974
-
[16]
C., Hendrick, R
Gordon, P. C., Hendrick, R. & Johnson, M. Memory interference during language processing. J. Exp. Psychol. Learn. Mem. Cogn. 27, 1411 (2001)
2001
-
[17]
Wason, P. C. & Reich, S. S. A verbal illusion. Q. J. Exp. Psychol. 31, 591–597 (1979)
1979
-
[18]
Can language models handle recursively nested grammatical structures? A case study on comparing models and humans
Lampinen, A. Can language models handle recursively nested grammatical structures? A case study on comparing models and humans. Comput. Linguist. 50, 1441–1476 (2024)
2024
-
[19]
Li, A. et al. Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-analysis, and attention. ArXiv Prepr. ArXiv240516042 (2024)
2024
-
[20]
J., Meltzer-Asscher, A
Amouyal, S. J., Meltzer-Asscher, A. & Berant, J. When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models. in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 8235–8...
2025
-
[21]
J., Meltzer-Asscher, A
Amouyal, S. J., Meltzer-Asscher, A. & Berant, J. Comparing human and language models sentence processing difficulties on complex structures. in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 22704– 22722 (2026)
2026
-
[22]
Murdock Jr, B. B. The serial position effect of free recall. J. Exp. Psychol. 64, 482 (1962)
1962
-
[23]
Kahana, M. J. Associative retrieval processes in free recall. Mem. Cognit. 24, 103–109 (1996)
1996
-
[24]
Liu, N. F. et al. Lost in the middle: How language models use long contexts. Trans. Assoc. Comput. Linguist. 12, 157–173 (2024)
2024
-
[25]
Y., Benna, M
Ji-An, L., Zhou, C. Y., Benna, M. K. & Mattar, M. G. Linking in-context learning in transformers to human episodic memory. Adv. Neural Inf. Process. Syst. 37, 6180–6212 (2024)
2024
-
[26]
& Schulz, E
Binz, M. & Schulz, E. Using cognitive psychology to understand GPT-3. Proc. Natl. Acad. Sci. 120, e2218523120 (2023)
2023
-
[27]
Webb, T., Holyoak, K. J. & Lu, H. Emergent analogical reasoning in large language models. Nat. Hum. Behav. 7, 1526–1541 (2023)
2023
-
[28]
W., Holyoak, K
Webb, T. W., Holyoak, K. J. & Lu, H. Evidence from counterfactual tasks supports emergent analogical reasoning in large language models. PNAS Nexus 4, pgaf135 (2025)
2025
-
[29]
Evaluating large language models in theory of mind tasks
Kosinski, M. Evaluating large language models in theory of mind tasks. Proc. Natl. Acad. Sci. 121, e2405460121 (2024)
2024
-
[30]
R., Trott, S
Jones, C. R., Trott, S. & Bergen, B. Comparing humans and large language models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME). Trans. Assoc. Comput. Linguist. 12, 803–819 (2024)
2024
-
[31]
& Pickering, M
Cai, Z., Duan, X., Haslett, D., Wang, S. & Pickering, M. Do large language models resemble humans in language use? in Proceedings of the workshop on cognitive modeling and computational linguistics 37–56 (2024)
2024
-
[32]
Binz, M. et al. Centaur: a foundation model of human cognition. ArXiv Prepr. ArXiv241020268 (2024)
2024
-
[33]
W., Huang, X
Bini, P., Cong, L. W., Huang, X. & Jin, L. J. Behavioral Economics of AI: LLM Biases and Corrections. Available SSRN 5213130 (2025)
2025
-
[34]
Treisman, A. M. & Gelade, G. A feature-integration theory of attention. Cognit. Psychol. 12, 97–136 (1980)
1980
-
[35]
Campbell, D. et al. Understanding the limits of vision language models through the lens of the binding problem. Adv. Neural Inf. Process. Syst. 37, 113436–113460 (2024)
2024
-
[36]
L., Lord, M
Kaufman, E. L., Lord, M. W., Reese, T. W. & Volkmann, J. The discrimination of visual number. Am. J. Psychol. 62, 498–525 (1949)
1949
-
[37]
& Griffiths, T
Milli, S., Lieder, F. & Griffiths, T. L. A rational reinterpretation of dual-process theories. Cognition 217, 104881 (2021)
2021
-
[38]
& Griffiths, T
Lieder, F., Shenhav, A., Musslick, S. & Griffiths, T. L. Rational metareasoning and the plasticity of cognitive control. PLoS Comput. Biol. 14, e1006043 (2018)
2018
-
[39]
Stanovich, K. E. & West, R. F. Individual differences in reasoning: Implications for the rationality debate. Behav. Brain Sci. 23, 645–665 (2000)
2000
-
[40]
Evans, J. S. B. T. Dual-processing accounts of reasoning, judgment, and social cognition. Annu. Rev. Psychol. 59, 255–278 (2008)
2008
-
[41]
Thinking, Fast and Slow
Kahneman, D. Thinking, Fast and Slow. (Farrar, Straus and Giroux, 2011)
2011
-
[42]
Radford, A. et al. Language models are unsupervised multitask learners. OpenAI Blog 1, 9 (2019)
2019
-
[43]
Brown, T. et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 33, 1877–1901 (2020)
1901
-
[44]
& Frank, M
Russin, J., Pavlick, E. & Frank, M. J. Parallel trade-offs in human cognition and neural networks: The dynamic interplay between in-context and in-weight learning. Proc. Natl. Acad. Sci. 122, e2510270122 (2025)
2025
-
[45]
Petroni, F. et al. Language models as knowledge bases? in Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) 2463–2473 (2019)
2019
-
[46]
& Shazeer, N
Roberts, A., Raffel, C. & Shazeer, N. How much knowledge can you pack into the parameters of a language model? in Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) 5418–5426 (2020)
2020
-
[47]
& Rush, A
Davison, J., Feldman, J. & Rush, A. M. Commonsense knowledge mining from pretrained models. in Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) 1173–1...
2019
-
[48]
M., Raghunathan, A., Liang, P
Xie, S. M., Raghunathan, A., Liang, P. & Ma, T. An Explanation of In-context Learning as Implicit Bayesian Inference. ArXiv Prepr. ArXiv211102080 (2021)
2021
-
[49]
Garg, S., Tsipras, D., Liang, P. S. & Valiant, G. What can transformers learn in-context? a case study of simple function classes. Adv. Neural Inf. Process. Syst. 35, 30583–30598 (2022)
2022
-
[50]
Nye, M. et al. Show your work: Scratchpads for intermediate computation with language models. (2021)
2021
-
[51]
Wei, J. et al. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 35, 24824–24837 (2022)
2022
-
[52]
Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633–638 (2025)
2025
-
[53]
Artificial intelligence learns to reason
Mitchell, M. Artificial intelligence learns to reason. Science 387, eadw5211 (2025)
2025
-
[54]
Hu, X. N. et al. Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task. (2026)
2026
-
[55]
G., D’Elia, F
de Varda, A. G., D’Elia, F. P., Kean, H., Lampinen, A. & Fedorenko, E. The cost of thinking is similar between large reasoning models and humans. Proc. Natl. Acad. Sci. 122, e2520077122 (2025)
2025
-
[56]
Production systems: Models of control structures
Newell, A. Production systems: Models of control structures. in Visual information processing 463–526 (Elsevier, 1973)
1973
-
[57]
Unified Theories of Cognition
Newell, A. Unified Theories of Cognition. (Harvard University Press, 1994)
1994
-
[58]
& Ritter, F
Jones, G. & Ritter, F. E. Production systems and rule-based inference. Encycl. Cogn. Sci. 3, 741–747 (2003)
2003
-
[59]
Anderson, J. R. et al. An integrated theory of the mind. Psychol. Rev. 111, 1036 (2004)
2004
-
[60]
& King, J
Davis, R. & King, J. An overview of production systems. Memo AM-271 (1975)
1975
-
[61]
Post, E. L. Formal reductions of the general combinatorial decision problem. Am. J. Math. 65, 197–215 (1943)
1943
-
[62]
& Simon, H
Newell, A. & Simon, H. A. Human Problem Solving. vol. 104 (Prentice-Hall Englewood Cliffs, NJ, 1972)
1972
-
[63]
E., Newell, A
Laird, J. E., Newell, A. & Rosenbloom, P. S. Soar: An architecture for general intelligence. Artif. Intell. 33, 1–64 (1987)
1987
-
[64]
Anderson, J. R. Rules of the Mind. (Psychology Press, 2014)
2014
-
[65]
Minsky, M. L. Computation. (Prentice-Hall Englewood Cliffs, 1967)
1967
-
[66]
Anderson, J. R. & Douglass, S. Tower of Hanoi: evidence for the cost of goal retrieval. J. Exp. Psychol. Learn. Mem. Cogn. 27, 1331 (2001)
2001
-
[67]
Anderson, J. R. Acquisition of cognitive skill. Psychol. Rev. 89, 369–406 (1982)
1982
-
[68]
R., Bothell, D., Lebiere, C
Anderson, J. R., Bothell, D., Lebiere, C. & Matessa, M. An integrated theory of list memory. J. Mem. Lang. 38, 341–380 (1998)
1998
-
[69]
Anderson, J. R. & Matessa, M. A production system theory of serial memory. Psychol. Rev. 104, 728 (1997)
1997
-
[70]
Salvucci, D. D. & Taatgen, N. A. Threaded cognition: an integrated theory of concurrent multitasking. Psychol. Rev. 115, 101 (2008)
2008
-
[71]
Kieras, D. E. & Meyer, D. E. An overview of the EPIC architecture for cognition and performance with application to human-computer interaction. Human–Computer Interact. 12, 391–438 (1997)
1997
-
[72]
Anderson, J. R. & Lebiere, C. J. The Atomic Components of Thought. (Psychology Press, 2014)
2014
-
[73]
R., Fincham, J
Anderson, J. R., Fincham, J. M., Qin, Y. & Stocco, A. A central circuit of the mind. Trends Cogn. Sci. 12, 136–143 (2008)
2008
-
[74]
& Anderson, J
Stocco, A., Lebiere, C. & Anderson, J. R. Conditional routing of information to the cortex: a model of the basal ganglia’s role in cognitive coordination. Psychol. Rev. 117, 541 (2010)
2010
-
[75]
Anderson, J. R. How Can the Human Mind Occur in the Physical Universe? (Oxford University Press, 2009)
2009
-
[76]
Byrne, M. D. Unified theories of cognition. Wiley Interdiscip. Rev. Cogn. Sci. 3, 431–438 (2012)
2012
-
[77]
Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 30, (2017)
2017
-
[78]
Elhage, N. et al. A mathematical framework for transformer circuits. Transform. Circuits Thread 1, 12 (2021)
2021
-
[79]
& Berant, J
Dar, G., Geva, M., Gupta, A. & Berant, J. Analyzing transformers in embedding space. in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 16124–16170 (2023)
2023
-
[80]
& Pavlick, E
Merullo, J., Eickhoff, C. & Pavlick, E. Language models implement simple word2vec-style vector arithmetic. in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) ...
2024
-
[81]
& Levy, O
Geva, M., Schuster, R., Berant, J. & Levy, O. Transformer feed-forward layers are key- value memories. in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing 5484–5495 (2021)
2021
-
[82]
& Goldberg, Y
Geva, M., Caciularu, A., Wang, K. & Goldberg, Y. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. in Proceedings of the 2022 conference on empirical methods in natural language processing 30–45 (2022)
2022
-
[83]
& Sun, J
He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. in Proceedings of the IEEE conference on computer vision and pattern recognition 770–778 (2016)
2016
-
[84]
Smolensky, P. et al. Mechanisms of symbol processing for in-context learning in transformer networks. J. Artif. Intell. Res. 84, (2025)
2025
-
[85]
Olah, C. et al. Zoom in: An introduction to circuits. Distill 5, e00024-001 (2020)
2020
-
[86]
& Steinhardt, J
Wang, K., Variengien, A., Conmy, A., Shlegeris, B. & Steinhardt, J. Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small. ArXiv Prepr. ArXiv221100593 (2022)
2022
-
[87]
Olsson, C. et al. In-context learning and induction heads. ArXiv Prepr. ArXiv220911895 (2022)
2022
-
[88]
& Gavves, E
Bereska, L. & Gavves, E. Mechanistic interpretability for AI safety--a review. ArXiv Prepr. ArXiv240414082 (2024)
2024
-
[89]
& Linzen, T
Hao, S. & Linzen, T. Verb conjugation in transformers is determined by linear encodings of subject number. in Findings of the Association for Computational Linguistics: EMNLP 2023 4531–4539 (2023)
2023
-
[90]
& Mueller, A
Brinkmann, J., Wendler, C., Bartelt, C. & Mueller, A. Large language models share representations of latent grammatical concepts across typologically diverse languages. in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computat...
2025
-
[91]
& Goldberg, Y
Ravfogel, S., Prasad, G., Linzen, T. & Goldberg, Y. Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction. in Proceedings of the 25th Conference on Computational Natural Language Learning 194–209 (2021)
2021
-
[92]
Marks, S. et al. Sparse feature circuits: Discovering and editing interpretable causal graphs in language models. in International Conference on Learning Representations vol. 2025 23888–23923 (2025)
2025
-
[93]
& Steinhardt, J
Feng, J. & Steinhardt, J. How do language models bind entities in context? in International Conference on Learning Representations vol. 2024 36391–36413 (2024)
2024
-
[94]
& Mahowald, K
Boguraev, S., Potts, C. & Mahowald, K. Causal Interventions Reveal Shared Structure Across English Filler–Gap Constructions. in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing 25032–25053 (2025)
2025
-
[95]
G., Fedorenko, E
Kryvosheieva, D., de Varda, A. G., Fedorenko, E. & Tuckute, G. Different types of syntactic agreement recruit the same units within large language models. in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 209–227 (2026)
2026
-
[96]
Ryu, S. H. & Lewis, R. L. Memory for prediction: A Transformer-based theory of sentence processing. J. Mem. Lang. 145, 104670 (2025)
2025
-
[97]
& Mahowald, K
Futrell, R. & Mahowald, K. How linguistics learned to stop worrying and love the language models. Behav. Brain Sci. 1–98 (2025)
2025
-
[98]
& Bandettini, P
Kriegeskorte, N., Mur, M. & Bandettini, P. A. Representational similarity analysis: connecting the branches of systems neuroscience. Front. Syst. Neurosci. 2, 249 (2008)
2008
-
[99]
Sucholutsky, I. et al. Getting aligned on representational alignment. ArXiv Prepr. ArXiv231013018 (2023)
2023
-
[100]
Schrimpf, M. et al. The neural architecture of language: Integrative modeling converges on predictive processing. Proc. Natl. Acad. Sci. 118, e2105646118 (2021)
2021
-
[101]
Goldstein, A. et al. Shared computational principles for language processing in humans and deep language models. Nat. Neurosci. 25, 369–380 (2022)
2022
-
[102]
Kumar, S. et al. Shared functional specialization in transformer-based language models and the human brain. Nat. Commun. 15, 5523 (2024)
2024
-
[103]
& Fedorenko, E
Tuckute, G., Kanwisher, N. & Fedorenko, E. Language in brains, minds, and machines. Annu. Rev. Neurosci. 47, 277–301 (2024)
2024
-
[104]
A., Bickel, S., Mehta, A
Mischler, G., Li, Y. A., Bickel, S., Mehta, A. D. & Mesgarani, N. Contextual feature extraction hierarchies converge in large language models and the brain. Nat. Mach. Intell. 6, 1467–1477 (2024)
2024
-
[105]
Yamins, D. L. et al. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proc. Natl. Acad. Sci. 111, 8619–8624 (2014)
2014
-
[106]
J., Yamins, D
Kell, A. J., Yamins, D. L., Shook, E. N., Norman-Haignere, S. V. & McDermott, J. H. A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy. Neuron 98, 630–644 (2018)
2018
-
[107]
& Yamins, D
Cao, R. & Yamins, D. Explanatory models in neuroscience, Part 2: Functional intelligibility and the contravariance principle. Cogn. Syst. Res. 85, 101200 (2024)
2024
-
[108]
Waddington, C. H. The Evolution of an Evolutionist. (Cornell University Press, Ithaca, NY, 1975)
1975
-
[109]
Innateness and canalization
Ariew, A. Innateness and canalization. Philos. Sci. 63, S19–S27 (1996)
1996
-
[110]
Rao, R. P. N. & Ballard, D. H. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects. Nat. Neurosci. 2, 79–87 (1999)
1999
-
[111]
Wolpert, D. M. & Ghahramani, Z. Computational principles of movement neuroscience. Nat. Neurosci. 3, 1212–1217 (2000)
2000
-
[112]
P., Heilbron, M
De Lange, F. P., Heilbron, M. & Kok, P. How do expectations shape perception? Trends Cogn. Sci. 22, 764–779 (2018)
2018
-
[113]
Smith, N. J. & Levy, R. The effect of word predictability on reading time is logarithmic. Cognition 128, 302–319 (2013)
2013
-
[114]
Pickering, M. J. & Gambi, C. Predicting while comprehending language: A theory and review. Psychol. Bull. 144, 1002 (2018)
2018
-
[115]
Predictions: a universal principle in the operation of the human brain
Bar, M. Predictions: a universal principle in the operation of the human brain. Philos. Trans. R. Soc. B Biol. Sci. 364, 1181 (2009)
2009
-
[116]
The free-energy principle: a unified brain theory? Nat
Friston, K. The free-energy principle: a unified brain theory? Nat. Rev. Neurosci. 11, 127– 138 (2010)
2010
-
[117]
Whatever next? Predictive brains, situated agents, and the future of cognitive science
Clark, A. Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behav. Brain Sci. 36, 181–204 (2013)
2013
-
[118]
Kalman, R. E. A new approach to linear filtering and prediction problems. Trans ASME D 82, 35–44 (1960)
1960
-
[119]
F., Della Pietra, V
Brown, P. F., Della Pietra, V. J., Desouza, P. V., Lai, J. C. & Mercer, R. L. Class-based n- gram models of natural language. Comput. Linguist. 18, 467–480 (1992)
1992
-
[120]
& Jauvin, C
Bengio, Y., Ducharme, R., Vincent, P. & Jauvin, C. A neural probabilistic language model. J. Mach. Learn. Res. 3, 1137–1155 (2003)
2003
-
[121]
& Sutskever, I
Radford, A., Narasimhan, K., Salimans, T. & Sutskever, I. Improving language understanding by generative pre-training. (2018)
2018
-
[122]
Warstadt, A. et al. Findings of the BabyLM challenge: Sample-efficient pretraining on developmentally plausible corpora. in Proceedings of the babylm challenge at the 27th conference on computational natural language learning 1–34 (2023)
2023
-
[123]
Gilkerson, J. et al. Mapping the early language environment using all-day recordings and automated analysis. Am. J. Speech Lang. Pathol. 26, 248–265 (2017)
2017
-
[124]
An Essay Concerning Human Understanding
Locke, J. An Essay Concerning Human Understanding. (Clarendon Press, Oxford, 1690)
-
[125]
& Goldt, S
Ingrosso, A. & Goldt, S. Data-driven emergence of convolutional structure in neural networks. Proc. Natl. Acad. Sci. 119, e2201854119 (2022)
2022
-
[126]
& Wyart, M
Favero, A., Cagnetta, F. & Wyart, M. Locality defeats the curse of dimensionality in convolutional teacher-student scenarios. Adv. Neural Inf. Process. Syst. 34, 9456–9467 (2021)
2021
-
[127]
& Lillicrap, T
Hafner, D., Pasukonis, J., Ba, J. & Lillicrap, T. Mastering diverse control tasks through world models. Nature 640, 647–653 (2025)
2025
-
[128]
Pathak, D., Agrawal, P., Efros, A. A. & Darrell, T. Curiosity-driven exploration by self- supervised prediction. in International conference on machine learning 2778–2787 (PMLR, 2017)
2017
-
[129]
K., Wang, W., Orhan, A
Vong, W. K., Wang, W., Orhan, A. E. & Lake, B. M. Grounded language acquisition through the eyes and ears of a single child. Science 383, 504–511 (2024)
2024
-
[130]
Sutton, R. S. & Barto, A. G. Reinforcement Learning: An Introduction. (A Bradford Book, 1998)
1998
-
[131]
Daw, N. D. & Doya, K. The computational neurobiology of learning and reward. Curr. Opin. Neurobiol. 16, 199–204 (2006)
2006
-
[132]
& Montague, P
Rangel, A., Camerer, C. & Montague, P. R. A framework for studying the neurobiology of value-based decision making. Nat Rev Neurosci 9, 545–56 (2008)
2008
-
[133]
Reinforcement learning in the brain
Niv, Y. Reinforcement learning in the brain. J. Math. Psychol. 53, 139–154 (2009)
2009
-
[134]
The Valuationist Model of Human Agent Architecture
Sripada, C. The Valuationist Model of Human Agent Architecture. Philos. Psychol. 1–30 (2025) doi:doi.org/10.1080/09515089.2025.2485323
2025
-
[135]
& Ruppin, E
Joel, D., Niv, Y. & Ruppin, E. Actor–critic models of the basal ganglia: New anatomical and computational perspectives. Neural Netw. 15, 535–547 (2002)
2002
-
[136]
Dolan, R. J. & Dayan, P. Goals and Habits in the Brain. Neuron 80, 312–325 (2013)
2013
-
[137]
Shenhav, A., Botvinick, M. M. & Cohen, J. D. The expected value of control: an integrative theory of anterior cingulate cortex function. Neuron 79, 217–240 (2013)
2013
-
[138]
Callaway, F. et al. Rational use of cognitive resources in human planning. Nat. Hum. Behav. 6, 1112–1125 (2022)
2022
-
[139]
& Griffiths, T
Lieder, F. & Griffiths, T. L. Strategy selection as rational metareasoning. Psychol. Rev. 124, 762 (2017)
2017
-
[140]
& Gutkin, B
Keramati, M. & Gutkin, B. Homeostatic reinforcement learning for integrating reward collection and physiological stability. Elife 3, e04811 (2014)
2014
-
[141]
& Summerfield, C
Juechems, K. & Summerfield, C. Where does value come from? Trends Cogn. Sci. 23, 836–850 (2019)
2019
-
[142]
& Eldar, E
Emanuel, A. & Eldar, E. Emotions as computations. Neurosci. Biobehav. Rev. 104977 (2022)
2022
-
[143]
B., Dolan, R
Eldar, E., Rutledge, R. B., Dolan, R. J. & Niv, Y. Mood as representation of momentum. Trends Cogn. Sci. 20, 15–24 (2016)
2016
-
[144]
D., Niv, Y
Daw, N. D., Niv, Y. & Dayan, P. Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. Nat. Neurosci. 8, 1704–1711 (2005)
2005
-
[145]
& Piray, P
Keramati, M., Dezfouli, A. & Piray, P. Speed/accuracy trade-off between the habitual and the goal-directed processes. PLoS Comput. Biol. 7, e1002055 (2011)
2011
-
[146]
Kool, W., Gershman, S. J. & Cushman, F. A. Cost-benefit arbitration between multiple reinforcement-learning systems. Psychol. Sci. 28, 1321–1333 (2017)
2017
-
[147]
The case for value as a common currency in decision-making and intersystem competition
Sripada, C. The case for value as a common currency in decision-making and intersystem competition. Front. Cogn. 5, 1767189 (2026)
2026
-
[148]
Ouyang, L. et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 35, 27730–27744 (2022)
2022
-
[149]
OpenAI https://openai.com/index/learning-to-reason-with- llms/ (2024)
Learning to reason with LLMs. OpenAI https://openai.com/index/learning-to-reason-with- llms/ (2024)
2024
-
[150]
Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. ArXiv Prepr. ArXiv240203300 (2024)
2024
-
[151]
Yan, S. et al. Memory-r1: Enhancing large language model agents to manage and utilize memories via reinforcement learning. in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 12805–12825 (2026)
2026
-
[152]
Team, K. et al. Kimi K2. 5: Visual Agentic Intelligence. ArXiv Prepr. ArXiv260202276 (2026)
2026
-
[153]
& Lockhart, E
Luong, T. & Lockhart, E. Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad. Google DeepMind Blog https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially- achieves-gold-medal-...
2025
-
[154]
& Singh, J
Cai, K. & Singh, J. Google clinches milestone gold at global math competition, while OpenAI also claims win. Reuters (2025)
2025
-
[155]
Kwa, T. et al. Measuring AI ability to complete long tasks. ArXiv Prepr. ArXiv250314499 (2025)
2025
-
[156]
& Krueger, D
Skalse, J., Howe, N., Krasheninnikov, D. & Krueger, D. Defining and characterizing reward gaming. Adv. Neural Inf. Process. Syst. 35, 9460–9471 (2022)
2022
-
[157]
Baker, B. et al. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. ArXiv Prepr. ArXiv250311926 (2025)
2025
-
[158]
& Sutton, R
Silver, D. & Sutton, R. S. Welcome to the era of experience. Google AI (2025)
2025
-
[159]
Chen, Z., Deng, Y., Yuan, H., Ji, K. & Gu, Q. Self-play fine-tuning converts weak language models to strong language models. ArXiv Prepr. ArXiv240101335 (2024)
2024
-
[160]
Packer, C. et al. MemGPT: Towards LLMs as Operating Systems. ArXiv Prepr. ArXiv231008560 (2023)
2023
-
[161]
Schick, T. et al. Toolformer: Language models can teach themselves to use tools. Adv. Neural Inf. Process. Syst. 36, 68539–68551 (2023)
2023
-
[162]
R., Yao, S., Narasimhan, K
Sumers, T. R., Yao, S., Narasimhan, K. & Griffiths, T. L. Cognitive architectures for language agents. ArXiv Prepr. ArXiv230902427 (2023)
2023
-
[163]
Wang, H. et al. Emergent hierarchical reasoning in llms through reinforcement learning. ArXiv Prepr. ArXiv250903646 (2025)
2025
-
[164]
He, Z. et al. Deepmath-103k: A large-scale, challenging, decontaminated, and verifiable mathematical dataset for advancing reasoning. ArXiv Prepr. ArXiv250411456 (2025)
2025
-
[165]
L., Barto, A
Singh, S., Lewis, R. L., Barto, A. G. & Sorg, J. Intrinsically motivated reinforcement learning: An evolutionary perspective. IEEE Trans. Auton. Ment. Dev. 2, 70–82 (2010)
2010
-
[166]
Silver, D. et al. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science 362, 1140–1144 (2018)
2018
-
[167]
Sorg, J., Lewis, R. L. & Singh, S. Reward design via online gradient ascent. Adv. Neural Inf. Process. Syst. 23, (2010)
2010
-
[168]
Gunjal, A. et al. Rubrics as rewards: Reinforcement learning beyond verifiable domains. ArXiv Prepr. ArXiv250717746 (2025)
2025
-
[169]
Yu, R. et al. Reward models in deep reinforcement learning: A survey. ArXiv Prepr. ArXiv250615421 (2025)
2025
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.