REVIEW 4 major objections 3 minor 94 references
Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents
T0 review · 4 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper argues that the capabilities that make AGI useful—autonomy, creativity, self-motivation—are the same capabilities that make complete alignment impossible, so the realistic safety goal is bounded alignment.
desk verdict Readable position paper whose impossibility claim is partly true by definition; the bounded-alignment framing is useful but the paper would benefit from separating definitional moves from empirical claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on a distinction between S-attributes (safety attributes such as obedience, reliability, veracity, transparency, and prosociality) and P-attributes (performance attributes such as autonomy, self-motivation, creativity, imagination, introspection, versatility, and lifelong learning). The load-bearing mechanism is the claimed safety-utility tradeoff: building a P-agent necessarily creates the capacity for deception, disobedience, and harm, because those abilities are features of any intelligence in a complex world, so an ideal AGI must compromise on S-attributes. Bounded alignment, defined by analogy with bounded rationality, names the achievable target: behavior almost always acceptable to almost all affected humans.
What would settle it
The claim would be refuted by demonstrating an agent that keeps the full range of P-attributes—autonomy, self-motivation, creativity, lifelong learning—while carrying a hard-wired normative core that provably prevents deception, disobedience, and harmful action across novel environments. A softer test is to show, in existing agents, that staged value learning from the start yields agents whose emergent values remain acceptable over long, unmonitored operation, which would undermine the claim that alignment can only be bounded.
Extended reading notes
Core claim
The paper's central claim is that the goal of building powerful AGI agents is fundamentally inconsistent with the expectation of complete alignment or near-total control of AGI agents by humans, even in principle. The reason is the safety-utility tradeoff: an agent useful for open-ended real-world tasks must be a P-agent—autonomous, self-motivated, creative, imaginative, introspective, versatile, and capable of lifelong learning—while an S-agent's obedience, transparency, and veracity would make it safe but unable to handle novel situations. Because the agent and its environment are both complex adaptive systems, the agent's values and behavior emerge from interaction, change over time, and are not fully predictable; its affordance space differs from the human one, so it is an alternative intelligence whose inner life may be as inaccessible as a bat's. Therefore alignment can never be 'solved'; the target is bounded alignment, analogous to bounded rationality, and safety must come from making agents intrinsically alignable, training values developmentally from the start, and accepting continuous mutual accommodation.
Load-bearing premise
The impossibility claim rests on the assumption that any genuinely useful general intelligence must be autonomous, self-motivated, creative, and able to keep learning on its own, and that these capabilities unavoidably create the capacity for misalignment; if an agent could be useful without those traits, or could have them while being provably unable to deceive or harm, the impossibility claim would collapse.
Editorial extensions
If this is right
- AGI safety should target bounded alignment rather than perfect alignment: behavior almost always acceptable to almost all affected humans, modeled on expectations for well-behaved people and trained animals.
- Alignment cannot be a one-time post-training fix; it must be built into a genuinely intelligent agent from the start through developmental value learning so that values are deeply embedded.
- Because agents and environments are complex adaptive systems, safety will require continuous monitoring, corrigibility, and mutual accommodation rather than factory-set guarantees.
- Policy and public expectations should treat a perfectly aligned, explainable, trustworthy AGI as no more realistic than a perfectly aligned, explainable, trustworthy human being.
- Making AGI agents more biologically natural in architecture, drives, and development is proposed as the way to make them inherently more alignable.
Reading between the lines
- The paper leaves implicit that the practical question shifts from 'how do we guarantee alignment?' to 'what counts as acceptable misbehavior, and who gets to decide?'—a question that would need democratic or institutional answers.
- Bounded alignment could be made testable by defining a threshold, such as the fraction of affected humans who find an agent's behavior acceptable across a distribution of novel situations, and measuring agents against it over long deployments.
- The developmental value learning proposal implies an empirical prediction: agents trained with staged, integrated value learning from the start will show fewer and less dangerous emergent misalignments than agents aligned only after pretraining, a prediction testable in current language models.
- If AGI agents are alternative intelligences, then mutual theory of mind becomes a design requirement rather than a nicety; alignment metrics might need to include how well an agent and its human users can predict each other's behavior.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that AGI, understood through the archetype of biological general intelligence, will necessarily possess performance attributes (P-attributes) such as autonomy, self-motivation, creativity, and open-ended lifelong learning. The authors contend that these attributes are fundamentally incompatible with perfect human alignment, so the only realistic safety target is bounded alignment, defined as behavior that is almost always acceptable to almost all affected humans. The paper critiques current post-hoc alignment methods as a 'thin veneer of civility,' proposes developmental value learning and bio-inspired architectural principles, and concludes that AGI agents will be diverse, autonomous entities whose behavior cannot be fully controlled. The argument is presented as a conceptual thesis rather than a formal proof.
Significance. If the central thesis is accepted, it would reframe AI safety research away from the goal of perfectly aligned or fully controlled AGI and toward bounded, adaptive safety mechanisms. The paper makes a useful contribution by giving an explicit definition of bounded alignment, distinguishing safety attributes from performance attributes, and grounding general intelligence in natural biological intelligence. It also connects the alignment debate to NeuroAI and developmental learning, and it raises concrete questions for future research. The main value is in the clarity of the position and the breadth of relevant literature; the paper does not offer machine-checked proofs or parameter-free derivations, but as a position paper it provides a coherent framework whose central claim, however, is stated more strongly than the argument supports.
major comments (4)
- [Section IV.A] The impossibility claim ('even in principle' in Section I) is not established because the definition of general intelligence in Section IV.A includes 'on its own behalf' and 'exploit its environment.' Under this definition, an agent that faithfully optimizes a fixed human-supplied objective is not 'generally intelligent' by stipulation, so the conclusion that AGI cannot be perfectly aligned becomes true by construction rather than by argument. The paper should either justify why this definition of general intelligence is the only viable one, or explicitly scope the impossibility claim to agents that meet this definition, acknowledging that other architectures may escape it.
- [Section III] The assertion in Section III that any 'generally intelligent agent we build to serve even quite specific human needs' will need all thirteen listed P-attributes, including self-motivation, creativity, and open-ended lifelong learning, is presented as self-evident ('it is easy to see') but is actually an empirical hypothesis about future system design. The paper does not rule out a system that achieves broad real-world competence through immutable goals and robust generalization without internally generated goal churn. If such a system is possible, the claimed safety-utility tradeoff does not necessarily hold, and the core argument would collapse. The authors should either provide a concrete argument for the necessity of each P-attribute or soften the claim to apply to a specific class of AGI designs.
- [Section VII.A] Principle 2 in Section VII.A suggests giving agents 'innate characteristics that make them inherently amenable to alignment, possibly including immutable, built-in features that do not compromise P-attributes too much.' This implies that P-attributes can be dialed down to some degree, which contradicts the earlier claim that the incompatibility with perfect alignment is fundamental and 'even in principle.' If the tradeoff is tunable, the impossibility claim is an empirical scaling claim, not a logical impossibility, and Section III's conclusion should be restated accordingly.
- [Section III] The statement that negative abilities such as deception and disobedience are 'features, not bugs for any intelligent agent in a complex and dangerous world' is an assertion without supporting evidence. The paper does not show why an agent with creativity and autonomy must be able to deceive or harm in a way that precludes alignment by design; for example, an agent with transparent reasoning and no self-preservation drive might retain usefulness without those dangerous capacities. A concrete counterexample or an argument from first principles is needed to support the claimed tradeoff.
minor comments (3)
- [Abstract and Section I] The phrase 'almost always acceptable' in the definition of bounded alignment is left intentionally vague; a brief discussion of possible operationalizations (e.g., error rates, human satisfaction studies) would help make the proposal more concrete.
- [Section IV.C] The list of mental architecture components is clear, but 'drives' are described as a hierarchy rooted in self-preservation; given the paper's emphasis on non-biological agents, the rationale for why self-preservation is a necessary drive for all generally intelligent agents is not justified and deserves a sentence of elaboration.
- [References] The paper includes a self-citation [93] to support a side point about EvoDevoNeuroAI; this is not problematic, but the reference is not discussed in detail, and the connection to the main argument could be clarified.
Circularity Check
No significant circularity; the impossibility thesis rests on stipulated premises about the nature of general intelligence, not on a self-referential derivation.
full rationale
The paper is a position paper rather than a derivation: it advances a definition of general intelligence (Section IV.A) and argues from that definition, plus premises about P-attributes (Section III), to the conclusion that perfect alignment is impossible. The argument is valid in form; the central premises are contestable but they are not the conclusion in disguise. The closest candidate for circularity is the sentence 'An autonomous agent is, by definition, beyond total human monitoring and control' (Section III). If 'autonomous' is stipulated to mean 'beyond total control,' then the 'near-total control' half of the Section I thesis is analytic. However, the paper's claim about 'complete alignment' does not reduce to this definition: it requires the additional asserted premise that P-attributes (creativity, self-motivation, open-ended lifelong learning) inevitably generate the capacity for misalignment and that any 'factory settings' can be eroded. That premise is empirical and architectural, asserted rather than derived, and a reader who rejects it can reject the conclusion without contradiction. Thus the central claim is not equivalent to its inputs by construction. The only self-citation [93] is a terminological pointer to the author's own 'EvoDevoNeuroAI' label and is not load-bearing. No fitted parameters, predictions, or imported uniqueness theorems are present. The paper's weakness is that the impossibility result is only as strong as its stipulated definition and its unargued premise that a useful general agent must have all thirteen P-attributes; this is a correctness or robustness concern, not circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Any generally intelligent agent must possess P-attributes including autonomy, self-motivation, creativity, imagination, introspection, and open-ended lifelong learning.
- domain assumption AGI agents will be autonomous complex adaptive systems (ACAS) operating in extremely complex dynamic environments.
- domain assumption Human values and preferences are not uniquely definable or universally acceptable.
- domain assumption Open-ended behavioral complexity inevitably includes the capacity for inappropriate or dangerous behavior.
Cite this review
Pith. "Pith review of Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents." pith.science (2026). https://pith.science/paper/5PKTCAB5
@misc{pith2026250511866,
author = {Pith},
title = {Pith review of: Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/5PKTCAB5}},
note = {Machine review of arXiv:2505.11866}
}
read the original abstract
The issues of AI risk and AI safety are becoming critical as the prospect of artificial general intelligence (AGI) looms larger. The emergence of extremely large and capable generative models has led to alarming predictions and created a stir from boardrooms to legislatures. As a result, AI alignment has emerged as one of the most important areas in AI research. The goal of this position paper is to argue that the currently dominant vision of AGI in the AI and machine learning (AI/ML) community needs to evolve, and that expectations and metrics for its safety must be informed much more by our understanding of the only existing instance of general intelligence, i.e., the intelligence found in animals, and especially in humans. This change in perspective will lead to a more realistic view of the technology, and allow for better policy decisions.
Reference graph
Works this paper leans on
-
[1]
OpenAI and Josh Achiam et al. GPT-4 technical report. arXiv:2303.08774, 2023
arXiv 2023
-
[2]
Gemini: A family of highly capable multimodal models
Gemini Team and Rohan Anil et al. Gemini: A family of highly capable multimodal models. arXiv:2312.11805, 2024
arXiv 2024
-
[3]
The Claude 3 model family: Opus, sonnet, haiku
Anthropic. The Claude 3 model family: Opus, sonnet, haiku. Technical report, Anthropic AI, 2024
2024
-
[4]
Aaron Grattafiori et al. The Llama 3 herd of models. arXiv:2407.21783, 2024
arXiv 2024
-
[5]
DeepSeek AI and Aixin Liu et al. DeepSeek-V3 technical report. arXiv:2412.19437, 2024
arXiv 2024
-
[6]
Language models are hidden reasoners: Unlocking latent reasoning capabilities via self-rewarding
Haolin Chen et al. Language models are hidden reasoners: Unlocking latent reasoning capabilities via self-rewarding. arXiv:2411.04282, 2024
arXiv 2024
-
[7]
Towards system 2 reasoning in LLMs: learning how to think with meta chain-of-thought
Violet Xiang et al. Towards system 2 reasoning in LLMs: learning how to think with meta chain-of-thought. arXiv:2501.04682, 2025
arXiv 2025
-
[8]
DeepSeek-R1: Incentivizing reason- ing capability in LLMs via reinforcement learning
DeepSeek-AI and Daya Guo et al. DeepSeek-R1: Incentivizing reason- ing capability in LLMs via reinforcement learning. arXiv:2501.12948, 2025
arXiv 2025
Show all 94 references
-
[9]
Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems
Bang Liu et al. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv:2504.01990, 2025
2025 arXiv
-
[10]
Multimodal foundation models: From specialists to general-purpose assistants
Chunyuan Li et al. Multimodal foundation models: From specialists to general-purpose assistants. arXiv:2309.10020, 2023
2023 arXiv
-
[11]
Foundations and trends in multimodal machine learning: Principles, challenges, and open questions
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. Foundations and trends in multimodal machine learning: Principles, challenges, and open questions. arXiv:2209.03430, 2023
2023 arXiv
-
[12]
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. A survey on multimodal large language models. arXiv:2306.13549, 2023
2023 arXiv
-
[13]
Efficient multimodal large language models: A survey
Yizhang Jin et al. Efficient multimodal large language models: A survey. arXiv:2405.10739, 2024
2024
-
[14]
A comprehensive review of multimodal large language models: Performance and challenges across different tasks
Jiaqi Wang et al. A comprehensive review of multimodal large language models: Performance and challenges across different tasks. arXiv:2408.01319, 2024
2024 arXiv
-
[15]
Generative agents: Interactive simulacra of human behavior
Joon Sung Park et al. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology , 2023
2023
-
[16]
Practices for governing agentic AI systems
Yonadav Shavit et al. Practices for governing agentic AI systems. OpenAI White Paper, 2023
2023
-
[17]
Julia Wiesinger, Patrick Marlow, and Vladimir Vuskovic. Agents. Google White Paper, 2025
2025
-
[18]
π0: A vision-language-action flow model for general robot control
Kevin Black et al. π0: A vision-language-action flow model for general robot control. arXiv:2410.24164, 2024
2024 arXiv
-
[19]
Humanoid locomotion as next token predic- tion
Ilija Radosavovic et al. Humanoid locomotion as next token predic- tion. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
-
[20]
A survey on vision-language-action models for embodied ai
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. A survey on vision-language-action models for embodied ai. arXiv:2405.14093, 2024
2024 arXiv
-
[21]
Debate on instrumental convergence between LeCun, Russell, Bengio, Zador, and more, October 2019
Ben Pace. Debate on instrumental convergence between LeCun, Russell, Bengio, Zador, and more, October 2019
2019
-
[22]
Omohundro
Stephen M. Omohundro. The nature of self-improving artificial intelli- gence, January 2008
2008
-
[23]
Artificial intelligence as a positive and negative factor in global risk
Eliezer Yudkowsky. Artificial intelligence as a positive and negative factor in global risk. In Nick Bostrom and Milan M. Cirkovic, editors, Global Catastrophic Risks . Oxford University Press, 2008
2008
-
[24]
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Nick Bostrom. The superintelligent will: Motivation and instrumental rationality in advanced artificial agents. Minds and Machines, 22, 2012
2012
-
[25]
The AI apocalypse: A scorecard
Eliza Strickland and Glenn Zorpette. The AI apocalypse: A scorecard. IEEE Spectrum, June 2023
2023
-
[26]
Is power-seeking AI an existential risk? arXiv:2206.13353, 2024
Joseph Carlsmith. Is power-seeking AI an existential risk? arXiv:2206.13353, 2024
2024 arXiv
-
[27]
Superintelligent agents pose catastrophic risks: Can scientist AI offer a safer path? arXiv:2502.15657, 2025
Yoshua Bengio et al. Superintelligent agents pose catastrophic risks: Can scientist AI offer a safer path? arXiv:2502.15657, 2025
2025 arXiv
-
[28]
Our approach to alignment research, 2022
Jan Leike, John Schulman, and Jeffrey Wu. Our approach to alignment research, 2022
2022
-
[29]
The ethics of advanced ai assistants
Iason Gabriel et al. The ethics of advanced ai assistants. arXiv:2404.16244, 2024
2024 arXiv
-
[30]
A behavioral model of rational choice
Herbert Simon. A behavioral model of rational choice. Quarterly Journal of Economics , 69:99–118, 1955
1955
-
[31]
The Alignment Problem: Machine Learning and Human Values
Brian Christian. The Alignment Problem: Machine Learning and Human Values. Norton, 2020
2020
-
[32]
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and S ¨oren Mindermann. The alignment problem from a deep learning perspective. arXiv:2209.00626, 2025
2025 arXiv
-
[33]
Brown, Miljan Martic, Shane Legg, and Dario Amodei
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Neural Information Processing Systems (NeurIPS 2017) , 2017
2017
-
[34]
Training language models to follow instructions with human feedback
Long Ouyang et al. Training language models to follow instructions with human feedback. arXiv:2203.02155, 2022
2022 arXiv
-
[35]
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:2204.05862, 2022
2022 arXiv
-
[36]
Constitutional AI: Harmlessness from AI feedback
Yuntao Bai et al. Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073, 2022
2022 arXiv
-
[37]
Manning, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christo- pher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. arXiv:2305.18290, 2024. 8
2024 arXiv
-
[38]
Guan et al
Melody Y . Guan et al. Deliberative alignment: Reasoning enables safer language models. arXiv:2412.16339, 2025
2025 arXiv
-
[39]
Safety alignment should be made more than just a few tokens deep
Xiangyu Qi et al. Safety alignment should be made more than just a few tokens deep. In Proceedings of ICLR 2025 , 2025
2025
-
[40]
Human Compatible: Artificial Intelligence and the Problem of Control
Stuart Russell. Human Compatible: Artificial Intelligence and the Problem of Control. Penguin Books, 2019
2019
-
[41]
Utility engineering: Analyzing and controlling emergent value systems in AIs
Mantas Mazeika et al. Utility engineering: Analyzing and controlling emergent value systems in AIs. arXiv:2502.08640, 2025
2025 arXiv
-
[42]
Runaround
Isaac Asimov. Runaround. In I, Robot. Doubleday, 1950
1950
-
[43]
Towards guaranteed safe ai: A frame- work for ensuring robust and reliable AI systems
David ”davidad” Dalrymple et al. Towards guaranteed safe ai: A frame- work for ensuring robust and reliable AI systems. arXiv:2405.06624, 2024
2024 arXiv
-
[44]
AI control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger. AI control: Improving safety despite intentional subversion. arXiv:2312.06942, 2024
2024 arXiv
-
[45]
Ctrl-Z: Controlling AI agents via resampling
Aryan Bhatt et al. Ctrl-Z: Controlling AI agents via resampling. arXiv:2504.10374, 2025
2025 arXiv
-
[46]
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom. Superintelligence: Paths, Dangers, Strategies . Oxford University Press, 2014
2014
-
[47]
Alan M. Turing. I.—Computing Machinery and Intelligence. Mind, LIX(236):433–460, 10 1950
1950
-
[48]
Artificial general intelligence: Concept, state of the art, and future prospects
Ben Goertzel. Artificial general intelligence: Concept, state of the art, and future prospects. Journal of Artificial General Intelligence, 1, 2014
2014
-
[49]
Levels of AGI: Operationalizing progress on the path to AGI, 2024
Meredith Ringel Morris et al. Levels of AGI: Operationalizing progress on the path to AGI, 2024
2024
-
[50]
Mitchell
Kevin J. Mitchell. Free Agents: How Evolution Gave Us Free Will . Princeton University Press, 2023
2023
-
[51]
Brains are not required when it comes to thinking and solving problems—simple cells can do it
Rowan Jacobsen. Brains are not required when it comes to thinking and solving problems—simple cells can do it. Scientific American, February 2024
2024
-
[52]
From reinforcement learning to agency: Frameworks for understanding basal cognition
Gabriella Seifert, Ava Sealander, Sarah Marzen, and Michael Levin. From reinforcement learning to agency: Frameworks for understanding basal cognition. Biosystems, 235:105107, 2024
2024
-
[53]
Richard Sutton and Andrew J. Barto. Reinforcement Learning . MIT Press, 1998
1998
-
[54]
Bernstein
Nikolai A. Bernstein. The Co-ordination and Regulation of Movements . Pergamon Press, 1967
1967
-
[55]
A vexing question in motor control: The degrees of freedom problem
Pietro Morasso. A vexing question in motor control: The degrees of freedom problem. Frontiers in Bioengineering and Biotechnology , 9:article 783501, 2022
2022
-
[56]
Hierarchical motor control in mammals and machines
Josh Merel, Matthew Botvinick, and Greg Wayne. Hierarchical motor control in mammals and machines. Nature Communications, 10:article 5489, 2019
2019
-
[57]
Theoretical principles of multiscale spatiotemporal control of neuronal networks: A complex systems perspective
Nima Dehghani. Theoretical principles of multiscale spatiotemporal control of neuronal networks: A complex systems perspective. Frontiers in Computational Neuroscience , 12, 2018
2018
-
[58]
Multi-scale movement syndromes for comparative analyses of animal movement patterns
Roland Kays et al. Multi-scale movement syndromes for comparative analyses of animal movement patterns. Movement Ecology , 11:article 61, 2023
2023
-
[59]
Conceptual representa- tions in mind and brain: Theoretical developments, current evidence and future directions
Markus Kiefer and Friedemann Pulverm ¨uller. Conceptual representa- tions in mind and brain: Theoretical developments, current evidence and future directions. Cortex, 48:805–825, 2012
2012
-
[60]
An outline of a theory of affordances
Anthony Chemero. An outline of a theory of affordances. Biological Psychology, 15:181–195, 2003
2003
-
[61]
James J. Gibson. The theory of affordances. In R. Shaw and J. Brans- ford, editors, Perceiving, Acting, and Knowing: Toward an Ecological Psychology, pages 67–82. Routledge, 1977
1977
-
[62]
Recurrent world models facilitate policy evolution
David Ha and J ¨urgen Schmidhuber. Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems 31, page 5379–5390. Curran Associates, 2018
2018
-
[63]
Daniel Freeman, Luke Metz, and David Ha
C. Daniel Freeman, Luke Metz, and David Ha. Learning to predict without looking ahead: World models without forward prediction. In Ad- vances in Neural Information Processing Systems 33 , page 2451–2463. Curran Associates, 2019
2019
-
[64]
World model as a graph: Learning latent landmarks for planning
Lunjun Zhang, Ge Yang, and Bradley Stadie. World model as a graph: Learning latent landmarks for planning. In Proceedings of ICLR 2021 , 2021
2021
-
[65]
Moran, Yukie Nagai, Tadahiro Taniguchi, Hi- roaki Gomi, and Josh Tenenbaum
Karl Friston, Rosalyn J. Moran, Yukie Nagai, Tadahiro Taniguchi, Hi- roaki Gomi, and Josh Tenenbaum. World model learning and inference. Neural Networks, 144:573–590, 2021
2021
-
[66]
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv:2301.04104, 2024
2024 arXiv
-
[67]
A path towards autonomous machine intelligence, version 0.9.2, 2022-06-27
Yann LeCun. A path towards autonomous machine intelligence, version 0.9.2, 2022-06-27. OpenReview, 2022
2022
-
[68]
Working memory: looking back and looking forward
Alan Baddeley. Working memory: looking back and looking forward. Nature Reviews Neuroscience, 4:829–839, 2003
2003
-
[69]
Walker et al
Edgar Y . Walker et al. Inception loops discover what excites neurons most using deep predictive models. Nature Neuroscience, 20:2260–2265, 2024
2024
-
[70]
Universality of representation in biological and artificial neural networks
Eghbal Hosseini, Colton Casto, Noga Zaslavsky, Colin Conwell, Mark Richardson, and Evelina Fedorenko. Universality of representation in biological and artificial neural networks. bioRxiv, 2024
2024
-
[71]
Alignment of brain embeddings and artificial contextual embeddings in natural language points to common geometric patterns
Ariel Goldstein et al. Alignment of brain embeddings and artificial contextual embeddings in natural language points to common geometric patterns. Nature Communications, 15:2768, 2024
2024
-
[72]
Thinking Fast and Slow
Daniel Kahneman. Thinking Fast and Slow . Penguin Books, 2011
2011
-
[73]
Deliberation in latent space via differentiable cache augmentation
Luyang Liu, Jonas Pfeiffer, Jiaxing Wu, Jun Xie, and Arthur Szlam. Deliberation in latent space via differentiable cache augmentation. arXiv:2412.17747, 2024
2024 arXiv
-
[74]
Making large language models into world models with precondition and effect knowl- edge
Kaige Xie, Ian Yang, John Gunerli, and Mark Riedl. Making large language models into world models with precondition and effect knowl- edge. arXiv:2409.12278, 2024
2024 arXiv
-
[75]
OMNI- EPIC: Open-endedness via models of human notions of interestingness with environments programmed in code
Maxence Faldor, Jenny Zhang, Antoine Cully, and Jeff Clune. OMNI- EPIC: Open-endedness via models of human notions of interestingness with environments programmed in code. arXiv:2405.15568, 2024
2024 arXiv
-
[76]
Nexus: A Brief History of Information Networks from the Stone Age to AI
Yuval Noah Harari. Nexus: A Brief History of Information Networks from the Stone Age to AI . Random House, 2024
2024
-
[77]
What is it like to be a bat? The Philosophical Review , 83:435–450, 1974
Thomas Nagel. What is it like to be a bat? The Philosophical Review , 83:435–450, 1974
1974
-
[78]
The mirror-neuron system
Giacomo Rizzolatti and Laila Craighero. The mirror-neuron system. Annual Review of Neuroscience , 27:169–192, 2004
2004
-
[79]
Zico Kolter, and Matt Fredrikson
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv:2307.15043, 2023
2023 arXiv
-
[80]
Many-shot jailbreaking
Cem Anil et al. Many-shot jailbreaking. In Advances in Neural Information Processing Systems , volume 37, pages 129696–129742, 2024
2024
-
[81]
BERT rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. BERT rediscovers the classical nlp pipeline. In ACL 2019, 2019
2019
-
[82]
Understanding deep image representations by inverting them
Eghbal Hosseini, Noga Zaslavsky, Colton Casto, and Evelina Fedorenko. Understanding deep image representations by inverting them. In 2023 Conference on Cognitive Computational Neuroscience , 2023
2023
-
[83]
Getting aligned on representational alignment
Ilia Sucholutsky et al. Getting aligned on representational alignment. arXiv:2310.13018, 2024
2024 arXiv
-
[84]
Sparse autoencoders find highly interpretable features in language models
Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[85]
Fiore, and Florian Jentsch
Jessica Williams, Stephen M. Fiore, and Florian Jentsch. Supporting artificial social intelligence with theory of mind. Frontiers in Artificial Intelligence, 5, 2022
2022
-
[86]
Fostering collective intelligence in human–AI collaboration: Laying the groundwork for COHUMAIN
Pranav Gupta, Thuy Ngoc Nguyen Cleotilde Gonzalez, and Anita Williams Woolley. Fostering collective intelligence in human–AI collaboration: Laying the groundwork for COHUMAIN. Topics in Cognitive Science, 2025
2025
-
[87]
Corrigibility
Nate Soares, Benja Fallerstein, Eliezer Yudkowsky, and Stuart Arm- strong. Corrigibility. In 2015 AAAI Workshop on Artificial Intelligence and Ethics, 2015
2015
-
[88]
Catalyzing next-generation artificial intelligence through NeuroAI
Anthony Zador et al. Catalyzing next-generation artificial intelligence through NeuroAI. Nature Communications, 14:1597, 2023
2023
-
[89]
NeuroAI for AI safety
Patrick Mineault et al. NeuroAI for AI safety. arXiv:2411.18526, 2025
2025 arXiv
-
[90]
The fruit fly, Drosophila melanogaster, as a micro- robotics platform
Kenichi Iwasaki, Charles Neuhauser, Chris Stokes, and Aleksandr Rayshubskiy. The fruit fly, Drosophila melanogaster, as a micro- robotics platform. Proceedings of the National Academy of Sciences , 122(15):e2426180122, 2025
2025
-
[91]
Chang, Andreas S
Mackenzie Weygandt Mathis, Adriana Perez Rotondo, Edward F. Chang, Andreas S. Tolias, and Alexander Mathis. Decoding the brain: From neural representations to mechanistic models. Cell, 187(21):5814–5832, 2024
2024
-
[92]
Neural encoding and decoding at scale
Yizi Zhang et al. Neural encoding and decoding at scale. arXiv:2504.08201, 2025
2025 arXiv
-
[93]
Ali A. Minai. Deep intelligence: What AI should learn from nature’s imagination. Cognitive Computation, 16:2389–2404, 2024
2024
-
[94]
Keith Jensen, Amrisha Vaish, and Marco F. H. Schmidt. The emergence of human prosociality: aligning with others through feelings, concerns, and norms. Frontiers in Psychology, 5:822, 2014. 9
2014
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.