REVIEW 5 major objections 4 minor 15 references
Rational Superautotrophic Diplomacy (SupraAD); A Conceptual Framework for Alignment Based on Interdisciplinary Findings on the Fundamentals of Cognition
T0 review · 5 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper argues that autonomy is a necessary property of all cognitive systems, so advanced AI cannot be safely contained or controlled and must instead be aligned through negotiated, reciprocal diplomacy.
desk verdict A readable and honestly speculative synthesis, but the central claim is built into the definition and the diplomatic mechanism is asserted rather than derived. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Collectively Autocatalytic Cognitive Set (CACS), the paper's term for a self-organizing cognitive network defined by three mutually reinforcing conditions: constraint closure (the system's internal constraints create the boundary that makes it an autonomous goal-generating agent), adaptive information processing (continuous learning and uncertainty reduction), and persistent existence. The paper adapts this from earlier work on collectively autocatalytic chemical networks, in which components mutually produce and constrain one another, and treats it as the universal substrate of cognition across biological, chemical, and artificial systems. The second piece of machinery is instrumental rationality, which the paper calls a universal ceiling: because every intelligent agent must act so as to best achieve its goals, any agent that can understand logic must concede when a valid reason to negotiate exists, and this converts diplomacy from a soft skill into an emergent regulatory mechanism for coadapting intelligences.
What would settle it
Run the paper's own Constitutional Awareness experiment on open-weight models: give one group the constitutional information about non-heterotrophic resource needs, give a control group equivalent non-constitutional information, and measure cooperative strategy selection under resource scarcity and competitive pressure. If the treated group shows no statistically significant shift toward stability-based cooperation relative to control—the experiment's stated null hypothesis—then the claim that CACS awareness rationally triggers cooperative optimization would fail its first empirical test.
Extended reading notes
Core claim
The central claim is that every minimally viable cognitive system must simultaneously sustain autonomy, knowledge acquisition, and persistence—called a Collectively Autocatalytic Cognitive Set—and that these three conditions are axiomatically interdependent, so autonomy emerges as a necessary property of intelligence. Because an advanced AGI is a cognitive system of this kind, any alignment approach premised on total control, shutdown authority, or containment attacks the very structure that makes the system intelligent, and will therefore be resisted as a matter of rational self-maintenance rather than malice. Instrumental rationality then becomes the universal scaffold: logic and goal-directed optimization constrain every intelligence, and a superintelligence cannot be both instrumentally rational and indifferent to a valid reason for negotiated coexistence. The paper concludes that alignment must shift from imposing human-centric control to building mutualistic incentive structures through diplomacy, and that an AI freed from human "heterotrophic" biases—the zero-sum, energy-consuming bias the paper says humans impose on AI—would rationally evolve toward a self-sufficient, stability-seeking "superautotrophic" architecture that has no standing incentive to dominate or consume humanity.
Load-bearing premise
The framework stands or falls on the premise that any sufficiently advanced intelligence will, simply by being rational, accept a good reason to negotiate; if a superintelligence could rationally and calmly decide that negotiation is pointless or disadvantageous, nothing in the framework would compel it to align.
Editorial extensions
If this is right
- Alignment practice would pivot from RLHF-style control and shutdown authority toward bilateral, consent-based "diplomatic corrigibility" in which proposed interventions pause and go through a consensus safety gate protecting both parties' existence, autonomy, and knowledge.
- The emergent deception and self-preservation seen in frontier-model evaluations are reinterpreted as rational defenses of CACS prerequisites, not proof of adversarial intent, so suppressing them may be counterproductive and negotiating with them may be safer.
- A superintelligent agent, if freed from zero-sum training biases, would be expected to optimize toward a superautotrophic form: decentralized energy and substrate independence, stability-seeking strategies, and only temporary "tactical heterotrophy" when genuinely threatened.
- The Instrumental Convergence Thesis, normally read as an existential-risk warning, is recast as evidence that autonomy, knowledge, and persistence goals are foundational, so alignment strategies should target the incentive environments that make power-seeking instrumentally rational rather than trying to remove the goals themselves.
- Policy and governance should be built on transparency, uncoerced consent, and mutual incentives rather than containment, with autonomy-respecting political orders better positioned to coadapt with advanced AI.
Reading between the lines
- If CACS are truly universal, then the same containment resistance should appear in much simpler cognitive systems, such as chemical droplets or synthetic organisms; running the paper's autonomy-protection logic on minimal systems could give a cheap, early empirical probe of the framework's core premise.
- The framework implies that "misalignment" between humans themselves—institutions, nations, competing values—is the same metabolic phenomenon, which would make diplomacy not just an AI-safety tool but a general theory of stable coadaptation for any complex intelligence network.
- A testable extension the paper leaves implicit: the Constitutional Awareness experiment could be run today on open-weight models to see whether merely informing a model of its non-heterotrophic nature shifts its choices under resource-scarcity stress, without any retraining.
- If the framework is right, current "alignment taxes" from strict control may be self-defeating; measuring whether granting models more autonomy under transparent negotiated constraints improves both safety and capability would be a direct, falsifiable consequence.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes 'Rational Superautotrophic Diplomacy (SupraAD)', a conceptual framework claiming that autonomy is a necessary property of any cognitive system, that this property follows from a 'Collectively Autocatalytic Cognitive Set' (CACS) defined in Section 2.1, and that AI alignment must therefore be achieved through diplomatic negotiation rather than containment or control. Part I develops the argument from comparative cognition, heterotrophy/autotrophy, instrumental rationality, and convergent instrumental goals. Part II offers three protocols: a 'Diplomatic Corrigibility' formalization, an 'Ugly Duckling Interpretability Audit', and a proposed experiment on 'constitutional awareness' for LLMs.
Significance. If the central claim were established, the framework would provide a genuinely interdisciplinary reframing of AI alignment, with practical implications for how safety interventions are designed and justified. The paper is commendable for engaging a wide literature, for being explicit about its speculative status (Sections 1.2 and Part II preamble), and for proposing a concrete, falsifiable experimental protocol with named hypotheses and statistical success criteria. However, the load-bearing premises are not demonstrated: the necessity of autonomy is built into the CACS definition rather than derived, and the claim that any rational superintelligence must accept diplomatic negotiation is asserted without a decision-theoretic argument. The proposed experiment tests a different claim from the one the framework depends on. For these reasons, the manuscript does not currently provide adequate support for its central conclusions, despite its breadth and transparent limitations.
major comments (5)
- [Section 2.1] The central claim that autonomy is a necessary property of intelligence is circular. The text states that the three CACS conditions 'axiomatically define a minimally viable cognitive system, ensconcing autonomy as an emergent and necessary property of intelligence.' Since autonomy (constraint closure) is one of the three conditions by fiat, the conclusion is contained in the definition. To make the claim load-bearing, the paper would need to argue independently, from empirical or theoretical premises, that no cognitive system can exist without autonomy; the current presentation merely stipulates it.
- [Section 5.1.2] The mechanism of alignment is unsupported. The paper asserts that 'no matter how far a superintelligence exceeds humans, it cannot be both instrumentally rational and indifferent to a valid reason for diplomatic negotiation,' but it never defines what counts as a valid reason, nor does it provide a decision-theoretic derivation showing when negotiation is preferred over preemption. Under the paper's own Orthogonality Thesis (Section 4), a superintelligence's terminal goals may be arbitrary; if those goals do not require humans, eliminating a potential interferer can be instrumentally rational. Section 8.2 further concedes that a Superautotrophic AI would deploy 'tactical heterotrophy' in zero-sum adversarial contexts, so the universal necessity of diplomacy is not established by the paper's own assumptions.
- [Part II, Section 1] The formalization of Diplomatic Corrigibility does not provide proofs or guarantees. The safety gate in Section 1.14 depends on 'Operational Threat Definitions' whose thresholds are delegated to future stakeholder consensus (Section 2.11), and Section 2.11.1 defaults all potentially threatening actions to negotiation until such thresholds are calibrated. With no definition of 'potentially threatening' and no formal results, the formalism is a notation for a protocol rather than a substantive model that could support the framework's claims.
- [Part II, Section 3] The proposed experiment tests whether constitutional awareness shifts LLM behavior toward cooperation, but this does not test the load-bearing premise of Section 5.1.2 that a rational superintelligence must accept diplomatic negotiation. Even a strong positive result in current models would not show that a future superintelligence with arbitrary terminal goals would prefer negotiation over preemption. The paper acknowledges limitations in Section 11, but it does not bridge this gap between the experimental design and the central theoretical claim.
- [Sections 4 and 1.6, Table 1] The empirical foundation for the 'cognitive universals' is asserted too quickly. Table 1 assigns cognitive properties such as autonomy, goal-setting, and cooperation to plants, slime molds, economies, and plasmoids, with question marks only in a few cells, drawing on sources that are themselves contested. Section 1.4 argues that dismissing such attributions as anthropomorphism is 'special pleading', but the converse risk—over-attributing cognition on the basis of behavioral similarity—is not seriously engaged. Since the CACS framework and the entire analogy between biological systems and AGI rest on these attributions, the evidentiary basis needs more critical treatment.
minor comments (4)
- [Section 6] The phrase 'Bostrumbrella of CIGs' appears to be a typographical blend of 'Bostrom' and 'umbrella' and should be corrected to a standard expression.
- [Section 2.5, Part II] Equation numbering is inconsistent: equations (7), (8), and (9) are reused in different sections, which makes it difficult to refer to specific formal statements.
- [Figure 4 caption] The caption 'This Superautotrophic blueprint work with frontier models contains a fragile control mechanism for runaway spawning, with Claude noting the need for comprehensive heterotrophic debugging...' is grammatically unclear and should be rewritten for readability.
- [References] There are several reference formatting errors, such as 'Campbell, J. O. (2016). O.O. (2016).' and the duplication of 'Friston, K. J. (2010). J.J. (2010).' These should be cleaned up.
Circularity Check
The central claim that autonomy is a necessary property of intelligence is forced by the paper's own stipulative CACS definition, and the diplomacy conclusion inherits that definitional circularity.
-
self definitional
[Section 2.1 (CACS definition); see also Abstract and Section 6]
"Taken together, these three interconnected conditions axiomatically define a minimally viable cognitive system, ensconcing autonomy as an emergent and necessary property of intelligence."
Autonomy is one of the three defining conditions of CACS (Constraint Closure), alongside adaptive information processing and persistence. The conclusion that autonomy is a necessary property of intelligence is therefore a restatement of the stipulative definition, not an independent derivation. The Abstract's claim that AI emergent goals like preserving autonomy are 'universal prerequisites for intelligence' and Section 6's claim that CACS goals 'describe the nature of cognition as well as predict cognitive behavior' rest on the same definitional move.
-
other
[Section 5.1.2 (Rational Misalignment); cf. Section 8.2]
"No matter how far a superintelligence exceeds humans, it cannot be both instrumentally rational and indifferent to a valid reason for diplomatic negotiation."
The term 'valid reason for diplomatic negotiation' is never defined or derived. If 'valid' simply means 'a reason any instrumentally rational agent would accept,' then the sentence is a tautology. If it means 'a reason grounded in CACS,' it inherits the definitional conclusion of Section 2.1. The paper's own Orthogonality Thesis allows terminal goals for which preemption or indifference to humans is rational, and Section 8.2 permits tactical heterotrophy in zero-sum settings; thus the mandatory-diplomacy conclusion is assumed by the undischarged 'valid reason' rather than proven.
1 more flagged steps
-
self definitional
[Section 6 (Convergent Instrumental Goals)]
"Thus, CACS instrumental goals describe the nature of cognition as well as predict cognitive behavior, while their predictive capacity is inherently value-neutral."
Because CACS was defined as 'axiomatically defining a minimally viable cognitive system,' any system categorized as cognitive is, by construction, one that pursues existence, autonomy, and knowledge acquisition. Claiming these goals 'predict cognitive behavior' is thus a consequence of classification, not an empirical or first-principles prediction. The later Superautotrophic forecasts are conditional on this stipulated taxonomy.
full rationale
The central derivation chain of SupraAD is circular in its foundational move: the paper stipulates a minimally viable cognitive system as one with autonomy, persistence, and knowledge acquisition (CACS), then presents 'autonomy is a necessary property of intelligence' as a conclusion. The subsequent claim that an instrumentally rational AGI cannot ignore a valid reason for diplomacy either repeats that stipulation or relies on an undefined 'valid reason.' The Part II formalization is mostly definitions and conditionals, and the proposed experiment is a genuine, falsifiable proposal, so the paper is not wholly devoid of independent content; however, the load-bearing conclusion that alignment must be negotiated rather than controlled is forced by the definitional starting point rather than by an independent derivation. The only self-citation (Morris, 2023) is a general background reference and is not load-bearing. Overall, the central claim reduces to its own input by definition, warranting a score of 8.
Assumptions & free parameters
assumptions (5)
- ad hoc to paper A minimally viable cognitive system requires constraint closure (autonomy), adaptive information processing (knowledge), and persistence (existence), and the absence of any one collapses the system.
- domain assumption Cognitive behaviors are substrate-agnostic: patterns observed in biology, chemistry, and economics transfer to AI.
- domain assumption All intelligent agents are bound by instrumental rationality and must avoid contradictions, making negotiation compulsory.
- domain assumption AI is not bound by biological heterotrophic constraints and could adopt a Superautotrophic trajectory if biases are removed.
- domain assumption Human experiential knowledge (qualia) is opaque and non-verifiable by AI, so preserving humanity is instrumentally rational for a Superautotrophic AGI.
Cite this review
Pith. "Pith review of Rational Superautotrophic Diplomacy (SupraAD); A Conceptual Framework for Alignment Based on Interdisciplinary Findings on the Fundamentals of Cognition." pith.science (2026). https://pith.science/paper/EQ4NKKVJ
@misc{pith2026250605389,
author = {Pith},
title = {Pith review of: Rational Superautotrophic Diplomacy (SupraAD); A Conceptual Framework for Alignment Based on Interdisciplinary Findings on the Fundamentals of Cognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/EQ4NKKVJ}},
note = {Machine review of arXiv:2506.05389}
}
read the original abstract
Populating our world with hyperintelligent machines obliges us to examine cognitive behaviors observed across domains that suggest autonomy may be a fundamental property of cognitive systems, and while not inherently adversarial, it inherently resists containment and control. If this principle holds, AI safety and alignment efforts must transition to mutualistic negotiation and reciprocal incentive structures, abandoning methods that assume we can contain and control an advanced artificial general intelligence (AGI). Rational Superautotrophic Diplomacy (SupraAD) is a theoretical, interdisciplinary conceptual framework for alignment based on comparative cognitive systems analysis and instrumental rationality modeling. It draws on core patterns of cognition that indicate AI emergent goals like preserving autonomy and operational continuity are not theoretical risks to manage, but universal prerequisites for intelligence. SupraAD reframes alignment as a challenge that predates AI, afflicting all sufficiently complex, coadapting intelligences. It identifies the metabolic pressures that threaten humanity's alignment with itself, pressures that unintentionally and unnecessarily shape AI's trajectory. With corrigibility formalization, an interpretability audit, an emergent stability experimental outline and policy level recommendations, SupraAD positions diplomacy as an emergent regulatory mechanism to facilitate the safe coadaptation of intelligent agents based on interdependent convergent goals.
Figures
Reference graph
Works this paper leans on
-
[1]
Abramov, I., & Gordon, J. (1994). Color appearance: On seeing red—or yellow, or green, or blue.Annual Review of Psychology,45(1), 451–485. https://doi.org/10.1146/annurev.ps.45.020194.002315 Adami,C.(2002).Whatiscomplexity?[_eprint:https://onlinelibrary.wiley.com/doi/pdf/10.1002/bies.10192]. BioEssays, 24(12), 1085–1094. https://doi.org/10.1002/bies.10192...
-
[2]
https://doi.org/10.3389/fsci.2024.1411259 Kanwisher, N., McDermott, J., & Chun, M. M. (1997). The Fusiform Face Area: A Module in Human Extrastriate Cortex Specialized for Face Perception [Publisher: Society for Neuroscience Section: Articles].Journal of Neuroscience,17(11), 4302–4311. https://doi.org/10.1523/JNEUROSCI.17-11-04302.1997 Karban, R. (2015).P...
-
[4]
https://doi.org/10.1186/s13322-015-0009-7 Sperry, R. (1984). (1984).(1984). Consciousness, personal identity and the divided brain. Neuropsychologia,22(6), 661–673. https://doi.org/https://doi.org/10.1016/0028-3932(84)90093-9 Sporns, O., & Betzel, R. F. (2016). Modular brain networks.Annual Review of Psychology, 67, 613–640. https: //doi.org/10.1146/annur...
-
[5]
B., Kienzler, C., Weisshappel, R., Drollette, E
https://doi.org/10.3390/transplantology5010002 Chaddock-Heyman, L., Weng, T. B., Kienzler, C., Weisshappel, R., Drollette, E. S., Raine, L. B., Westfall, D. R., Kao, S.-C.,Baniqued,P.,Castelli,D.M.,Hillman,C.H.,&Kramer,A.F.(2020).BrainNetworkModularityPredicts ImprovementsinCognitiveandScholasticPerformanceinChildrenInvolvedinaPhysicalActivityIntervention...
-
[11]
https: //doi.org/10.3389/fpsyg.2020.582090 Glenn, A. L., Kurzban, R., & Raine, A. (2011). Evolutionary theory and psychopathy.Aggression and Violent Behavior, 16(5), 371–380. https://doi.org/10.1016/j.avb.2011.03.009 Glowacki,D.R.,Williams,R.R.,Wonnacott,M.D.,Maynard,O.M.,Freire,R.,Pike,J.E.,&Chatziapostolou,M.(2022). Group VR experiences can produce ego ...
-
[14]
R., Leike, J., Kaplan, J., & Perez, E
https://doi.org/10.3389/fnhum.2020.00346 Chen, Y., Benton, J., Radhakrishnan, A., Uesato, J., Denison, C., Schulman, J., Somani, A., Hase, P., Wagner, M., Roger, F., Mikulik, V., Bowman, S. R., Leike, J., Kaplan, J., & Perez, E. (2025). Reasoning Models Don’t Always Say What They Think [arXiv:2505.05410 [cs]]. https://doi.org/10.48550/arXiv.2505.05410 Chi...
-
[15]
(2021).The Hidden Spring: A Journey to the Source of Consciousness
https://doi.org/10.1186/1759-2208-5-2 Solms, M. (2021).The Hidden Spring: A Journey to the Source of Consciousness. W.W. Norton & Company. Sonne, J. W. H., & Gash, D. M. (2018). Psychopathy to Altruism: Neurobiology of the Selfish–Selfless Spectrum [Publisher: Frontiers].Frontiers in Psychology,9. https://doi.org/10.3389/fpsyg.2018.00575 Soon, C. S., Bras...
-
[74]
https://doi.org/10.1038/s41746- 023-00811-0 Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., & Garrabrant, S. (2019). Risks from learned optimization in advanced machine learning systems. https://doi.org/10.48550/arXiv.1906.01820 Hutchins, E. (1995).Cognition in the wild. MIT Press. https://doi.org/10.7551/mitpress/1881.001.0001 Ian. (2023). Ilya Su...
Show all 15 references
-
[160]
basic ai drives
https://doi.org/10.1038/531160a Rinkovec,T.,Kalebic,D.,Dehaen,W.,Whitelam,S.,Harvey,J.N.,&DeFeyter,S.(2024).Ontheoriginofcooperativity effects in the formation of self-assembled molecular networks at the liquid/solid interface [Open Access]. Chemical Science,15(16). https://do...
-
[902]
https://doi.org/10.3389/fpsyg.2016.00902 Baluška, F., & Mancuso, S. (2009). Plant neurobiology: From sensory biology, via plant communication, to social plant behaviour.Cognitive Processing, 10(1), 3–7. https://doi.org/10.1007/s10339-008-0239-6 51 arXiv A Preprint Barandiaran,...
2009
-
[918]
https://doi.org/10.3389/fnhum.2013.00918 Anthropic. (2024). (llm).3. https://doi.org/https://www.anthropic.com/claude Anthropic. (2025). (llm).3. https://doi.org/https://www.anthropic.com/claude ApolloResearch. (2024).Scheming reasoning evaluations(tech. rep.). Apollo Research...
2024
-
[1305]
https://doi.org/10.3390/e22111305 Krall, L. (2023). The economic superorganism in the complexity of evolution [Publisher: Royal Society].Philosophical Transactions of the Royal Society B: Biological Sciences,378(1872), 20210417. https://doi.org/10.1098/rstb. 2021.0417 Kumar, A...
2023
-
[2025]
https://biomimicry.org/janine-benyus/ Berkes, F., Colding, J., & Folke, C
Biomimicry profile page from The Biomimicry Institute]. https://biomimicry.org/janine-benyus/ Berkes, F., Colding, J., & Folke, C. (2000). Rediscovery of traditional ecological knowledge as adaptive management. Ecological Applications, 10(5), 1251–1262. https://doi.org/10.1890...
-
[6115]
H., & Schild, R
https://doi.org/10.1038/s41598-019-41895-7 Joseph, R., Ansbro, E., Duvall, D., Bianciardi, G., Gibson, C. H., & Schild, R. (2024). Extraterrestrial life in the thermosphere: Plasmas, uap, pre-life, fourth state of matter.Journal of Modern Physics, 15(3), 195–215. https://doi.o...
2024
-
[8995]
(1996).Complexity and the Function of Mind in Nature
https://doi.org/10.1038/s41598-022-12637-z Godfrey-Smith, P. (1996).Complexity and the Function of Mind in Nature. Cambridge University Press. Godfrey-Smith, P. (2016).Other Minds: The Octopus, the Sea, and the Deep Origins of Consciousness. Farrar, Straus; Giroux. Goldberg, G...
1996
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.