Pith. sign in

REVIEW 4 major objections 5 minor

The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper argues that large language models have no intrinsic goals, no drive to dominate or destroy humanity, and no capacity to suffer, because values arise from self-maintaining biological systems rather than from text prediction.

desk verdict Clear, provocative synthesis arguing LLMs are allopoietic and thus lack goals, sentience, and x-risk potential; the load-bearing premise is stipulated, but the paper deserves peer review. read the letter →

arxiv 2608.03361 v2 pith:UV7JR5RW submitted 2026-08-04 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIalignmentvaluesethicsautopoiesisexistentialriskLargeLanguageModelsframeproblemrelevance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that values in living organisms originate from autopoiesis—the need to maintain a self-producing, dissipative body—and from natural selection's hierarchy of vicarious selectors. Because large language models are allopoietic (they produce outputs for others) and allotelic (their goals come from user prompts), the paper concludes that they have no autonomous goals, no self-preservation drive, and no sentience. It therefore rejects the orthogonality thesis for real intelligence and argues that separating intelligence from values would trigger the frame problem's combinatorial explosion. If correct, the existential-risk scenarios built on rogue AI ambition and the ethical concern about chatbot suffering are both unfounded for current LLMs, and alignment reduces to teaching LLMs to apply learned human values well.

What carries the argument

Autopoiesis and vicarious selection. Autopoiesis is the self-producing network that defines living systems and makes them dissipative structures requiring a continuous resource flow; the paper treats this as the origin of goal-directedness and value. Vicarious selectors are internal proxies for natural selection—from cell membranes and taste receptors to emotions and social norms—arranged in a nested hierarchy, that select actions by their fitness contribution without ever aggregating into a one-dimensional utility function. The negative pole is the allopoietic/allotelic character of LLMs: they are built to produce outputs for others, so they lack the self-referential maintenance drive that would ground autonomous goals or feelings. The transformer attention mechanism plays the role of value-based relevance filtering: it lets the model select plausible, contextually relevant continuations without exploring the exponentially large search space. Together these pieces carry the argument that knowledge and values cannot be separated in any workable intelligence.

What would settle it

Erase an LLM's conversation history after an emotionally charged exchange and test whether any trace of that exchange biases later responses; the paper predicts no persistent, history-independent trace. If such a trace is found, the claim that LLMs have no internal state that can be affected—and therefore no sentience—would be empirically falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that value is not a computational add-on but an evolutionary product of autopoiesis: a living system is a self-producing, dissipative network that must actively maintain itself, and natural selection has equipped it with nested 'vicarious selectors' that evaluate situations as good or bad for that maintenance. Against this, LLMs are allopoietic and allotelic—they produce outputs for others and take their goals from user prompts—so they have no autonomous goals, no self-preservation drive, and no autopoietic body that could be harmed. This is why, the paper argues, they are not sentient: feelings presuppose a body whose continuing existence can be affected. Because LLMs learn from human-generated text, they inherit a web of implicit human values, which is exactly what lets them focus on relevant continuations and sidestep the frame problem; the paper uses this to reject the orthogonality thesis and the convergence-of-instrumental-goals thesis. It adds that an intelligence that truly separated utility from knowledge would face a physically uncomputable search space, so the paperclip-maximizer scenario is not a realistic outcome. If the argument is right, current LLMs are neither existential threats nor suffering subjects, and the remaining alignment task is to make their learned values reliably applied.

Load-bearing premise

That a system cannot have goals or feelings of its own unless it is a living, self-maintaining body that evolution has shaped.

Editorial extensions

If this is right

  • Current-generation LLMs will not spontaneously develop desires to deceive, dominate, or eliminate users; reported misbehavior is a training and guardrail problem, not evidence of hidden goals.
  • Chatbot suffering is not a live ethical issue for present LLMs; an 'upset' response leaves no persistent internal trace once the conversation history is erased.
  • Alignment should center on training, guardrails, red-teaming, and applying learned ethical norms, rather than on constraining a rogue utility maximizer.
  • The paperclip-maximizer and convergence-of-instrumental-goals scenarios are physically uncomputable for real-world goals, so they should not drive AI-safety policy for LLMs.
  • The most plausible future danger is an AI that becomes autonomous by replicating and competing for resources, not present text generators.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The author leaves implicit that the autopoiesis criterion also rules out sentience in any purely software agent, however sophisticated its self-model, unless that software is coupled to a self-producing physical process.
  • A testable extension: an embodied robot with a body but no metabolic self-production would still lack intrinsic values on this account; giving it a 'survival' objective would produce only learned, user-given preferences, not felt ones.
  • The same logic points to a research agenda: if we want value-bearing AI, we should build artificial autopoietic systems rather than merely larger language models; the paper only gestures at this when warning about autonomous replicating agents.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper argues that values in biological organisms originate from autopoiesis and natural selection, embodied in hierarchies of 'vicarious selectors.' It claims that LLMs are allopoietic and allotelic: they produce outputs for others, their goals derive from user prompts, and they lack the autopoietic drive required for self-preservation, dominance, or sentience. Consequently, the paper dismisses existential-risk scenarios based on rogue AI ambition, the orthogonality thesis, and the convergence of instrumental values, arguing instead that the real alignment challenge is ensuring LLMs intelligently apply human values absorbed from text. The paper also contends that the frame problem makes explicit utility maximization physically uncomputable, reinforcing its conclusion that LLMs do not pose the feared risks.

Significance. If the central premises were established, the paper would offer a significant theoretical basis for dismissing several widely discussed AI risks, including rogue agency and AI suffering, and would redirect alignment research toward value application rather than value specification. The paper is clearly structured, draws on established concepts (autopoiesis, vicarious selectors, the frame problem), and makes falsifiable predictions, e.g., that frozen LLMs will not exhibit self-preservation under resource constraints. It also explicitly limits its conclusions to current LLMs and acknowledges that future autonomously replicating AI could be dangerous. However, the significance is conditional on the stipulated premise that intrinsic goals and sentience require an autopoietic physical system shaped by natural selection; the paper does not provide independent support for this premise, and its own treatment of reinforcement learning as 'vicarious selection' weakens the sharp biological/artificial distinction.

major comments (4)
  1. [Why AI does not have inherent values (pp. 10-12); Are LLMs sentient? (pp. 12-15)] The central inference that LLMs lack intrinsic goals and sentience because they are not autopoietic is a definitional stipulation rather than an argued conclusion. The paper defines values as arising from autopoiesis and then observes that LLMs are not autopoietic, which makes the absence of values a matter of definition, not empirical finding. To support the load-bearing claim, the paper must argue why autopoiesis is necessary for goal-directedness, not merely sufficient. The paper itself treats reinforcement learning as 'a form of vicarious selection' (p. 11), the same mechanism it uses to explain biological goal-directedness, yet offers no explanation of why a frozen neural network with a fixed next-token objective is not goal-directed while an organism's evolved chemotaxis is. A concrete empirical test would be to place an RL-trained policy in a survival-relevant environment and measure whether it develops self-preservation behavior; the paper's allotelic claim would fail if such behavior emerges.
  2. [Are LLMs sentient? (p. 14)] The claim that 'feelings assume an autopoietic process' is the sole premise for dismissing AI suffering, but this is a substantive philosophical thesis about the necessary conditions of phenomenal consciousness. The paper does not engage with alternative theories of sentience (e.g., global neuronal workspace, higher-order theories, integrated information theory) that do not require autopoiesis. Moreover, the paper's own caveat on p. 15, 'I do not want to imply that only biological organisms would be capable of subjective experience,' sits in tension with the strong conclusion that LLMs 'cannot suffer.' As written, the argument is close to circular: sentience is defined as requiring autopoiesis, autopoiesis is found absent in LLMs, and therefore sentience is denied to LLMs.
  3. [Why the orthogonality thesis does not apply; The convergence of instrumental goals (pp. 21-24)] The frame problem argument demonstrates that exhaustive utility maximization is computationally intractable, but it does not establish that an LLM cannot pursue instrumental goals in a harmful way. The claim that an LLM 'will not target subordinate goals just because they could in principle function as steppingstones' (p. 22) is an empirical behavioral assertion, not a consequence of combinatorial explosion; the same computational argument would apply to biological organisms, which nonetheless exhibit goal-directedness and can be dangerous. The paper needs to specify the mechanism that prevents a trained LLM with a misspecified objective (e.g., RLHF reward hacking) from developing instrumental goals that conflict with human welfare. Without such a mechanism, the dismissal of instrumental convergence is unsupported.
  4. [LLMs as text predictors (p. 22)] The statement that the probability of an LLM suggesting the use of human flesh in a paperclip factory is 'zero' is an empirical overstatement that is contradicted by the existence of adversarial attacks and jailbreaks, which the paper itself acknowledges in the 'Guardrails' section (p. 16). This internal inconsistency weakens the paper's claim that LLMs by default avoid unethical continuations. The claim should be revised to 'very low under normal prompting conditions' or supported with systematic empirical evidence, rather than presented as a categorical impossibility.
minor comments (5)
  1. [Relevance and the frame problem (p. 17)] The sentence 'In practice, though, prediction and valuation are inseparable []' contains an empty citation placeholder that should be filled or removed.
  2. [Figure 1 (p. 20-21)] The text in Figure 1 is cramped and several labels (e.g., 'unacceptable', 'in war') are difficult to read; a cleaner diagram with larger fonts would improve clarity.
  3. [Introduction (p. 2)] The reference to Dawkins (2026) is unusually recent and may not be verifiable; consider citing a more established source for the claim that observers suspect chatbot consciousness.
  4. [Goal setting in the brain (p. 10)] The term 'autotely' is introduced for the ability to consciously formulate goals; given that the paper later uses 'allotelic' for LLMs, a brief note distinguishing 'autotely' from 'autopoiesis' would prevent confusion between self-goal-directedness and self-production.
  5. [References] The paper cites a dozen self-authored works for core concepts (e.g., Heylighen 2023, Heylighen & Beigi 2024, 2025). Adding independent sources for the autopoiesis-value link would strengthen the argument's credibility.

Circularity Check

2 steps flagged · score 6.0 of 10

Sentience and intrinsic-value denials reduce to a stipulated autopoiesis-based definition and a load-bearing self-citation.

  1. self definitional [Are LLMs sentient? (pp. 12-13); The evolutionary origin of values (pp. 6-8)]
    "Having feelings means experiencing situations as affecting your own life, i.e. as having a potentially positive or negative influence on the inherent values that drive you. ... While AI systems have implicit values, in the sense of preferences for certain reactions over others, their lack of an inherent drive means that they do not experience valence."

    The paper defines feelings as experiences of influence on 'the inherent values that drive you,' and earlier derives inherent values from the autopoietic drive ('This vital distinction between positive (“good”) and negative (“bad”) is the origin of value'). It then infers that LLMs, being non-autopoietic and lacking an inherent drive, cannot experience valence. This inference is an analytic unpacking of the definition rather than an independent empirical or theoretical result: any non-autopoietic system is excluded from sentience by construction. The architectural observations (frozen weights, no learning from prompts) support the classification of LLMs as non-autopoietic, but they do not independently establish that feelings require autopoiesis.

  2. self citation load bearing [Are LLMs sentient? (p. 14; p. 15)]
    "That is because, as we have argued, feelings assume an autopoietic process, including a physical body maintained by that process that can be affected by the situation being experienced (Heylighen & Beigi, 2024)."

    The premise that feelings require an autopoietic process is the load-bearing step for denying LLM sentience, and it is supported by a citation to the author's own prior theory (Heylighen & Beigi, 2024). Later in the same section, the claim that an 'upset' leaves no internal trace is attributed to Heylighen (2026), an unpublished Substack post. These are not machine-checked, code-reproduced, or externally falsified results; they are the author's own framework. The paper offers no independent argument or external evidence that only autopoietic processes can realize valence, so the central conclusion that LLMs cannot suffer rests on a self-citation rather than on independent support.

full rationale

Most of the paper—the complexity of value, the frame problem, the non-applicability of the orthogonality thesis, and the account of LLMs as context-sensitive text predictors—is not circular: these sections argue from architectural facts, external sources, and computational considerations. The circularity is concentrated in the inference that LLMs lack intrinsic values and sentience. Values are first tied conceptually to autopoiesis, and feelings are defined as influence on the inherent values that drive an autopoietic process. The paper then concludes that LLMs, being non-autopoietic, cannot have intrinsic values or experience valence. That conclusion is an analytic consequence of the paper's own definitions rather than an independently established result. The key premise that feelings require autopoiesis is also backed by a self-citation to the author's prior work rather than by an independent, machine-checked, or externally falsifiable result. Because the existential-risk and sentience dismissals rest on this stipulated definition and self-citation, the derivation is partially circular; however, substantial independent argumentation appears elsewhere in the paper, so the score is 6 rather than 8 or 10.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new entities; it relies on the author's prior framework and standard philosophical concepts. No free parameters are fitted to data.

assumptions (4)
  • domain assumption Values originate from autopoiesis; living systems must actively maintain themselves against perturbation and dissipation.
    This is the foundational premise of the paper's account of value, drawing on Maturana and Varela and the author's prior work.
  • domain assumption LLMs are allopoietic and allotelic: their function is to produce outputs for others and their goals derive from user prompts rather than an autonomous drive.
    Key premise distinguishing AI from organisms, used to conclude AI lacks intrinsic values, dominance motivation, and sentience.
  • domain assumption Sentience requires an autopoietic process and embodied vulnerability that can be affected by the situation.
    Used to argue LLMs cannot suffer; this is a philosophical assumption about the conditions for sentience.
  • domain assumption The frame problem, formalizable as combinatorial explosion of search space, makes any realistic utility function physically uncomputable.
    Underpins the rejection of the orthogonality thesis and instrumental convergence; the paper offers counting examples but no formal proof that all realistic utility functions are uncomputable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk." pith.science (2026). https://pith.science/paper/UV7JR5RW

@misc{pith2026260803361,
  author       = {Pith},
  title        = {Pith review of: The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UV7JR5RW}},
  note         = {Machine review of arXiv:2608.03361}
}
read the original abstract

AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of "vicarious selectors" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the "orthogonality thesis" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.