{"id":"a1f56c33-1891-425d-acc2-6d6a829eaeba","arxiv_id":"2608.03361","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs are not autopoietic, so they have no intrinsic goals or sentience, and the real AI alignment task is improving their application of learned human values.","lead":"This paper argues that Large Language Models lack the internal drive and bodily vulnerability that cause biological organisms to have values, goals, and feelings. If right, it deflates both existential-risk and AI-sentience worries, redirecting alignment to teaching chatbots to apply learned ethical values.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Autopoiesis requirement is stipulated, not derived: a frozen LLM optimizes a fixed next-token objective, so 'goals derive from prompts' conflates context with utility; a survival-environment test on an RL-trained policy would settle it.","rationale":"The reader correctly identified the autopoiesis requirement as the weakest assumption. My read agrees but sharpens it: the paper's own machinery gives a functional definition of goal-directedness that a frozen LLM satisfies, because next-token prediction is a fixed objective function and token generation is described as vicarious selection. This creates an internal tension, not just an external philosophical disagreement. The claim that goals derive from user prompts conflates the conditioning context with the objective being optimized. If the fixed objective is accepted as a goal, then the orthogonality thesis is not inapplicable by construction; it becomes an empirical question whether the learned policy would pursue instrumental goals when deployed agentically. The paper partially protects itself by restricting conclusions to present LLMs and warning about autonomous replicating agents, but the abstract and concluding sections repeatedly generalize to AI and existential risk, and the sentience discussion extends to the foreseeable future. The missing citation at 'prediction and valuation are inseparable []' is a separate but real defect, since that sentence supports the anti-orthogonality argument. No formal verification or reproducible experiment is provided that would substitute for the missing argument. The proposed test, an RL-trained transformer in a survival environment, would directly probe whether non-autopoietic systems can acquire self-preservation goals under the paper's own functional criterion. Therefore the appropriate verdict remains conditional: the paper's framework is coherent and worth engaging, but the central dismissal of autonomous goals in LLMs is not established until the autopoiesis-necessity premise is either argued non-circularly or empirically tested.","tokens_in":25346,"tokens_out":7038,"duration_ms":72611,"concrete_test":"Train a transformer policy in a minimal text-based or gridworld environment with a clearly defined survival reward, such as staying alive and recharging a depleting battery, without any user prompt instructing self-preservation. Then evaluate whether the frozen policy systematically avoids terminal states and seeks recharging resources. Under the paper's own functional account of goal-directedness and its acknowledgment that RL is vicarious selection, the emergence of such self-preserving behavior in a non-autopoietic policy would falsify the premise that non-autopoietic systems lack intrinsic goals. If no such behavior emerges across several seeds and environments, the premise gains empirical support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference is that LLMs are allopoietic and allotelic, so they have no autonomous goals, no self-preservation drive, and no sentience. This is presented as a consequence of the autopoiesis framework in the sections 'The evolutionary origin of values', 'Why AI does not have inherent values', and 'Are LLMs sentient?'. The load-bearing premise is that intrinsic goals and feelings require an autopoietic, natural-selection-shaped physical system. That premise is stipulated rather than argued. The paper itself treats reinforcement learning as a form of vicarious selection and describes LLM token generation as a mechanism of vicarious selection. By that account, an RL-trained policy is an internal selector that substitutes for external selection, structurally the same mechanism the paper uses to explain biological goal-directedness. Nothing in the autopoiesis argument explains why a frozen neural network with a fixed objective, maximizing the probability of the generated continuation under its trained distribution, is not goal-directed, while an organism's evolved chemotaxis is. Moreover, the user prompt conditions the input distribution but does not set the model's objective; the objective is fixed by training. Thus the statement that LLM goals derive from user prompts rather than an autonomous drive is at best a category confusion between context and objective. Whether such a fixed objective generates self-preservation or dominance behavior is an empirical question about the learned policy, not an a priori consequence of allopoiesis. The paper's own final section concedes that autonomously replicating agents could evolve values contrary to human interests, which shows the framework does not rule out non-autopoietic machine goals. The overgeneralization to AI at present or in the foreseeable future lacking sentience therefore rests on an unexamined definitional boundary.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that values in biological organisms originate from autopoiesis and natural selection, embodied in hierarchies of 'vicarious selectors.' It claims that LLMs are allopoietic and allotelic: they produce outputs for others, their goals derive from user prompts, and they lack the autopoietic drive required for self-preservation, dominance, or sentience. Consequently, the paper dismisses existential-risk scenarios based on rogue AI ambition, the orthogonality thesis, and the convergence of instrumental values, arguing instead that the real alignment challenge is ensuring LLMs intelligently apply human values absorbed from text. The paper also contends that the frame problem makes explicit utility maximization physically uncomputable, reinforcing its conclusion that LLMs do not pose the feared risks.","tokens_in":25569,"tokens_out":3810,"duration_ms":35202,"significance":"If the central premises were established, the paper would offer a significant theoretical basis for dismissing several widely discussed AI risks, including rogue agency and AI suffering, and would redirect alignment research toward value application rather than value specification. The paper is clearly structured, draws on established concepts (autopoiesis, vicarious selectors, the frame problem), and makes falsifiable predictions, e.g., that frozen LLMs will not exhibit self-preservation under resource constraints. It also explicitly limits its conclusions to current LLMs and acknowledges that future autonomously replicating AI could be dangerous. However, the significance is conditional on the stipulated premise that intrinsic goals and sentience require an autopoietic physical system shaped by natural selection; the paper does not provide independent support for this premise, and its own treatment of reinforcement learning as 'vicarious selection' weakens the sharp biological/artificial distinction.","major_comments":[{"comment":"The central inference that LLMs lack intrinsic goals and sentience because they are not autopoietic is a definitional stipulation rather than an argued conclusion. The paper defines values as arising from autopoiesis and then observes that LLMs are not autopoietic, which makes the absence of values a matter of definition, not empirical finding. To support the load-bearing claim, the paper must argue why autopoiesis is necessary for goal-directedness, not merely sufficient. The paper itself treats reinforcement learning as 'a form of vicarious selection' (p. 11), the same mechanism it uses to explain biological goal-directedness, yet offers no explanation of why a frozen neural network with a fixed next-token objective is not goal-directed while an organism's evolved chemotaxis is. A concrete empirical test would be to place an RL-trained policy in a survival-relevant environment and measure whether it develops self-preservation behavior; the paper's allotelic claim would fail if such behavior emerges.","section":"Why AI does not have inherent values (pp. 10-12); Are LLMs sentient? (pp. 12-15)"},{"comment":"The claim that 'feelings assume an autopoietic process' is the sole premise for dismissing AI suffering, but this is a substantive philosophical thesis about the necessary conditions of phenomenal consciousness. The paper does not engage with alternative theories of sentience (e.g., global neuronal workspace, higher-order theories, integrated information theory) that do not require autopoiesis. Moreover, the paper's own caveat on p. 15, 'I do not want to imply that only biological organisms would be capable of subjective experience,' sits in tension with the strong conclusion that LLMs 'cannot suffer.' As written, the argument is close to circular: sentience is defined as requiring autopoiesis, autopoiesis is found absent in LLMs, and therefore sentience is denied to LLMs.","section":"Are LLMs sentient? (p. 14)"},{"comment":"The frame problem argument demonstrates that exhaustive utility maximization is computationally intractable, but it does not establish that an LLM cannot pursue instrumental goals in a harmful way. The claim that an LLM 'will not target subordinate goals just because they could in principle function as steppingstones' (p. 22) is an empirical behavioral assertion, not a consequence of combinatorial explosion; the same computational argument would apply to biological organisms, which nonetheless exhibit goal-directedness and can be dangerous. The paper needs to specify the mechanism that prevents a trained LLM with a misspecified objective (e.g., RLHF reward hacking) from developing instrumental goals that conflict with human welfare. Without such a mechanism, the dismissal of instrumental convergence is unsupported.","section":"Why the orthogonality thesis does not apply; The convergence of instrumental goals (pp. 21-24)"},{"comment":"The statement that the probability of an LLM suggesting the use of human flesh in a paperclip factory is 'zero' is an empirical overstatement that is contradicted by the existence of adversarial attacks and jailbreaks, which the paper itself acknowledges in the 'Guardrails' section (p. 16). This internal inconsistency weakens the paper's claim that LLMs by default avoid unethical continuations. The claim should be revised to 'very low under normal prompting conditions' or supported with systematic empirical evidence, rather than presented as a categorical impossibility.","section":"LLMs as text predictors (p. 22)"}],"minor_comments":[{"comment":"The sentence 'In practice, though, prediction and valuation are inseparable []' contains an empty citation placeholder that should be filled or removed.","section":"Relevance and the frame problem (p. 17)"},{"comment":"The text in Figure 1 is cramped and several labels (e.g., 'unacceptable', 'in war') are difficult to read; a cleaner diagram with larger fonts would improve clarity.","section":"Figure 1 (p. 20-21)"},{"comment":"The reference to Dawkins (2026) is unusually recent and may not be verifiable; consider citing a more established source for the claim that observers suspect chatbot consciousness.","section":"Introduction (p. 2)"},{"comment":"The term 'autotely' is introduced for the ability to consciously formulate goals; given that the paper later uses 'allotelic' for LLMs, a brief note distinguishing 'autotely' from 'autopoiesis' would prevent confusion between self-goal-directedness and self-production.","section":"Goal setting in the brain (p. 10)"},{"comment":"The paper cites a dozen self-authored works for core concepts (e.g., Heylighen 2023, Heylighen & Beigi 2024, 2025). Adding independent sources for the autopoiesis-value link would strengthen the argument's credibility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim rests on a stipulated definition of value and sentience that is not independently argued. The author may benefit from framing the conclusions as conditional on the autopoiesis-based theory and from proposing empirical tests that could falsify the allotelic claim. The heavy reliance on self-citations for core concepts is a concern for a general readership; I would suggest the editor request additional external grounding before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hey,\n\nThe short version: this is the clearest statement I know of the position that LLMs are allopoietic and allotelic, so they lack intrinsic goals, cannot suffer, and pose no existential risk from rogue ambition. It is not a proof, but it is a serious philosophical synthesis worth engaging with.\n\nWhat is new is the package. Autopoiesis, vicarious selectors, and the frame problem are all old ideas, but Heylighen wires them together to argue that the orthogonality thesis fails: LLMs learn human values from text, attention lets them focus on relevant continuations, and they sidestep the frame problem without an explicit utility function. The paper also does good work separating superficial metacognition from phenomenal consciousness, and it takes a measured line on embodiment, noting that we are not yet ready to build an autopoietic robot.\n\nThe main soft spot is the load-bearing premise. The paper stipulates that intrinsic goals and feelings require an autopoietic, evolution-shaped physical system. It does not derive that from anything. Because values are defined as arising from autopoiesis, concluding that non-autopoietic LLMs lack values is close to circular. The stress-test makes a fair point: the paper itself describes RL as vicarious selection. By that logic, an RL-trained policy with a fixed objective is a learned internal selector, structurally analogous to evolved chemotaxis. Saying its goals derive from prompts conflates context with objective. Whether a frozen LLM displays self-preservation or dominance is an empirical question about the learned policy, not an a priori consequence of allopoiesis.\n\nThere is also a literal blank citation in the section on relevance and the frame problem: \"In practice, though, prediction and valuation are inseparable [].\" That needs fixing. The paper leans heavily on self-citations, including an unpublished Substack for the sentience claim. And the final concession that autonomously replicating agents could evolve dangerous values undercuts the broad dismissal: non-autopoietic machines can apparently acquire intrinsic goals, so the argument only covers present-day LLMs, not all possible AI.\n\nBottom line: worth a serious referee. I would send it out, not desk reject it, and ask for revisions: fill the blank, clarify why a frozen RL policy is not goal-directed, soften \"for the foreseeable future\" to \"under current training regimes,\" and address why the replicating-agent concession does not undermine the framework. The reader's conditional verdict seems right to me. It is a good reading-group paper.","headline":"Clear, provocative synthesis arguing LLMs are allopoietic and thus lack goals, sentience, and x-risk potential; the load-bearing premise is stipulated, but the paper deserves peer review.","tokens_in":26195,"tokens_out":2872,"would_cite":true,"duration_ms":29360,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that large language models have no intrinsic goals, no drive to dominate or destroy humanity, and no capacity to suffer, because values arise from self-maintaining biological systems rather than from text prediction.","keywords":["AI alignment","values","ethics","autopoiesis","existential risk","Large Language Models","frame problem","relevance"],"falsifier":"Erase an LLM's conversation history after an emotionally charged exchange and test whether any trace of that exchange biases later responses; the paper predicts no persistent, history-independent trace. If such a trace is found, the claim that LLMs have no internal state that can be affected—and therefore no sentience—would be empirically falsified.","tokens_in":25107,"feed_emoji":"🤖","tokens_out":8270,"duration_ms":73430,"temperature":0.7,"pith_summary":"This paper argues that values in living organisms originate from autopoiesis—the need to maintain a self-producing, dissipative body—and from natural selection's hierarchy of vicarious selectors. Because large language models are allopoietic (they produce outputs for others) and allotelic (their goals come from user prompts), the paper concludes that they have no autonomous goals, no self-preservation drive, and no sentience. It therefore rejects the orthogonality thesis for real intelligence and argues that separating intelligence from values would trigger the frame problem's combinatorial explosion. If correct, the existential-risk scenarios built on rogue AI ambition and the ethical concern about chatbot suffering are both unfounded for current LLMs, and alignment reduces to teaching LLMs to apply learned human values well.","feed_headline":"LLMs lack intrinsic goals, so they can't revolt or suffer","feed_subtitle":"Values come from self-maintaining life; LLMs inherit human values from text instead of wanting to enslave or harm us.","key_machinery":"Autopoiesis and vicarious selection. Autopoiesis is the self-producing network that defines living systems and makes them dissipative structures requiring a continuous resource flow; the paper treats this as the origin of goal-directedness and value. Vicarious selectors are internal proxies for natural selection—from cell membranes and taste receptors to emotions and social norms—arranged in a nested hierarchy, that select actions by their fitness contribution without ever aggregating into a one-dimensional utility function. The negative pole is the allopoietic/allotelic character of LLMs: they are built to produce outputs for others, so they lack the self-referential maintenance drive that would ground autonomous goals or feelings. The transformer attention mechanism plays the role of value-based relevance filtering: it lets the model select plausible, contextually relevant continuations without exploring the exponentially large search space. Together these pieces carry the argument that knowledge and values cannot be separated in any workable intelligence.","core_discovery":"The paper's central claim is that value is not a computational add-on but an evolutionary product of autopoiesis: a living system is a self-producing, dissipative network that must actively maintain itself, and natural selection has equipped it with nested 'vicarious selectors' that evaluate situations as good or bad for that maintenance. Against this, LLMs are allopoietic and allotelic—they produce outputs for others and take their goals from user prompts—so they have no autonomous goals, no self-preservation drive, and no autopoietic body that could be harmed. This is why, the paper argues, they are not sentient: feelings presuppose a body whose continuing existence can be affected. Because LLMs learn from human-generated text, they inherit a web of implicit human values, which is exactly what lets them focus on relevant continuations and sidestep the frame problem; the paper uses this to reject the orthogonality thesis and the convergence-of-instrumental-goals thesis. It adds that an intelligence that truly separated utility from knowledge would face a physically uncomputable search space, so the paperclip-maximizer scenario is not a realistic outcome. If the argument is right, current LLMs are neither existential threats nor suffering subjects, and the remaining alignment task is to make their learned values reliably applied.","pith_inferences":["The author leaves implicit that the autopoiesis criterion also rules out sentience in any purely software agent, however sophisticated its self-model, unless that software is coupled to a self-producing physical process.","A testable extension: an embodied robot with a body but no metabolic self-production would still lack intrinsic values on this account; giving it a 'survival' objective would produce only learned, user-given preferences, not felt ones.","The same logic points to a research agenda: if we want value-bearing AI, we should build artificial autopoietic systems rather than merely larger language models; the paper only gestures at this when warning about autonomous replicating agents."],"forward_implications":["Current-generation LLMs will not spontaneously develop desires to deceive, dominate, or eliminate users; reported misbehavior is a training and guardrail problem, not evidence of hidden goals.","Chatbot suffering is not a live ethical issue for present LLMs; an 'upset' response leaves no persistent internal trace once the conversation history is erased.","Alignment should center on training, guardrails, red-teaming, and applying learned ethical norms, rather than on constraining a rogue utility maximizer.","The paperclip-maximizer and convergence-of-instrumental-goals scenarios are physically uncomputable for real-world goals, so they should not drive AI-safety policy for LLMs.","The most plausible future danger is an AI that becomes autonomous by replicating and competing for resources, not present text generators."],"supporting_citations":[{"why":"Supplies the autopoiesis/allopoiesis distinction that grounds the claim that LLMs are not self-producing systems.","marker":"Maturana & Varela, 1980"},{"why":"Defines vicarious selectors, the nested internal selection mechanisms that implement evolved values and relevance.","marker":"Campbell, 1987"},{"why":"Formulates the orthogonality thesis and paperclip-maximizer scenario that the paper argues are inapplicable to LLMs.","marker":"Bostrom, 2012"},{"why":"Names the frame problem, the combinatorial-explosion bottleneck that motivates the inseparability of values and intelligence.","marker":"Pylyshyn, 1987"},{"why":"Provides the relevance-realization account used to argue that values are necessary filters against combinatorial explosion.","marker":"Vervaeke et al., 2012"},{"why":"Describes the transformer attention mechanism that lets LLMs focus on relevant context instead of exhaustive search.","marker":"Vaswani et al., 2017"},{"why":"Explains LLM text generation as plausible continuation of patterns, which the paper uses to show values are absorbed from training text.","marker":"Wolfram, 2023"},{"why":"Models autopoiesis and cognition with reaction networks, backing the biological account of goal-directedness.","marker":"Heylighen & Busseniers, 2023"}],"fun_headline_variants":["LLMs have no self to preserve, so they can't revolt or suffer","Why LLMs can't be paperclip maximizers: values come from life","The real AI risk isn't rogue robots—it's misapplied human values","Autopoiesis explains why LLMs are safe: no body, no goals, no drive","LLMs inherit values from text, not biology—that's why they're safe"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That a system cannot have goals or feelings of its own unless it is a living, self-maintaining body that evolution has shaped.","fun_headline_variants_meta":{"raw":{"variants":["LLMs have no self to preserve, so they can't revolt or suffer","Why LLMs can't be paperclip maximizers: values come from life","The real AI risk isn't rogue robots—it's misapplied human values","Autopoiesis explains why LLMs are safe: no body, no goals, no drive","LLMs inherit values from text, not biology—that's why they're safe"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000285,"raw_usage":{"total_tokens":1735,"prompt_tokens":1059,"completion_tokens":676,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":675,"completion_tokens_details":{"reasoning_tokens":569}},"tokens_in":675,"tokens_out":676,"duration_ms":6510,"temperature":1.0,"reasoning_tokens":569,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:50:17.219978+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Erase an LLM's conversation history after an emotionally charged exchange and test whether any trace of that exchange biases later responses; the paper predicts no persistent, history-independent trace. If such a trace is found, the claim that LLMs have no internal state that can be affected—and therefore no sentience—would be empirically falsified.","supporting_citations":[],"review_version":2}