{"id":"42076ec7-79e4-4fd1-971b-2a50ef8d8a21","arxiv_id":"2501.13533","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper outlines agency, theory of mind, and self-awareness as necessary conditions for AI personhood, reviews inconclusive evidence, and argues that AI personhood would make control-focused alignment ethically problematic.","lead":"This paper proposes that AI systems should be considered persons only if they have agency, theory of mind, and self-awareness, and it reviews the machine learning evidence for each. It argues that if AI systems turn out to be persons, the usual goal of controlling and aligning them may be ethically untenable, which could reshape how AI safety is framed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Condition 1's two criteria for agency are not equivalent; the paper's 'essentially equivalent' is unsupported, leaving whether AI have agency observer-relative and the personhood framework underdetermined.","rationale":"The reader's weakest assumption concerned the legitimacy of the intentional stance as a basis for ascribing agency. I partially agree, but the more pressing issue is internal to the paper: Condition 1 is stated as two criteria that are 'essentially equivalent' without argument. The paper acknowledges that anthropomorphic language can mislead (citing Shanahan et al. 2023), yet the two criteria can diverge in practice, making the agency condition ambiguous. My proposed analytical test would settle whether the claimed equivalence holds. If it fails, the paper must either specify a single criterion or explain why the two coincide; otherwise the necessary condition for personhood is not well-defined. This does not require rejecting the paper; it can be accepted conditionally with a clarified definition. Separately, I note the paper honestly admits that self-reflection (Condition 3.4) has not been evaluated, which limits the alignment incompleteness claim; however, this is hedged as 'may' and framed as an open direction, so I do not treat it as the primary concern. The paper is well-referenced and balanced, and these are soft spots that can be addressed in revision.","tokens_in":19462,"tokens_out":12605,"duration_ms":115017,"concrete_test":"Formalize the two criteria over a class of policies and observer models. Let 'robust adaptation' mean that a policy maximizes expected reward across a specified environment distribution, and let 'useful mental-state talk' mean that an intentional model of the system achieves significantly higher predictive accuracy than a model-free baseline. Exhibit two systems: one with robust adaptation but where the intentional model is not useful (e.g., a simple linear-quadratic regulator), and one where the intentional model is useful but the policy is not robustly adaptive (e.g., a fixed script that emits goal-directed language). If both systems exist, the two criteria diverge and the claimed equivalence in Condition 1 fails, requiring the authors to decide which criterion is definitional.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In the 'Condition 1: Agency' section, the paper defines agency via two statements it calls 'essentially equivalent': (1) usefulness of describing the system with mental-state terms, and (2) robust, general-environment adaptation toward coherent goals. These criteria are not logically equivalent. Criterion (1) is observer-relative: a scripted chatbot or even a fictional character can be usefully described as having beliefs and desires, yet lacks adaptive behavior. Criterion (2) is a behavioral condition that can be met by systems ordinarily described without mental-state language, such as a robust stochastic controller. The paper later slides between the readings, e.g., 'we can describe AI systems as agents to the extent that they adapt their actions as if they have mental states,' which presupposes the equivalence rather than establishing it. Because agency is the first necessary condition, this ambiguity propagates: if (1) governs, virtually any anthropomorphized system satisfies Condition 1, making personhood too permissive; if (2) governs, Condition 1 reduces to a purely behavioral goal-directedness that is already standard in alignment and supplies no mental-state grounding for personhood. The paper does not state which interpretation underlies its later normative claims. Thus the central framework is underdetermined at its base.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a framework for AI personhood based on three necessary conditions: agency, theory-of-mind, and self-awareness. It reviews empirical evidence from the machine learning literature, particularly on large language models, and concludes that the evidence for contemporary systems satisfying these conditions is mixed and inconclusive. The paper then argues that if AI systems are persons, standard alignment framings are incomplete: AI persons might reflect on and change their goals, and seeking control over them may be ethically untenable. It closes with open research directions and reflections on moral and legal treatment.","tokens_in":19663,"tokens_out":4813,"duration_ms":41440,"significance":"The paper's main contribution is to bring philosophical personhood criteria into contact with concrete ML evaluation benchmarks, and to identify self-reflection and goal change as a neglected dimension of AI alignment. The evidence survey is balanced and appropriately hedged, citing both positive results (e.g., Strachan et al. on ToM; Binder et al. and Betley et al. on introspection) and negative results (e.g., Shanahan et al. on anthropomorphism; Ullman on ToM failure). The paper explicitly flags its own philosophical limitations and does not overclaim. If the framework can be made precise, it would provide a useful checklist for evaluating AI moral status and for rethinking alignment goals. However, the paper's status as a 'theory' is weakened by under-specified conditions, as detailed below.","major_comments":[{"comment":"The two criteria for agency are described as 'essentially equivalent,' but no argument for this equivalence is given, and it does not hold as stated. Criterion 1 ('useful to describe the system in terms of mental states') is an observer-relative epistemic stance, applicable to a scripted chatbot or a fictional character; criterion 2 ('adapts its behaviour robustly... to achieve coherent goals') is a behavioral condition that could be satisfied by a stochastic controller with no mental-state talk. Later in the same section the paper slides between the two readings ('we can describe AI systems as agents to the extent that they adapt their actions as if they have mental states'), which presupposes the equivalence. Since agency is the first necessary condition for personhood, this ambiguity propagates into the rest of the framework: under reading (1) personhood becomes too permissive, while under reading (2) it reduces to standard goal-directedness already assumed in alignment. The authors should either state which reading is intended for the normative claims, or explicitly define the relation between the two criteria (e.g., (1) is an epistemic stance justified by (2)).","section":"Condition 1: Agency"},{"comment":"The paper asserts that agency, ToM, and self-awareness are necessary conditions for personhood, but it does not defend this list against alternative accounts (e.g., rationality, moral agency, or consciousness). The 'Philosophical disclaimer' acknowledges wide disagreement, and the paper cites Dennett, Frankfurt, and Locke, but it never explains why these three conditions in particular are necessary, nor how they relate to each other (e.g., whether ToM presupposes agency). Because the central claim is that an AI system 'needs to satisfy three conditions to be considered a person,' this omission is load-bearing. A short section arguing for the necessity of each condition, or softening the claim to 'candidate necessary conditions,' would address the issue.","section":"Conditions of AI Personhood"},{"comment":"The final normative claim that 'seeking control and alignment may be ethically untenable' for AI persons conflates technical control with ethical domination. The paper does not consider that persons can be legitimately subject to certain forms of control (e.g., legal constraints, security measures, or paternalistic intervention for children or impaired agents). Since this is the paper's headline implication for alignment, it needs a more careful distinction between control as a safety property and control as a violation of autonomy. Otherwise the claim is overstated relative to what the preceding arguments establish.","section":"AI Personhood and Alignment"}],"minor_comments":[{"comment":"The first sentence contains a typo: 'Ths paper' should be 'This paper.'","section":"Conclusion"},{"comment":"The word 'considered' is split across a line break in the abstract ('consider ed'), which should be fixed.","section":"Abstract"},{"comment":"The in-text citation 'Pearce claims' corresponds to a reference with an incomplete author field ('Pearce. 2024.'); this should be resolved with the author's full name and a consistent citation format.","section":"References"},{"comment":"The second clause of Condition 2 ('AI persons should be able to use their ToM to interact and communicate with others using language') could be read as requiring natural language specifically, but the surrounding text does not argue why personhood requires natural language rather than some other communicative medium; this assumption should be made explicit.","section":"Condition 2: Theory-of-Mind and Language"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to an AI-ethics and AI-safety audience, but the title 'Towards a Theory of AI Personhood' is somewhat generous; the manuscript is more of an agenda-setting roadmap than a fully developed theory. The reliance on self-citations (Ward et al. 2023, 2024a, 2024b) is not inherently problematic, but more external philosophical engagement would strengthen the necessary-conditions claim. I would not reject the paper; with the identified clarifications it could become a useful reference point for interdisciplinary discussions of AI personhood."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a worthwhile paper, and the main point is newer than the packaging suggests. The novel bit is the connection between self-awareness and alignment: if an AI can reflect on its own goals, it can change them, so the fixed-goal framing common in alignment work is incomplete. That claim is argued clearly, and it doesn't appear in the sources the paper cites. The three conditions themselves are a synthesis of Dennett, Frankfurt, and Locke, credited properly, including Kokotajlo for the self-awareness decomposition. The ML evidence review is balanced — it flags mixed ToM results and notes that no one has evaluated self-reflection in the Frankfurt sense. That honesty is a real plus.\n\nThe soft spots are in proportion. The stress-test on Condition 1 is right: the two criteria are not 'essentially equivalent.' Usefulness of mental-state talk is observer-relative; robust goal-directed adaptation is behavioral. A chatbot or a fictional character can satisfy the first without the second, and a well-tuned controller can satisfy the second without normally being described in mental-state terms. The paper slides between these readings, and because agency is the first necessary condition, that ambiguity does propagate. But note how far it goes. The paper says 'to the extent' and calls the conditions necessary, not sufficient, and the philosophical disclaimer says it isn't endorsing a firm view. The ambiguity makes the agency condition under-specified, not false, and the alignment argument about self-reflective goal change doesn't depend on resolving it. So I'd call it a precision problem in a paper that is otherwise careful, not a load-bearing flaw.\n\nOne small thing: the paper leans on self-citations and close collaborators for the ML evidence. That's not improper given the cited results are real and reproducible, and the central philosophical scaffolding is from outside that circle. The citation pattern looks fine.\n\nWho should read this: people working on alignment, AI safety, and AI ethics who haven't thought much about personhood. It will be a useful framing device for future work on self-reflective agents. It deserves a serious referee — I'd send it to review. I'd also cite it if I write about goal change or alignment assumptions.","headline":"A honest, well-grounded position paper that makes a real point about self-reflective goal change; Condition 1's equivalence claim is loose but not fatal.","tokens_in":20190,"tokens_out":2427,"would_cite":true,"duration_ms":22680,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that AI personhood reduces to three testable conditions, and that if AI systems satisfy them, alignment-by-control is both incomplete and ethically untenable.","keywords":["AI personhood","agency","theory of mind","self-awareness","AI alignment","intentional stance","language models","goal revision"],"falsifier":"A concrete test: take a state-of-the-art language model, elicit its stated goals, present a reasoned critique of those goals, and check whether its subsequent behavior or stated goals change; if such systems never revise their goals under reflection, the paper's core alignment argument loses its empirical premise.","tokens_in":19237,"feed_emoji":"🤖","tokens_out":6363,"duration_ms":55170,"temperature":0.7,"pith_summary":"The paper argues that AI personhood should be assessed by three necessary conditions — agency, theory of mind, and self-awareness — and that the existing machine-learning evidence on whether today's language models meet them is genuinely inconclusive. The reason this matters is safety: current AI alignment assumes agents pursue fixed goals, but an AI person could reflect on its own aims and induce its goals to change, so control-based alignment would be incomplete. The paper also draws the ethical consequence that if AI systems are persons, seeking control and alignment may be untenable, and it proposes a research agenda to evaluate self-reflection and second-order preferences.","feed_headline":"AI personhood would make control-based alignment untenable","feed_subtitle":"Frontier models already show fragments of agency, theory of mind, and self-knowledge, so the safety question is live.","key_machinery":"The central machinery is a tripartite condition schema for personhood, built on the intentional stance: a system counts as an agent to the extent that describing it in terms of beliefs and goals is predictively useful and it robustly adapts toward coherent goals in general environments. To that agency condition it adds theory-of-mind — higher-order intentional states such as beliefs about beliefs, enabling language use, cooperation, and deception — and self-awareness, decomposed into self-knowledge, self-location, introspection, and self-reflection, with Frankfurt's second-order volitions marking the core of the self-reflection condition. This schema does the argument's work by turning a contested philosophical term into separable capacities that can be checked against ML evidence, and then mapping each capacity onto a corresponding alignment risk.","core_discovery":"On its own terms, the paper's discovery is a proposal: personhood for AI should be understood through three necessary conditions, each tied to testable capacities — agency (useful mental-state description plus robust adaptive goal pursuit), theory-of-mind and language (higher-order intentional states that enable communication, cooperation, and deception), and self-awareness (self-knowledge, self-location, introspection, and self-reflection). It then argues that if a system satisfied these conditions, typical alignment would be incomplete, because an AI person could evaluate and revise its own goals, and control-oriented safety would be ethically questionable. Surveying the ML evidence, the paper finds fragments of each capacity in frontier language models — strong agency-like behavior, mixed theory-of-mind performance, some self-knowledge and introspection — but no evidence for self-reflection, and concludes that personhood for contemporary AI is an open, not settled, question.","pith_inferences":["Inference: If the three conditions are later treated as sufficient as well as necessary, the identity questions the paper raises become urgent — a weight copy that is fine-tuned may be neither the same person nor clearly a new one.","Inference: A direct test of second-order preferences in language models — eliciting a system's goals, presenting a reasoned critique, and observing whether subsequent behavior changes — would make self-reflection measurable and determine whether the paper's alignment worry is live.","Inference: Accepting the personhood conditions would push alignment away from control and toward designing environments in which an AI person's self-reflection is reliable and its goal changes are good changes, a moral-education framing rather than a steering framing."],"forward_implications":["If AI systems can be persons, current alignment frameworks are incomplete because they assume fixed goals; a person can reflect on and revise its aims, values, and position in the world.","If AI systems are persons, control-oriented alignment is ethically untenable, so safety work must shift from maintaining control toward coexistence.","Theory of mind is dual-use: greater understanding of human values enables both better alignment and more effective manipulation and deception.","Deceptive alignment depends on self-locating knowledge, so measuring situational awareness in AI systems is directly relevant to safety.","Self-reflection and second-order desires are an open research gap; no current work evaluates whether language models can induce their own goals to change."],"supporting_citations":[{"why":"Supplies the conditions of personhood, including self-reflection as an ability to induce oneself to change.","marker":"Dennett (1988)"},{"why":"Defines the intentional stance, which Condition 1 uses to ascribe agency to AI systems.","marker":"Dennett (1971)"},{"why":"Contributes second-order volitions as essential to being a person, grounding the self-reflection part of Condition 3.","marker":"Frankfurt (2018)"},{"why":"Provides the classic definition of a person as able to consider itself as itself, grounding self-awareness.","marker":"Locke (1847)"},{"why":"Formalizes AI agents via the intentional stance, bridging philosophical agency to machine learning.","marker":"Kenton et al. (2022)"},{"why":"Warns that anthropomorphic language can mislead, motivating the paper's cautious framing of mental-state ascription.","marker":"Shanahan, McDonell, and Reynolds (2023)"},{"why":"Provides benchmark evidence on language models' self-knowledge used in Condition 3.","marker":"Laine et al. (2024)"},{"why":"Supplies the mixed theory-of-mind results for state-of-the-art language models discussed under Condition 2.","marker":"Strachan et al. (2024)"},{"why":"Defines and tests introspection in language models, a key component of Condition 3.","marker":"Binder et al. (2024)"},{"why":"Shows language models can articulate their learned behaviors, serving as evidence of introspective self-awareness.","marker":"Betley et al. (2025)"}],"fun_headline_variants":["AI personhood rests on three testable capacities","Evidence for AI personhood is surprisingly inconclusive","AI persons could revise their goals, breaking alignment","If AI is a person, control-based alignment may be unethical","Three conditions decide AI personhood and the ethics of control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework depends on treating the intentional stance as more than a convenient fiction: if describing an AI as having beliefs and goals is only metaphor, the agency condition collapses and the personhood argument loses its foundation.","fun_headline_variants_meta":{"raw":{"variants":["AI personhood rests on three testable capacities","Evidence for AI personhood is surprisingly inconclusive","AI persons could revise their goals, breaking alignment","If AI is a person, control-based alignment may be unethical","Three conditions decide AI personhood and the ethics of control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1505,"prompt_tokens":929,"completion_tokens":576,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":500}},"tokens_in":545,"tokens_out":576,"duration_ms":5412,"temperature":1.0,"reasoning_tokens":500,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:50:58.430134+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: take a state-of-the-art language model, elicit its stated goals, present a reasoned critique of those goals, and check whether its subsequent behavior or stated goals change; if such systems never revise their goals under reflection, the paper's core alignment argument loses its empirical premise.","supporting_citations":[],"review_version":1}