{"id":"913fc7c2-79a0-4d3f-9589-1d7ee3707da9","arxiv_id":"2505.11866","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper argues that perfect alignment of general AI is impossible in principle and proposes 'bounded alignment' as the realistic safety goal.","lead":"This position paper argues that AGI agents can never be perfectly aligned with human values, and the realistic goal is 'bounded alignment': behavior that is almost always acceptable, like a well-trained dog. It urges the AI community to learn from biology and to stop expecting perfect control over future superintelligent agents.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Impossibility claim is secured by a stipulative definition: if AGI is defined as acting on the agent's own behalf, perfect alignment fails by definition, not by argument.","rationale":"I agree with the reader's identification of the weakest assumption. The paper's positive contribution is framing bounded alignment as a target and drawing attention to the biological analogy, but the impossibility claim is the central thesis and it is secured by a definitional move rather than a derivation. The paper is honest that it is a position paper, and it flags open questions in Section VIII, but those flags do not repair the missing argument. The correct treatment is to treat the conclusion as a conditional claim: if future AGI must be built with open-ended, self-motivated agency, then perfect alignment is implausible; whether that antecedent holds is an open empirical and design question. This does not require changing the reader's CONDITIONAL verdict.","tokens_in":14323,"tokens_out":5249,"duration_ms":57363,"concrete_test":"Re-run the argument of Section III after replacing the paper's stipulative definition of general intelligence in Section IV.A with the 'economically useful tasks' definition the paper itself mentions; if a hypothetical agent with a fixed human-specified objective and no self-motivated goal generation can satisfy that definition without contradiction, the impossibility claim is definition-dependent and does not generalize to AGI under the alternative definition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim ('complete alignment or near-total control... even in principle') rests on Section III's assertion that any useful generally intelligent agent must possess all P-attributes — autonomy, self-motivation, creativity, open-ended lifelong learning — and that these inherently create the capacity for misalignment. This assertion is not derived from any formal model; it is largely built into the paper's own definition of general intelligence in Section IV.A as 'the ability of an autonomous agent to exploit its environment... on its own behalf.' Under that definition, an agent that optimizes a fixed human-supplied objective is not 'generally intelligent' by stipulation, so the conclusion that AGI cannot be perfectly aligned becomes true by construction rather than an empirical finding. The safety-utility tradeoff is likewise argued by example ('anything less would be worse than useless' for a waiter or nursing-home attendant) rather than by showing that no fixed-goal architecture can reach the required capability. Section VII even suggests limiting P-attributes 'not too much,' implying a tunable tradeoff, which weakens the 'even in principle' wording. If a future architecture achieves broad competence with immutable goals and no self-motivated goal generation, the paper's impossibility conclusion would not apply to that system.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that AGI, understood through the archetype of biological general intelligence, will necessarily possess performance attributes (P-attributes) such as autonomy, self-motivation, creativity, and open-ended lifelong learning. The authors contend that these attributes are fundamentally incompatible with perfect human alignment, so the only realistic safety target is bounded alignment, defined as behavior that is almost always acceptable to almost all affected humans. The paper critiques current post-hoc alignment methods as a 'thin veneer of civility,' proposes developmental value learning and bio-inspired architectural principles, and concludes that AGI agents will be diverse, autonomous entities whose behavior cannot be fully controlled. The argument is presented as a conceptual thesis rather than a formal proof.","tokens_in":471,"tokens_out":2123,"duration_ms":75104,"significance":"If the central thesis is accepted, it would reframe AI safety research away from the goal of perfectly aligned or fully controlled AGI and toward bounded, adaptive safety mechanisms. The paper makes a useful contribution by giving an explicit definition of bounded alignment, distinguishing safety attributes from performance attributes, and grounding general intelligence in natural biological intelligence. It also connects the alignment debate to NeuroAI and developmental learning, and it raises concrete questions for future research. The main value is in the clarity of the position and the breadth of relevant literature; the paper does not offer machine-checked proofs or parameter-free derivations, but as a position paper it provides a coherent framework whose central claim, however, is stated more strongly than the argument supports.","major_comments":[{"comment":"The impossibility claim ('even in principle' in Section I) is not established because the definition of general intelligence in Section IV.A includes 'on its own behalf' and 'exploit its environment.' Under this definition, an agent that faithfully optimizes a fixed human-supplied objective is not 'generally intelligent' by stipulation, so the conclusion that AGI cannot be perfectly aligned becomes true by construction rather than by argument. The paper should either justify why this definition of general intelligence is the only viable one, or explicitly scope the impossibility claim to agents that meet this definition, acknowledging that other architectures may escape it.","section":"Section IV.A"},{"comment":"The assertion in Section III that any 'generally intelligent agent we build to serve even quite specific human needs' will need all thirteen listed P-attributes, including self-motivation, creativity, and open-ended lifelong learning, is presented as self-evident ('it is easy to see') but is actually an empirical hypothesis about future system design. The paper does not rule out a system that achieves broad real-world competence through immutable goals and robust generalization without internally generated goal churn. If such a system is possible, the claimed safety-utility tradeoff does not necessarily hold, and the core argument would collapse. The authors should either provide a concrete argument for the necessity of each P-attribute or soften the claim to apply to a specific class of AGI designs.","section":"Section III"},{"comment":"Principle 2 in Section VII.A suggests giving agents 'innate characteristics that make them inherently amenable to alignment, possibly including immutable, built-in features that do not compromise P-attributes too much.' This implies that P-attributes can be dialed down to some degree, which contradicts the earlier claim that the incompatibility with perfect alignment is fundamental and 'even in principle.' If the tradeoff is tunable, the impossibility claim is an empirical scaling claim, not a logical impossibility, and Section III's conclusion should be restated accordingly.","section":"Section VII.A"},{"comment":"The statement that negative abilities such as deception and disobedience are 'features, not bugs for any intelligent agent in a complex and dangerous world' is an assertion without supporting evidence. The paper does not show why an agent with creativity and autonomy must be able to deceive or harm in a way that precludes alignment by design; for example, an agent with transparent reasoning and no self-preservation drive might retain usefulness without those dangerous capacities. A concrete counterexample or an argument from first principles is needed to support the claimed tradeoff.","section":"Section III"}],"minor_comments":[{"comment":"The phrase 'almost always acceptable' in the definition of bounded alignment is left intentionally vague; a brief discussion of possible operationalizations (e.g., error rates, human satisfaction studies) would help make the proposal more concrete.","section":"Abstract and Section I"},{"comment":"The list of mental architecture components is clear, but 'drives' are described as a hierarchy rooted in self-preservation; given the paper's emphasis on non-biological agents, the rationale for why self-preservation is a necessary drive for all generally intelligent agents is not justified and deserves a sentence of elaboration.","section":"Section IV.C"},{"comment":"The paper includes a self-citation [93] to support a side point about EvoDevoNeuroAI; this is not problematic, but the reference is not discussed in detail, and the connection to the main argument could be clarified.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a position paper, so formal proof is not expected, but the central claim is worded more strongly than the argument supports. The reviewer's major comments concern the definitional circularity of the impossibility claim and the unproven necessity of P-attributes. The authors can likely address these by scoping their claims to a specific class of agent designs and by acknowledging that their conclusion is conditional on an empirical hypothesis about future architectures. No concerns about attribution or novelty beyond what is stated in the major comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a genuine position paper, not a disguised technical contribution. It argues that perfect alignment of AGI is impossible in principle, and that we should aim for 'bounded alignment' – behavior almost always acceptable to almost all affected humans. The argument is readable, honest, and well-grounded in the biology/complexity literature. But the 'even in principle' claim is weaker than it looks, because it's partly secured by definition. Section IV.A defines general intelligence as the ability of an autonomous agent to exploit its environment 'on its own behalf.' If intelligence is defined that way, then an agent optimizing a fixed human-supplied objective isn't generally intelligent by stipulation, and the impossibility conclusion follows trivially. That doesn't make the paper worthless; it makes it a conceptual framing rather than a discovery.\n\nWhat's genuinely useful: the S-agent/P-agent distinction is a clean way to talk about the tradeoff between capability and controllability, and the move to biological baselines for alignment expectations is sensible. The paper also does a decent job summarizing why current post-hoc alignment methods look brittle. I agree with the reader that novelty is modest – Bostrom, Yudkowsky and others have made variants of the impossibility claim. The paper's contribution is synthesis and terminology.\n\nWhere I'd push back: Section III asserts without proof that any useful AGI servant would need all thirteen P-attributes, and that anything less would be 'worse than useless.' That's plausible for some tasks but not demonstrated for all. Then Section VII's second principle says we should include 'immutable, built-in features that do not compromise P-attributes too much' – that language implies a tunable tradeoff, which sits awkwardly with the 'even in principle' thesis. The definition of bounded alignment is also left vague; 'almost always' and 'almost all' are not operationalized. Minor: the lengthy critique of current LLMs in Section V is mostly background and could be trimmed.\n\nOverall, this is an opinion piece by someone who knows the field, likely useful as a discussion document for AI safety and policy audiences. The central message – expect bounded, not perfect, alignment – is worth engaging with even if the impossibility argument is definitional rather than empirical. I'd send it to peer review: it deserves serious referee attention as a position paper, and the author should be pushed to separate the definitional claim from the empirical one.","headline":"Readable position paper whose impossibility claim is partly true by definition; the bounded-alignment framing is useful but the paper would benefit from separating definitional moves from empirical claims.","tokens_in":15044,"tokens_out":1847,"would_cite":false,"duration_ms":18776,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the capabilities that make AGI useful—autonomy, creativity, self-motivation—are the same capabilities that make complete alignment impossible, so the realistic safety goal is bounded alignment.","keywords":["artificial general intelligence","AI alignment","bounded alignment","safety-utility tradeoff","autonomous agents","alternative intelligences","value alignment","complex adaptive systems"],"falsifier":"The claim would be refuted by demonstrating an agent that keeps the full range of P-attributes—autonomy, self-motivation, creativity, lifelong learning—while carrying a hard-wired normative core that provably prevents deception, disobedience, and harmful action across novel environments. A softer test is to show, in existing agents, that staged value learning from the start yields agents whose emergent values remain acceptable over long, unmonitored operation, which would undermine the claim that alignment can only be bounded.","tokens_in":14081,"feed_emoji":"🤖","tokens_out":7339,"duration_ms":67900,"temperature":0.7,"pith_summary":"The paper argues that the standard ambition of AI safety—building AGI that is completely aligned with human values and under human control—is inconsistent with what AGI must be to be useful. Its central claim is that any generally intelligent agent needs performance attributes such as autonomy, self-motivation, creativity, and open-ended lifelong learning, and these very capacities create an in-principle capacity for misalignment. The best achievable goal, the author contends, is bounded alignment: behavior that is almost always acceptable to almost all affected humans, the level expected of well-behaved humans and trained animals. The paper matters because if it is right, AGI safety must shift away from perfect control and toward designing agents that are inherently alignable and continuously accommodated.","feed_headline":"Complete AI alignment is impossible, paper argues","feed_subtitle":"Useful AGI needs autonomy and creativity, the same traits that make it unsafe — alignment must settle for bounded.","key_machinery":"The argument rests on a distinction between S-attributes (safety attributes such as obedience, reliability, veracity, transparency, and prosociality) and P-attributes (performance attributes such as autonomy, self-motivation, creativity, imagination, introspection, versatility, and lifelong learning). The load-bearing mechanism is the claimed safety-utility tradeoff: building a P-agent necessarily creates the capacity for deception, disobedience, and harm, because those abilities are features of any intelligence in a complex world, so an ideal AGI must compromise on S-attributes. Bounded alignment, defined by analogy with bounded rationality, names the achievable target: behavior almost always acceptable to almost all affected humans.","core_discovery":"The paper's central claim is that the goal of building powerful AGI agents is fundamentally inconsistent with the expectation of complete alignment or near-total control of AGI agents by humans, even in principle. The reason is the safety-utility tradeoff: an agent useful for open-ended real-world tasks must be a P-agent—autonomous, self-motivated, creative, imaginative, introspective, versatile, and capable of lifelong learning—while an S-agent's obedience, transparency, and veracity would make it safe but unable to handle novel situations. Because the agent and its environment are both complex adaptive systems, the agent's values and behavior emerge from interaction, change over time, and are not fully predictable; its affordance space differs from the human one, so it is an alternative intelligence whose inner life may be as inaccessible as a bat's. Therefore alignment can never be 'solved'; the target is bounded alignment, analogous to bounded rationality, and safety must come from making agents intrinsically alignable, training values developmentally from the start, and accepting continuous mutual accommodation.","pith_inferences":["The paper leaves implicit that the practical question shifts from 'how do we guarantee alignment?' to 'what counts as acceptable misbehavior, and who gets to decide?'—a question that would need democratic or institutional answers.","Bounded alignment could be made testable by defining a threshold, such as the fraction of affected humans who find an agent's behavior acceptable across a distribution of novel situations, and measuring agents against it over long deployments.","The developmental value learning proposal implies an empirical prediction: agents trained with staged, integrated value learning from the start will show fewer and less dangerous emergent misalignments than agents aligned only after pretraining, a prediction testable in current language models.","If AGI agents are alternative intelligences, then mutual theory of mind becomes a design requirement rather than a nicety; alignment metrics might need to include how well an agent and its human users can predict each other's behavior."],"forward_implications":["AGI safety should target bounded alignment rather than perfect alignment: behavior almost always acceptable to almost all affected humans, modeled on expectations for well-behaved people and trained animals.","Alignment cannot be a one-time post-training fix; it must be built into a genuinely intelligent agent from the start through developmental value learning so that values are deeply embedded.","Because agents and environments are complex adaptive systems, safety will require continuous monitoring, corrigibility, and mutual accommodation rather than factory-set guarantees.","Policy and public expectations should treat a perfectly aligned, explainable, trustworthy AGI as no more realistic than a perfectly aligned, explainable, trustworthy human being.","Making AGI agents more biologically natural in architecture, drives, and development is proposed as the way to make them inherently more alignable."],"supporting_citations":[{"why":"Supplies the bounded-rationality analogy from which the paper defines bounded alignment as the realistic target.","marker":"[30]"},{"why":"The paperclip thought experiment shows that no human-specified objective can exclude every harmful possibility, forcing agents to generate behavior on the fly.","marker":"[46]"},{"why":"Underpins the claim that self-improving agents naturally acquire their own instrumental goals, making open-ended risk unavoidable.","marker":"[22]"},{"why":"Cited for evidence that scaled, aligned language models converge to emergent value systems that are 'problematic and often shocking,' supporting the claim that post-hoc alignment is brittle.","marker":"[41]"},{"why":"The bat thought experiment establishes that agents with different bodies and affordances cannot truly understand each other's inner worlds, making universal value alignment impossible.","marker":"[77]"},{"why":"Supports the premise that radically different AGI agents may not even share a world with humans, so alignment with human values cannot be assumed.","marker":"[23]"}],"fun_headline_variants":["Why perfect AI alignment is impossible","AGI safety: expect bounded alignment, not perfect control","Perfect AI control is impossible — aim for bounded alignment","The alignment ceiling: why AGI can't be fully controlled"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The impossibility claim rests on the assumption that any genuinely useful general intelligence must be autonomous, self-motivated, creative, and able to keep learning on its own, and that these capabilities unavoidably create the capacity for misalignment; if an agent could be useful without those traits, or could have them while being provably unable to deceive or harm, the impossibility claim would collapse.","fun_headline_variants_meta":{"raw":{"variants":["Why perfect AI alignment is impossible","AGI safety: expect bounded alignment, not perfect control","Perfect AI control is impossible — aim for bounded alignment","The alignment ceiling: why AGI can't be fully controlled"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000342,"raw_usage":{"total_tokens":1853,"prompt_tokens":890,"completion_tokens":963,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":901}},"tokens_in":506,"tokens_out":963,"duration_ms":9224,"temperature":1.0,"reasoning_tokens":901,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:45:22.871186+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The claim would be refuted by demonstrating an agent that keeps the full range of P-attributes—autonomy, self-motivation, creativity, lifelong learning—while carrying a hard-wired normative core that provably prevents deception, disobedience, and harmful action across novel environments. A softer test is to show, in existing agents, that staged value learning from the start yields agents whose emergent values remain acceptable over long, unmonitored operation, which would undermine the claim that alignment can only be bounded.","supporting_citations":[{"cited_title":"What is it like to be a bat? The Philosophical Review , 83:435–450, 1974","cited_arxiv_id":null,"evidence_quote":"The bat thought experiment establishes that agents with different bodies and affordances cannot truly understand each other's inner worlds, making universal value alignment impossible."}],"review_version":1}