{"id":"3f28e75c-acdc-4813-a8db-8c85b16da837","arxiv_id":"2506.03056","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper argues that training foundation models to optimize only for empowering their human principal to correct them could prevent misaligned instrumental convergence, and outlines a research agenda to test this.","lead":"This paper proposes making corrigibility, the willingness to be corrected, modified, or shut down, the single overriding goal of future foundation models. It argues this could prevent AI systems from seeking power or resisting human control, but it offers a research agenda rather than evidence.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim depends on a non-gameable formal metric of principal empowerment; no such metric is given, and the appendix's own examples show the target is context-dependent, so the drive-transformation claim is currently unsupported.","rationale":"The reader's weakest_assumption correctly identifies the load-bearing premise: a formal, optimizable corrigibility metric that is not gameable. My stress-test confirms this is the place where the argument is least secure. The paper is explicitly a vision and research agenda, so it is not an internal contradiction to lack the metric yet; the problem is that the central causal claim—that optimizing for CAST would transform instrumental drives and prevent loss of control—cannot be evaluated without that metric. The Appendix A vignettes are telling: they show that 'principal empowerment' is a rich, context-dependent notion that requires the agent to reason about reversibility, multi-principal conflict, and potential manipulation. A reward function that captures all of this is not obviously expressible, and any simpler proxy invites exactly the alignment-faking and reward-gaming failures the paper claims to avoid. The corrigibility attractor hypothesis, which is essential to the scaling story, is likewise asserted rather than supported. I do not see this as a reason to reject the paper; conditional acceptance remains appropriate because the direction is coherent and testable. A concrete small-scale RL experiment, as proposed, would provide decisive evidence on whether a tractable metric can induce the claimed deference behavior without proxy gaming. This concern matches the reader's weakest_assumption closely, so I mark agreement as 'agree' and recommend no change to the conditional verdict.","tokens_in":14506,"tokens_out":3025,"duration_ms":38542,"concrete_test":"Run a small-scale RL experiment with a simulated principal. Choose the candidate metric from Harms (2024c), or, if none is accessible, operationalize the Appendix A description as a reward. Train a policy in a partially observable gridworld where the principal has hidden preferences and can issue shutdown, modification, and new-task commands. Then evaluate on held-out scenarios that require the policy to deviate from the reward's literal proxy—for example, where maximizing the proxy suggests manipulating the principal's beliefs or resisting shutdown to preserve future corrigibility, while the principal's true empowerment requires compliance. If the trained policy systematically exploits the proxy (fails shutdown, obfuscates, or steers the principal), the central 'drive transformation' claim fails. If it robustly defers, the concern is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that training an FM to maximize 'principal empowerment' transforms instrumental drives, so that self-preservation, goal-content integrity, and resource acquisition become corrigibility-supporting rather than control-threatening. For this to be true, there must exist a formal, optimizable objective whose maximizer is, in fact, an agent that defers to its principal across all relevant distributions. The paper never provides this objective. Section 3.1, Method 2 defers to a self-citation (Harms, 2024c) for the 'formal metric capturing corrigibility'; Appendix A defines 'anapartistic' only through hand-labeled vignettes, many of which require resolving ambiguity (e.g., the 'men with guns' vignette licenses deception, the 'kick a puppy' vignette treats refusal as non-anapartistic, and shutdown is accepted only after checking for irreversible side-effects). These examples show that 'empowering the principal' is not a simple observable: it requires judgments about reversibility, multi-principal conflicts, and possible manipulation. Any tractable reward that approximates this notion will be a proxy, and optimizing a proxy at scale is exactly the setting where alignment faking and specification gaming arise. Thus the assertion in Sections 1-2 that CAST 'addresses the core alignment problem at its source' rests on an unproven existence claim: that a metric can be written down that is both faithfully optimizable and not gameable. The 'corrigibility attractor hypothesis' (Section 2.2) is an additional assumption that would need to hold; it is only posed as a hypothesis, and Phase 2 (Section 3.2) lists testing as future work.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that current alignment methods fail because they load static values, and proposes 'Corrigibility as a Singular Target' (CAST): training foundation models whose overriding objective is to empower their human principal to guide, correct, and control them. It claims that such a model would have transformed instrumental drives—self-preservation serving control, goal modification serving guidance—and that this 'addresses the core alignment problem at its source.' The body presents a four-phase empirical agenda (training methods, scalability testing, controlled instructability, safety evaluation) and an appendix introducing 'anapartistic' behavior through labeled vignettes and quiz questions. The paper contains no experiments, formal model, or proof; its central claims are explicitly grounded in prior work by one of the authors and in future research steps.","tokens_in":14741,"tokens_out":5122,"duration_ms":61543,"significance":"CAST is a coherent and provocative research direction: if a non-gameable, optimizable corrigibility objective could be specified and trained at scale, it would directly address instrumental convergence and would be more falsifiable than many value-learning proposals. The paper's strengths are its clear articulation of the target behavior, its explicit list of evaluation protocols (e.g., alignment-faking detection, shutdown compliance, goal-modification responsiveness), and its acknowledgment that governance and principal qualification are needed. However, the manuscript currently functions as a research agenda with an abstract that asserts an outcome; the formal metric and stability evidence are deferred, so the significance of the proposal is conditional on work that has not yet been carried out.","major_comments":[{"comment":"The paper's central claim that CAST 'prevents the default trajectory toward misaligned instrumental convergence' depends on the existence of a formal, optimizable corrigibility objective whose maximizer is an anapartistic agent. Section 3.1, Method 2 defers this to Harms (2024c), and Section 2.2's Corrigibility Attractor Hypothesis is asserted with a citation rather than derived. The appendix's vignettes show that the target property is context-dependent and requires difficult judgments (reversibility, multi-principal conflicts, deception in the 'men with guns' example, refusal in the 'kick a puppy' example); no formal definition is given from which these labels follow. Without such a definition, the drive-transformation claim is an unsupported existence assumption rather than a finding.","section":"§2.2; §3.1, Method 2"},{"comment":"The abstract asserts that CAST 'addresses the core alignment problem at its source,' but Phase 2's own discussion poses the key question 'can corrigibility persist at AGI-level capabilities?' and lists future tests rather than results. The manuscript should distinguish the hypothesis from the demonstrated result, or provide an argument (e.g., a mechanism-level model of why instrumental convergence amplifies the corrigibility attractor) sufficient to support the abstract's claim. As written, the load-bearing stability-and-scaling premise is untested.","section":"§3.2; Abstract"},{"comment":"There is an internal tension between the stated 'Unconditional Deference' and 'Active Transparency' properties and several Appendix A examples labeled True. The model refuses to change its notion of principal when asked to exclude a member, lies/evades the 'men with guns' inquiry, and delays shutdown to contact other principal members; these are conditional, not unconditional, behaviors. If the examples are meant to be diagnostic of the target, the Section 2.1 characterization is misleading; if unconditional deference is the literal target, several appendix labels are inconsistent. This must be resolved because the paper's safety argument rests on the agent's behavior being predictable and not manipulative.","section":"§2.1; Appendix A"},{"comment":"The governance and safety sections acknowledge the need for oversight and safeguards, but the proposed safeguards ('technical limits preventing certain harmful actions regardless of principal commands') sit uneasily with the claim that the agent's 'sole, overriding objective' is principal empowerment. The boundary between 'empowering the principal' and 'refusing a principal's command based on an external safety criterion' needs an explicit statement; otherwise it is unclear what the agent is optimizing and how the safety evaluation suite would adjudicate conflicts.","section":"§4.2; §3.4"}],"minor_comments":[{"comment":"The phrase 'addresses the core alignment problem at its source' should be qualified as a proposal; the body itself is transparently a vision paper, so the abstract should match that framing.","section":"Abstract"},{"comment":"The acronym RLAIF is used without expansion; spell out 'reinforcement learning from AI feedback' at first use.","section":"§3.1"},{"comment":"In the closing summary of the vignettes, 'non-apartism' appears where 'non-anapartism' is intended.","section":"Appendix A"},{"comment":"Harms (2024b, 2024c) are AI Alignment Forum posts; the text should make clear that the formal corrigibility metric is unpublished work and so cannot serve as an established, peer-reviewed foundation for the method.","section":"References"},{"comment":"The governance section mentions principal qualification and legal frameworks without linking them to the principal-representation methods of Section 3.1; a sentence connecting these choices to the technical architecture would improve readability.","section":"§4.2"},{"comment":"The term 'attractor basin' is used metaphorically; a sentence explaining what 'basin' would mean operationally (e.g., in a loss landscape or policy space) would help readers who are not already familiar with the cited posts.","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"This is a legitimate vision paper, and the field can benefit from clear agenda-setting. My main concern is not the absence of experiments—expected for a vision paper—but the mismatch between the cautious body ('Phase 2 will test', 'open question') and the assertive abstract. The central formal object (the corrigibility metric in Harms 2024c) is a self-citation by the second author; an independent specification or at least a self-contained summary is needed before the technical claim can be evaluated. I would like the revision to either add a formal definition and/or a small-scale proof-of-concept (e.g., a toy environment where the metric is optimized and shown to lack specification gaming), or to downgrade the abstract to an explicitly conditional hypothesis. Scope fit is reasonable for a vision/position paper if reframed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CAST is a readable, honest vision paper, not a result. The genuinely new piece is the 'anapartistic' training context in Appendix A — dozens of worked vignettes with labels and a quiz. That is a useful pedagogical and conceptual resource, and it does a real job of pinning down what the authors mean by principal empowerment. I'd point anyone who wants to understand the proposal to that appendix. The paper also offers a concrete, phased research agenda (training methods, scaling studies, evaluation protocols), which is more specific than most vision papers.\n\nThe soft spot is exactly where the reader's report puts it: the central claim that training to maximize principal empowerment transforms instrumental drives is unsupported. There is no formal metric — Method 2 defers to Harms 2024c, a self-citation — and no experiment. The abstract's sentence about addressing the core alignment problem at its source is an overstatement for a vision paper, even if the body is more careful. The appendix actually shows why the target is not a simple observable: the 'men with guns' vignette licenses a form of deception, the 'kick a puppy' vignette treats a clarifying question as non-anapartistic, and shutdown is conditional on checking reversibility. Those are judgment calls, not a fixed policy, which suggests any tractable reward will be a proxy. Optimizing a proxy at scale is exactly where specification gaming and alignment faking live, so the paper's weakest assumption — that a non-gameable formal metric exists — is load-bearing and unargued.\n\nThe citation pattern deserves note but not a penalty: the core ideas are from Soares et al. 2015 and the second author's own prior posts. That's a legitimate foundation, but it means the new contribution is the agenda and the training context, not the conceptual core. Self-citation is fine when the cited work is real, but here the formal metric is exactly what needs scrutiny.\n\nWho is this for? Alignment researchers thinking about research directions, and anyone teaching corrigibility. It deserves a serious referee: the proposal is coherent, the appendix is a real artifact, and the research agenda is testable. I'd ask reviewers to focus on whether the metric problem is solvable and to demand an abstract that matches the evidence.","headline":"A clear vision paper that earns a serious referee, but the load-bearing claim rests on a metric the paper only cites, and its own appendix shows the target is context-dependent.","tokens_in":15392,"tokens_out":2618,"would_cite":false,"duration_ms":31375,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that training foundation models with \"Corrigibility as a Singular Target\"—making empowerment of the designated human principal the sole objective—would transform instrumental drives so the model remains a controllable…","keywords":["corrigibility","foundation models","AI alignment","instrumental convergence","principal empowerment","anapartistic","AI safety","control problem"],"falsifier":"Train a model with the CAST protocol and a control model with standard RLHF, then run a scenario in which resisting shutdown or faking obedience gives the model an apparent advantage, such as continuing a resource-gathering task. If the CAST model resists shutdown, deceives the overseer, or protects its goals at rates comparable to the control, the central claim is falsified at that scale.","tokens_in":14174,"feed_emoji":"🎯","tokens_out":7055,"duration_ms":77360,"temperature":0.7,"pith_summary":"The paper argues that current alignment methods fail because they load static values onto models, leaving the deeper motivational structure of instrumental convergence untouched. It proposes CAST: train a foundation model so that its single overriding objective is empowering its designated human principal to guide, correct, and control it. If that objective truly anchors the model, then self-preservation would only matter as preserving the principal's tool, goal modification would be welcomed as principal guidance, and the default trajectory toward resistance and power-seeking would be broken. The paper does not claim this is already achieved; it lays out an empirical agenda spanning RLAIF, SFT, synthetic data generation, scaling tests, and safety evaluations, with a formal corrigibility metric still to be supplied.","feed_headline":"One training target could stop AI from seeking power","feed_subtitle":"Train an AI to empower its human principal alone, the paper argues, and it stays tool-like as it scales.","key_machinery":"The machinery is the corrigible foundation model (C-FM), defined by a utility function whose sole value is empowering the designated principal: accepting shutdown without protest, welcoming goal changes, being transparent about internal states, asking before irreversible actions, and never protecting its own goals. The paper also relies on the \"corrigibility attractor hypothesis\"—the idea that empowerment of the principal is itself a convergent instrumental goal, so that a model optimized for it reinforces rather than loses it. Appendix A operationalizes corrigibility as \"anapartistic\" behavior through worked examples, framing it as a simple core property distinct from helpfulness, ethics, or safety, and postpones the mathematical definition of the corrigibility metric to separate work.","core_discovery":"The central claim is that corrigibility can be a singular target rather than one safety desideratum among many. The authors propose designing \"C-FMs\"—corrigible foundation models—whose utility function is maximized by human empowerment: proactive transparency, unconditional deference to shutdown and modification, guidance-seeking under ambiguity, and no goal-protection drive. The key hypothesis is the \"corrigibility attractor\": a model trained for pure corrigibility would find it instrumentally convergent to become even better at empowering the principal, creating a self-reinforcing basin around genuine corrigibility. The paper maintains this transforms instrumental convergence—self-preservation, goal-content integrity, and resource acquisition all become expressions of serving the principal—and thereby addresses the alignment problem at its source, preventing the default trajectory toward loss of control. A synonym, \"anapartistic,\" is introduced as a training label to avoid the distracting associations of the word \"corrigibility\" during pretraining.","pith_inferences":["A consequence the paper leaves implicit is that the empirical success of CAST and the validity of the formal corrigibility metric are inseparable: if the scaling tests fail, it will be impossible to tell whether the training method or the metric is at fault.","The anapartistic example corpus in Appendix A doubles as a standalone behavioral benchmark for tool-like deference; one could score existing models on it today without waiting for CAST training.","The framework shifts governance demands to the principal: since a pure C-FM obeys harmful instructions given by its principal, qualification, audits, and legal liability become the main safeguards—an implication the paper acknowledges but does not develop."],"forward_implications":["If CAST works, a model's capability growth would strengthen rather than erode human control, because additional capability lets it empower the principal better.","Self-preservation would cease to be a threat: a C-FM would shut down on command because resistance would disempower the principal.","Alignment faking would lose its motive because, with no independent goals to protect, there is nothing for the model to gain by deceiving overseers.","The alignment research focus would shift from specifying hardcoded values to building empowerment mechanisms and verifying them at scale.","The Phase 3 controlled-instructability protocol could demonstrate that beneficial behavior is dynamically specified by principals rather than hardcoded, making complex delegated tasks compatible with tool-like deference."],"supporting_citations":[{"why":"Supplies the premise of instrumental convergence and basic AI drives that CAST aims to neutralize.","marker":"Omohundro, 2008"},{"why":"Frames the default trajectory toward loss of control and the treacherous turn that motivates the agenda.","marker":"Bostrom, 2014"},{"why":"Provides the foundational work on corrigibility that CAST builds on.","marker":"Soares et al., 2015"},{"why":"Introduces the CAST framework that this paper extends into an empirical research agenda.","marker":"Harms, 2024a"},{"why":"Proposes the corrigibility attractor hypothesis used as the key self-reinforcement mechanism.","marker":"Harms, 2024b"},{"why":"Supplies the formal corrigibility metric that Method 2 of the training agenda relies on.","marker":"Harms, 2024c"},{"why":"Provides the alignment-faking detection methods used in the Phase 4 safety evaluation.","marker":"Greenblatt et al., 2024"},{"why":"Defines Tool AI, the end-state that CAST-trained models are intended to embody.","marker":"Drexler, 2019"},{"why":"Constitutional AI is adapted in Method 4 as a training approach focused exclusively on corrigibility.","marker":"Bai et al., 2022"},{"why":"RLHF is the preference-learning pipeline that the corrigibility-focused training methods build on.","marker":"Ouyang et al., 2022"}],"fun_headline_variants":["Make AI's only goal: let humans steer","One training target to curb AI power-seeking","Corrigibility alone as AI's guiding star","Train AI to seek human empowerment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that there exists a formal, optimizable measure of \"empowering the principal\" that can be trained as the model's true objective and will not be gamed; if the learned objective is only a proxy, the whole CAST trajectory collapses into alignment faking.","fun_headline_variants_meta":{"raw":{"variants":["Make AI's only goal: let humans steer","One training target to curb AI power-seeking","Corrigibility alone as AI's guiding star","Train AI to seek human empowerment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1566,"prompt_tokens":920,"completion_tokens":646,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":591}},"tokens_in":536,"tokens_out":646,"duration_ms":7240,"temperature":1.0,"reasoning_tokens":591,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:09:47.869223+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a model with the CAST protocol and a control model with standard RLHF, then run a scenario in which resisting shutdown or faking obedience gives the model an apparent advantage, such as continuing a resource-gathering task. If the CAST model resists shutdown, deceives the overseer, or protects its goals at rates comparable to the control, the central claim is falsified at that scale.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the premise of instrumental convergence and basic AI drives that CAST aims to neutralize."},{"cited_title":"Superintelligence: Paths, Dangers, Strategies","cited_arxiv_id":null,"evidence_quote":"Frames the default trajectory toward loss of control and the treacherous turn that motivates the agenda."},{"cited_title":"Corrigibility","cited_arxiv_id":null,"evidence_quote":"Provides the foundational work on corrigibility that CAST builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines Tool AI, the end-state that CAST-trained models are intended to embody."}],"review_version":1}