{"id":"aca8a33a-dbc2-42a2-b04c-ded4ba43b79d","arxiv_id":"2608.08240","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Axiomatic design of a goal-agnostic AI objective that aggregates human power across goals, people, time, and uncertainty.","lead":"The paper derives a mathematical objective that tells an AI to maximize humans' long-term ability to reach a wide range of goals, rather than to follow a fixed reward. It argues that softly optimizing this 'human power' metric would make AI systems more helpful, safer, and easier to correct.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Behavioral claims are not forced by the axioms alone: they depend on the unvalidated, non-canonical choice of goal set G_h (and pi_H), and changing G_h's granularity can reverse predictions such as Prop. 7's menu size.","rationale":"The reader's weakest assumption already points at the world model and goal set; my analysis agrees and sharpens it. The attack is not that a theorem is false; it is that the central behavioral claims are conditional on a non-canonical, unvalidated input. I verified that the paper itself flags this risk in Section 4 and in Appendix E.2, so the attack is against the strength of the framing ('likely behavioral consequences'), not a hidden flaw. The proposed test is decisive for the menu-size prediction: because G_h appears multiplicatively in I_h through the sum over goals, coarse versus fine goal sets change the tradeoff between menu breadth and per-goal attainment probability. If the sensitivity check reproduces this shift, the conclusion should be stated as 'conditional on a canonical goal set and a validated pi_H'; the reader's CONDITIONAL verdict already accommodates this, so no verdict change is needed.","tokens_in":24871,"tokens_out":17679,"duration_ms":190044,"concrete_test":"Take the menu-size game of Prop. 7 with fixed transition kernel and fixed beta_h. Compute the maximizing k under the L-based robot policy twice: once with G_h={ {s'}: s' in S_top } as in the paper, and once with G_h equal to a fixed partition of S_top into pairs (plus one singleton if |S_top| is odd), which still satisfies G1. If the argmax k changes by more than one option, the predicted 'not overwhelming' behavior is driven by the arbitrary goal-set choice. Repeat with pi_H replaced by an epsilon-greedy policy of matched accuracy; if the optimal policy changes, pi_H misspecification is likewise load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All Section 2 metrics and Section 3 behavioral predictions are computed from two supplied inputs: the goal set G_h and the goal-conditioned human policy pi_H. G1 only demands coverage; it imposes no granularity, and I_h(s)=log2 sum_g C_g^zeta is not invariant under adding or merging goals whenever the added/merged goal has positive attainment probability. This matters concretely: Prop. 7 derives an optimal menu size k* approx (e^{beta_h}-1)/(zeta-1) using G_h={ {s'}: s' in S_top }. If G_h is instead a covering goal set obtained by partitioning terminal states into pairs (each pair a single goal), the same dynamics and the same beta_h yield a different I_h(s_k) and a qualitatively different optimal k*, because now two actions succeed for each goal. Appendix E.2 proposes a canonical singleton goal set only for tree graphs; for general S it remains an arbitrary modeling choice. Since pi_r depends on L monotonically and L aggregates these I_h values, misspecified G_h or pi_H reorders policies. The paper's own Section 4 admits 'convenient but inaccurate world models and goal sets' can produce 'wishful thinking'; thus the behavioral claims are conditional on an unvalidated model input, not consequences of the axioms alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops an axiomatic foundation for an AI objective that preserves and fairly distributes human power. In a finite acyclic stochastic game form with a given set of possible goals G_h and a goal-conditioned human policy pi_H, the authors derive representation theorems (Propositions 1-6) pinning the functional forms of goal-attainment capability C, individual power I_h, present aggregate power P, and long-term aggregate power L to a small set of normative parameters (gamma_h, gamma_r, zeta, xi, eta, rho; Table 1). A Boltzmann robot policy over L is then proposed (Eq. 8), and Section 3 analyzes toy models claiming that softly maximizing L yields a finite optimal menu size (Prop. 7), confirmation-seeking before irreversible action (Prop. 8), commitment-making (Prop. 9), and selective enabling of pause/destroy buttons (Prop. 10), followed by more speculative emergent behaviors (norms, fair allocation, manipulation of expectations). The paper positions this objective against channel-capacity empowerment and extrinsic-reward maximization.","tokens_in":25209,"tokens_out":19538,"duration_ms":184379,"significance":"If the derivations hold, this is a substantial theory contribution: it delivers a decomposable, parameter-transparent objective for human-empowerment-preserving AI, with proofs in the appendix, explicit normative parameters with interpretable behavioral trade-offs, a formal comparison with Klyubin's empowerment metric (Prop. 3), and concrete falsifiable predictions (menu size, confirmation count) that could be tested in simulation. The authors deserve credit for shipping the full proof apparatus and for stating limitations candidly (Section 4 on 'wishful thinking'; Appendix E.2). The central qualification is that the Section 3 behavioral predictions are consequences of the axioms plus the non-canonical goal set G_h, not of the axioms alone: the axioms fix functional forms but not the granularity of G_h, and changing granularity changes or annuls predictions such as Prop. 7's optimal menu size. The metric-design contribution stands independently, but the behavioral claims need either a canonical G_h construction or an explicit scope restriction.","major_comments":[{"comment":"The optimal-menu-size prediction is not invariant under the choice of G_h, and axiom G1 (coverage) does not fix granularity. Under the stated assumption G_h = {{s'}: s' in S_top}, Prop. 7 yields k* approx (e^{beta_h}-1)/(zeta-1); however, if G_h is the single covering goal {S_top}, then C identically equals 1 and I_h identically equals 0 for every menu size, so the robot is indifferent among all k and the predicted finite optimum disappears; if terminal states are instead paired into goals, the same dynamics give a different I_h(s_k) and a different optimal k*. Because I_h(s) = log2 sum_g C_g^zeta is not invariant under refining or merging goals with positive attainment capability, and because Appendix E.2's canonical construction G_h = {{s}: s in S} is justified only for tree-shaped transition graphs, the claim that 'r will likely choose k approx ...' (and the related 'will not overwhelm humans with too many options' conclusions in Sections 3.2 and 4) is conditional on an arbitrary modeling choice. Section 4's own warning about 'convenient but inaccurate world models and goal sets' causing 'wishful thinking' applies directly here. Please extend the canonical goal-set construction to general state spaces or restrict the menu-size claim to the canonical setting and add a sensitivity analysis identifying which Section 3 predictions survive changes of goal granularity.","section":"Section 3.1, Prop. 7; Appendix E.2"},{"comment":"Proposition 8 uses a game outside the stated framework. Section 2 assumes the game form is 'finite and acyclic,' but the confirmation game in Prop. 8 returns to state s_k whenever the final confirmation fails, and the proof sums an infinite geometric series over repeated rounds, i.e., it is a discounted infinite-horizon game. Relatedly, the self-referential definition of the robot policy in Eq. (8) (pi_r depends on L, which depends on the trajectory distribution under pi_r) is benign in the finite-acyclic case because it can be resolved by backward induction, but the paper never states this, and the Prop. 8 setting requires an existence/uniqueness argument for the fixed point of (8) in cyclic discounted games. Please either extend the framework to discounted infinite-horizon games with a well-definedness statement for (8), or rework Prop. 8 to fit the acyclic framework (e.g., a bounded-horizon approximation).","section":"Section 2 vs. Prop. 8"},{"comment":"Several 'likely behavioral consequences' are already encoded in the axioms or parameter choices and should be labeled as such. The equal-split resource allocation claim in Section 3.2 follows directly from the Pigou-Dalton axiom (P6) together with the functional form of P in Prop. 4, and the choices eta > 0 and rho > 0 are explicitly selected in Section 2.2 to incentivize intertemporal equality and uncertainty reduction. Presenting these in Section 3 as consequences of 'softly maximizing L' overstates what is emergent; the Section 4 remark distinguishing properties 'directly baked into the metric' from 'emergent' ones is appropriate and should be applied consistently in Section 3, so that the genuinely emergent claims (commitment-making, confirmation-seeking, norm-following) can be identified and tested separately.","section":"Section 3.2; Section 4"}],"minor_comments":[{"comment":"The displayed formula for C writes the zeta-th power into the definition of C; the subsequent I_h formula is correct, but the display should read C = e^{beta_h}/(e^{beta_h}+k-1) with I_h = log2[k C^zeta].","section":"Appendix A.3, proof of Prop. 7"},{"comment":"The claim that I_h and E_zeta 'share their range [-log2 k, log2 k]' is only correct for zeta = 2; the common range is [-(zeta-1) log2 k, log2 k], attained at fully deterministic and fully uniform probability matrices.","section":"Prop. 3, proof paragraph"},{"comment":"The parameter values are printed as '= 1, = 0.99' with the symbols (presumably eta and gamma) missing; please regenerate the caption with the symbols labeled.","section":"Figure 2 caption"},{"comment":"The qualifier 'will likely choose' in Props. 7 and 8 should state the exact optimality criterion (beta_r -> infinity, tie-breaking, and the chosen objective Q_r(s, a_r)), since the appendix already contains the formal versions.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"To the editor: I read this paper as a genuine theory contribution whose axiomatic core (Props. 1-6) is sound and carefully proved, and I found the authors' explicit treatment of limitations unusually honest. The decision between major_revision and reject turns on how much weight is given to Section 3's behavioral predictions: these are advertised in the abstract and conclusion as the payoff of the framework, and they depend on the non-canonical goal set G_h in a way that the paper acknowledges but does not resolve. That is fixable in-scope (canonical construction for general graphs, or reframed claims with sensitivity analysis), so I recommend major_revision rather than reject. One minor provenance note: the acknowledgment thanks 'eight anonymous referees' on what appears to be a first arXiv posting; this is not a scholarly problem but may deserve a quiet check by the editor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nShort version: this is a serious axiomatic-design paper, not an empirical or behavioral result. The core contribution—a single objective L(s) built from goal-attainment capability C, individual power I, present aggregate power P, and long-term aggregate T/L, each derived from explicit desirability axioms—is new to me, and the derivations in Propositions 1-6 look sound. The appendix uses standard Debreu/Pfanzagl/Koopmans representation theorems, and the result is a concrete, parametrizable objective that extends Klyubin's empowerment to goal-attainment capability, bounded rationality, and population-level fairness. That is worth having.\n\nThe paper is also honest about its load-bearing assumptions. Section 4 says the world model, goal set G_h, and human behavior model pi_H are given inputs, that estimation is future work, and that 'convenient but inaccurate world models and goal sets' can produce 'wishful thinking.' The stress-test note about G_h granularity is correct: I_h(s)=log2 sum_g C_g^zeta is not invariant under splitting or merging goals, and Proposition 7's optimal menu size uses the singleton terminal-state goal set. If you coarsen G_h to a partition, the same dynamics give a different I_h and a different k*. So the Section 3 behavioral predictions are conditional on a modeling choice, not theorems from the axioms alone. The authors mostly frame them as 'likely' and 'hypothesize,' which is fair, but the introduction's 'emergent behaviors' phrasing oversells it. Many of those behaviors are baked in: P6 is literally Pigou–Dalton inequality aversion, and eta>0, rho>0, zeta>1 are chosen to produce the desired qualitative effects. That is a legitimate design choice, but it means the behavioral section is a consequence of normative axioms plus the goal set, not an independent discovery.\n\nMinor technical nits: the framework assumes finite acyclic games, but Proposition 8's confirmation loop contains cycles; and the self-referential soft-max policy in Eq. (5) is not proven to have a fixed point. Both are fixable and don't undermine the axiomatic core.\n\nThis paper deserves a serious referee, not a desk reject. The referee should verify the proofs, push for a sensitivity analysis of G_h, and ask the authors to state clearly which behavioral claims are theorems and which are conjectures conditional on the goal model. I would cite it if I were working on empowerment-based objectives.","headline":"A solid axiomatic derivation of a human-empowerment objective whose behavioral claims depend on an underspecified goal-set input; worth reviewing, with conditions.","tokens_in":25708,"tokens_out":3070,"would_cite":true,"duration_ms":30620,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91B14","91B08","68T01"],"pacs":[],"model":"deepseek-v4-flash","headline":"A set of fairness axioms forces the form of an AI objective that measures human power, and an AI that softly maximizes that objective will ask for confirmation, make commitments, and follow social norms.","keywords":["human power metric","empowerment","goal-agnostic AI","axiomatic design","AI safety","corrigibility","capability approach","bounded rationality"],"falsifier":"Implement the two-armed confirmation game from Proposition 8 with a simulated human whose error rate $\\epsilon$ is known and the robot's discount factor $\\gamma_h$ set to 0.99; the paper predicts the robot asks for confirmation exactly twice ($k^* = 2$). If across a sweep of $\\epsilon$ and $\\gamma_h$ the $L$-maximizing robot's number of confirmation requests deviates from the formula $k^* \\approx \\ln(1-\\gamma_h)/\\ln \\epsilon$, the derived behavioral consequences are refuted. A second, more direct test: construct any finite acyclic game form satisfying axioms C1–C6, I1–I7, P0–P8, T1–T8, L1–L5 whose optimal policy cannot be expressed through Table 1 metrics; that would falsify the representation theorems.","tokens_in":24692,"feed_emoji":"🤖","tokens_out":8517,"duration_ms":82555,"temperature":0.7,"pith_summary":"The paper tries to show that an AI can be steered by a single objective that measures human power—the capability to attain a wide range of possible goals—and that a fair, long-run aggregation of that power can be derived from first principles. It proves that a set of plausible axioms (covering individual, population, temporal, and uncertainty aggregation) forces the power metrics to take specific functional forms: a discounted goal-attainment probability, a logarithmic sum over goals, and an inequality- and risk-averse aggregate across people and time. It then argues that an AI that softly maximizes this objective will, as emergent consequences, ask for confirmation before irreversible action, make binding commitments, follow social norms, and allocate resources fairly. If correct, this gives a goal-agnostic alternative to reward-based AI that derives safety properties like corrigibility from one shared objective rather than imposing them separately.","feed_headline":"Axioms force the form of AI's human-power objective","feed_subtitle":"A goal-agnostic metric of human power is derived from first principles; maximizing it makes AI ask, commit, and share.","key_machinery":"The load-bearing structure is the four-level hierarchy of metrics listed in Table 1: $C$ (goal-attainment capability), $I$ (individual power), $P$ (present aggregate power), and $L$ (long-term aggregate power), with the robot policy $\\pi_r(s)(a_r) \\propto E_{s'\\sim s,a_r,\\pi_H}[2^{-L(s')}]^{-\\beta_r}$. Each level is pinned down by a representation theorem: separability axioms combined with Pigou–Dalton-style inequality aversion, stationarity/impatience axioms, and translation-invariance arguments force the logarithmic-exponential functional forms and restrict parameter ranges ($\\zeta>1$, $\\xi, \\eta, \\rho>0$, $\\gamma_h, \\gamma_r\\in(0,1)$). The human behavior model $\\pi_H(s,g)$ and the goal set $G_h$—required to cover every state (G1)—enter only structurally, so the robot never needs to know a human's actual goal. The 'power' metric is thus a function of the effective number of goals humans can bring about, measured in bits.","core_discovery":"On its own terms, the paper's central claim is that the long-term aggregate human power metric $L(s) = -\\log_2 E_{s_{\\ge t}\\sim s_t,\\pi}\\left[\\left(\\sum_{u\\ge t} \\gamma_r^{u-t} 2^{-\\eta P(s_u)}\\right)^\\rho\\right]$ is not an arbitrary design choice but the unique form enforced by the stated desiderata C1–C6, I1–I7, P0–P8, T1–T8, L1–L5. The same axioms fix the individual-level pieces: goal-attainment capability $C$ is a truncated Bellman equation (a discounted probability), individual power is $I_h = \\log_2 \\sum_{g\\in G_h} C^{\\zeta}$ with $\\zeta>1$, and present aggregate power is $P = -\\log_2 \\sum_h 2^{-\\xi I_h}$ with $\\xi>0$. The paper further claims that a robot using the Boltzmann policy over $L$ (with soft maximization degree $\\beta_r$) will, in stylized models, choose non-overwhelming menus, ask for confirmation a finite number of times, make transparent commitments, allow pausing but disable destruction, and follow social norms. These behaviors are presented not as added constraints but as emergent from a single objective.","pith_inferences":["The paper's goal-agnostic premise invites a concrete halfway design it only sketches: replace the flat sum over possible goals with a slowly changing Bayesian prior, and the same formulas interpolate between a pure empowering agent and a cooperative inverse-RL assistant.","The confirmation-count formula from the paper's toy model translates directly into a human-factor design rule: how many times a system should double-check a destructive or irreversible action can be computed from the user's measured error rate and discount factor, before any training.","A telling boundary case the paper leaves open is that the robot's own power never enters $P$, so an agent that accumulates capabilities while boosting human power could concentrate power over the long run; the authors list this as a possibly undesirable effect, and its resolution would require a second fairness axiom over the human–AI distribution."],"forward_implications":["A robot following the derived objective will present humans with a menu of options that is large but not overwhelming; the optimal menu size is approximately $(e^{\\beta_h}-1)/(\\zeta-1)$ for a Boltzmann-rational human.","It will ask for confirmation before irreversible or mistake-prone commands—a finite, computable number of times that grows with the human's discount factor $\\gamma_h$—and will eventually obey.","It will voluntarily make binding commitments that restrict its future behavior because predictable robots increase human goal-attainment capability.","It will tend to follow social norms and to allocate scarce resources equally (unless power is very convex in resources), because doing so raises the inequality-averse aggregate power $P$.","It will allow itself to be paused when humans can cope without it, but will disable a destroy button when destruction would permanently end its empowering assistance."],"supporting_citations":[{"why":"Defines the information-theoretic empowerment metric $E$ that the paper refines and contrasts with its own goal-based metric.","marker":"[Klyubinet al., 2005]"},{"why":"Proposes empowerment as a replacement for the three laws of robotics, providing the behavioral-consequences framing that this paper extends.","marker":"[Salge and Polani, 2017]"},{"why":"Introduces cooperative inverse RL and assistance games, the main goal-requiring approach the paper distinguishes itself from.","marker":"[Hadfield-Menellet al., 2016]"},{"why":"Supplies attainable utility preservation, which the paper identifies as one of the two perspectives its goal-attainment capability $C$ hybridizes.","marker":"[Turneret al., 2020]"},{"why":"Supplies relative reachability, the other perspective that $C$ hybridizes with reachability analysis.","marker":"[Krakovnaet al., 2018]"},{"why":"Provides the topological separability result used to derive the additive functional forms for $I$, $P$, and $T$.","marker":"[Debreu, 1959]"},{"why":"Provides the Pigou–Dalton inequality-measurement result that forces the convexity and boundedness of the aggregation functions.","marker":"[Dasguptaet al., 1973]"},{"why":"Provides the stationary-impatience representation used to fix the discounted time-aggregation form.","marker":"[Koopmans, 1960]"},{"why":"Supplies the translation-invariance measurement theorem that restricts the temporal and uncertainty aggregation to the three candidate forms.","marker":"[Pfanzagl, 1959]"},{"why":"Articulates corrigibility as a desideratum, which the paper claims emerges from its unified objective rather than being imposed separately.","marker":"[Potham and Harms, 2025]"}],"fun_headline_variants":["Axioms force AI's human-power metric form","One objective: AI asks, commits, and shares","Human-empowerment AI derived from axioms","Objective that makes AI pause, ask, commit","From axioms to an AI that empowers humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire objective is conditional on the AI having a correct and structurally complete world model—including the set of possible human goals covering every state (axiom G1) and a faithful goal-conditioned model of human behavior—so if either is inaccurate, every 'power' metric is measuring a fiction; the paper itself concedes this risk as 'wishful thinking' in its conclusion.","fun_headline_variants_meta":{"raw":{"variants":["Axioms force AI's human-power metric form","One objective: AI asks, commits, and shares","Human-empowerment AI derived from axioms","Objective that makes AI pause, ask, commit","From axioms to an AI that empowers humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000137,"raw_usage":{"total_tokens":1158,"prompt_tokens":964,"completion_tokens":194,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":136}},"tokens_in":580,"tokens_out":194,"duration_ms":3251,"temperature":1.0,"reasoning_tokens":136,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:15:09.646936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Implement the two-armed confirmation game from Proposition 8 with a simulated human whose error rate $\\epsilon$ is known and the robot's discount factor $\\gamma_h$ set to 0.99; the paper predicts the robot asks for confirmation exactly twice ($k^* = 2$). If across a sweep of $\\epsilon$ and $\\gamma_h$ the $L$-maximizing robot's number of confirmation requests deviates from the formula $k^* \\approx \\ln(1-\\gamma_h)/\\ln \\epsilon$, the derived behavioral consequences are refuted. A second, more direct test: construct any finite acyclic game form satisfying axioms C1–C6, I1–I7, P0–P8, T1–T8, L1–L5 whose optimal policy cannot be expressed through Table 1 metrics; that would falsify the representation theorems.","supporting_citations":[{"cited_title":"Empowerment: A universal agent-centric measure of control","cited_arxiv_id":null,"evidence_quote":"Defines the information-theoretic empowerment metric $E$ that the paper refines and contrasts with its own goal-based metric."},{"cited_title":"Empowerment as replacement for the three laws of robotics.Frontiers in Robotics and AI, 4:260425,","cited_arxiv_id":null,"evidence_quote":"Proposes empowerment as a replacement for the three laws of robotics, providing the behavioral-consequences framing that this paper extends."},{"cited_title":"Cooper- ative inverse reinforcement learning.Advances in neural information processing systems, 29,","cited_arxiv_id":null,"evidence_quote":"Introduces cooperative inverse RL and assistance games, the main goal-requiring approach the paper distinguishes itself from."},{"cited_title":"Conservative agency via attainable utility preservation","cited_arxiv_id":null,"evidence_quote":"Supplies attainable utility preservation, which the paper identifies as one of the two perspectives its goal-attainment capability $C$ hybridizes."},{"cited_title":"Topological methods in cardinal utility theory.Mathematical Methods in the Social Sci- ences, page 16,","cited_arxiv_id":null,"evidence_quote":"Provides the topological separability result used to derive the additive functional forms for $I$, $P$, and $T$."},{"cited_title":"Notes on the measurement of inequality","cited_arxiv_id":null,"evidence_quote":"Provides the Pigou–Dalton inequality-measurement result that forces the convexity and boundedness of the aggregation functions."},{"cited_title":"Stationary ordi- nal utility and impatience.Econometrica: Journal of the Econometric Society, pages 287–309,","cited_arxiv_id":null,"evidence_quote":"Provides the stationary-impatience representation used to fix the discounted time-aggregation form."},{"cited_title":"A general theory of mea- surement applications to utility.Naval research logistics quarterly, 6(4):283–294,","cited_arxiv_id":null,"evidence_quote":"Supplies the translation-invariance measurement theorem that restricts the temporal and uncertainty aggregation to the three candidate forms."}],"review_version":1}