Pith. sign in

REVIEW 5 minor 24 references

Will artificial agents pursue power by default?

T0 review · 0 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Power is a convergent AI goal—but only an incomplete one.

desk verdict A careful formal treatment of power-seeking that delivers a new incompleteness result and honestly disclaims its own limits. read the letter →

arxiv 2506.06352 v1 pith:DIHDZZOK submitted 2025-06-02 cs.AI cs.CY

classification cs.AIcs.CY
keywords instrumentalconvergencepower-seekingAIalignmentdecisiontreesexpectedutilitytheoryorthogonalitythesisabsolutepowerrisk
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether artificial agents will pursue power by default because power helps almost any final goal. It formalizes this question in finite decision trees where an agent's final goals are drawn as a random utility function under minimal symmetry constraints. It proves that several natural power rankings—more available lotteries, and more available lotteries in every direction—are instrumentally convergent, so random agents are more likely, and sometimes certain, to choose them. But it also proves that no complete power ranking can be instrumentally convergent under those assumptions, so power-seeking is not a fully general prediction. The practical upshot is that instrumentally convergent power-seeking is most predictive for agents that can plausibly attain absolute or near-absolute power.

What carries the argument

The central object is a finite decision tree whose root offers a set of lotteries over final outcomes, with an agent modeled as an expected utility maximizer whose utility function is drawn from a distribution that is symmetric under permutations of outcomes, regular in the sense of assigning nonzero probability to every open region of utility assignments, and almost surely non-uniform. Power is defined set-theoretically on the collections of lotteries available at subtrees: inclusion ($L(t_j) \subseteq L(t_i)$), interiority ($L(t_j)$ inside the interior of $L(t_i)$), and their permutation-invariant forms ($\pi(L(t_j)) \subseteq L(t_i)$ and $\pi(L(t_j))$ inside the interior of $L(t_i)$). Instrumental convergence is defined by whether a random agent is at least as likely, strictly more likely, or certain to choose the action leading to the higher-ranked subtree. The positive proofs use convexity and the hyperplane separation theorem, while Proposition 5 works by constructing a pair of subtrees whose relative attractiveness reverses under nearly binary utility distributions, so no distribution-independent total ranking can exist.

What would settle it

Exhibit a total reflexive relation on finite decision trees that is weakly instrumentally convergent and whose strict part is strictly instrumentally convergent while satisfying the paper's symmetry, regularity, and almost-sure non-uniformity assumptions; Proposition 5 asserts that no such relation exists.

Watch

Extended reading notes

Core claim

The paper's central claim is that power, formalized as the ability to achieve a wider set of probability distributions over outcomes, is a convergent instrumental goal in a nontrivial but limited sense. Four power relations are defined on decision trees: inclusion, where one tree's available lotteries contain another's; interiority, where one tree's lottery set contains another's in its interior; and their permutation-invariant analogues. The first is weakly instrumentally convergent, the second is absolutely instrumentally convergent, the third is weakly but not strictly convergent, and the fourth is strictly but not absolutely convergent. The negative result, Proposition 5, says that no total reflexive relation on decision trees can be weakly instrumentally convergent while its irreflexive part is strictly instrumentally convergent, so any complete ranking of options by power must either depend on substantive assumptions about the distribution of an agent's goals or fail to predict choice. The author concludes that the claim that power is a convergent instrumental goal contains an element of truth but has limited predictive utility, except for agents who can approach absolute or near-absolute power.

Load-bearing premise

The paper treats total ignorance of an agent's final goals as a probability distribution over utility functions that treats all outcomes symmetrically, never rules out any utility function, and almost never assigns the same utility to every outcome; if real AI training produces goal distributions with systematic biases or correlations, the probability claims can fail.

Editorial extensions

If this is right

  • An agent with randomly generated final goals is more likely than not to choose an option that gives it a strictly larger set of available lotteries, and is almost certain to choose an option whose lottery set contains the alternative in its interior.
  • Survival, information acquisition, and generic resource accumulation are vindicated as convergent instrumental goals in the special cases where they expand, or interior-expand, the set of available lotteries.
  • No power ranking derived purely from ignorance of final goals can be complete: there are pairs of subtrees where which option a random agent prefers depends on features of the utility distribution beyond symmetry, regularity, and almost-sure non-uniformity.
  • When an agent can approach absolute power, the incompleteness is less damaging, because near-absolute power is near-maximally attractive for almost all non-uniform expected utility maximizers.
  • In a multipolar world where no agent can approach absolute control and power-seeking is costly, the fact of instrumental convergence gives relatively little basis for predicting that agents will seek power.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the incompleteness result pushes the practical question onto the actual training distribution; quantitative predictions about power-seeking require modeling the specific process that generates an agent's utility function, not just modeling ignorance of its goals.
  • Inference: the framework suggests a testable prediction: agents trained with broad goals in rich environments will show stronger power-seeking when they can achieve near-deterministic control, and weaker power-seeking when they face multipolar competition with costly failures.
  • Inference: because the paper's notion of power is not inherently rivalrous, in zero-sum settings the convergent goal may translate into disempowering other agents; this connection, left open by the paper, could be formalized by adding conflicts among multiple agents.
  • Inference: if real agents have non-consequentialist preferences, absolute instrumental convergence would fail, but weak and strict convergence might still hold under ignorance if intrinsic aversions to power-seeking are no more probable than intrinsic attractions to it.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. This paper develops a decision-theoretic formalization of instrumental convergence and power-seeking. The author defines a "random agent" as an expected-utility maximizer whose utility function is drawn from a symmetric, regular, and almost surely non-uniform distribution, and defines weak, strict, and absolute instrumental convergence as comparative likelihood properties of relations on decision trees. Four notions of power are considered: inclusion (≻p1), interiority (≻p2), permuted inclusion (≻p3), and permuted interiority (≻p4). Propositions 1–4 establish that these relations are convergent in the appropriate senses, with ≻p2 absolutely convergent and ≻p4 strictly but not absolutely convergent. Proposition 5 shows that no total reflexive relation on decision trees can be weakly instrumentally convergent with a strictly convergent irreflexive part, implying that power rankings are either incomplete or non-predictive without distributional information. Section 6 argues informally that absolute or near-absolute power is a more predictive target, and Section 7 explicitly disclaims any defense of the orthogonality thesis or alignment difficulty.

Significance. The paper makes a valuable and original contribution to the instrumental-convergence debate. It provides a general framework that avoids the specific RL assumptions of earlier power-seeking theorems, and it cleanly separates the logical relations among different notions of convergence. The positive results (Propositions 1–4) are elementary but nontrivial, and the negative result (Proposition 5) is a genuinely non-circular limitation result that addresses the predictive-utility question raised by critics such as Gallow and Thorstad. The paper is also commendably explicit about the limits of its framework, particularly the dependence on the "random agent" prior and the absence of any defense of the orthogonality thesis. These features make the paper suitable for a broad AI-safety or philosophy audience.

minor comments (5)
  1. [§5, Proposition 5] The proof of Proposition 5 needs an explicit construction of regular distributions that approximate the two point-mass distributions. The sentence 'they can be approximated arbitrarily closely be distributions that do' asserts the key step without demonstration; a concrete construction (e.g., a mixture of the symmetric orbit of the point mass with a small full-support continuous component) would make the proof complete.
  2. [Appendix, Proposition 4] The proof of Proposition 4 is compressed at the final step. To establish strict instrumental convergence, the author should explicitly state that any utility function satisfying u(π(t_j)) > u(t_j) also satisfies u(t_i) > u(t_j), because π(L(t_j)) ⊆ L(t_i), and that the positive-measure set of functions preferring t_j to π(t_j) and t_i to t_j breaks the symmetry equality between the two orderings.
  3. [§2] The definition of 'interior' for sets of lotteries should be clarified as relative interior with respect to the probability simplex, since a set of lotteries is a subset of a lower-dimensional affine space and may have empty Euclidean interior.
  4. [§4, Proposition 1 proof] In the proof of Proposition 1, the sentence 'if there is an expected-utility-maximizing lottery in L(t_j), that same lottery is in L(t_i), so the agent is equally likely to choose ai or aj' is potentially misleading because such a lottery always exists; the intended condition is that the maximal expected utility is equal in the two subtrees. Please rephrase.
  5. [Abstract] In the abstract, 'aconservative' should be 'a convergent'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the formal results are valid consequences of stated definitions, the near-tautological positive results are explicitly acknowledged, and the negative result is independent.

full rationale

The paper's central claims are theorems derived from explicitly stated definitions rather than empirical predictions, so the derivation chain is self-contained. The only step that could be viewed as definitional is Power 1/Power 2, where 'more power' is defined as a larger set of available lotteries, making the preference for more power a near-immediate consequence of expected-utility maximization. However, the author transparently acknowledges this by quoting Gallow's tautology objection and noting that Propositions 1 and 2 are 'formally fairly elementary' rather than substantive empirical discoveries. The paper does not present these as predictions about real agents; instead, it emphasizes their limited predictive utility and proves an independent negative result (Proposition 5) showing that no total instrumental-convergence ranking exists under the stated ignorance assumptions. Section 7 explicitly disclaims any defense of the orthogonality thesis or alignment difficulty, so the 'random agent' prior is not smuggled in as a conclusion. There is no load-bearing self-citation, fitted parameter renamed as a prediction, or imported uniqueness theorem. The derivation is therefore not circular.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central results rest on the random-utility prior and standard decision-theoretic assumptions; no free parameters are fitted. No new physical or metaphysical entities are postulated; the power relations P1-P4 are formal definitions.

assumptions (5)
  • domain assumption Agents are expected utility maximizers with a utility function over final outcomes.
    Section 2 states this as an assumption about instrumental rationality, not about the content of final goals.
  • domain assumption The set of outcomes O is finite with cardinality at least 3.
    Finite outcomes ensure compactness and convexity of accessible lottery sets, used in the proofs.
  • domain assumption A random agent's utility function is drawn from a symmetric, regular, and almost surely non-uniform distribution.
    Section 3 defines 'randomly generated' this way to represent ignorance about final goals; all probability claims are relative to this prior.
  • domain assumption Agents are certain their utility will not change over time and will continue to maximize expected utility at future choice nodes.
    Section 3 makes this assumption to allow backward-induction style evaluation of subtrees.
  • standard math Standard results from convex analysis, including the hyperplane separation theorem, are valid.
    Used in the proofs of Propositions 1, 4, and 5.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Will artificial agents pursue power by default?." pith.science (2026). https://pith.science/paper/DIHDZZOK

@misc{pith2026250606352,
  author       = {Pith},
  title        = {Pith review of: Will artificial agents pursue power by default?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DIHDZZOK}},
  note         = {Machine review of arXiv:2506.06352}
}
read the original abstract

Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that is useful for a wide range of final goals. Others have recently expressed skepticism of these claims. This paper aims to formalize the concepts of instrumental convergence and power-seeking in an abstract, decision-theoretic framework, and to assess the claim that power is a convergent instrumental goal. I conclude that this claim contains at least an element of truth, but might turn out to have limited predictive utility, since an agent's options cannot always be ranked in terms of power in the absence of substantive information about the agent's final goals. However, the fact of instrumental convergence is more predictive for agents who have a good shot at attaining absolute or near-absolute power.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

24 extracted references · 18 canonical work pages

  1. [1]

    D’Alessandro, and C

    Bales, A., W . D’Alessandro, and C. D. Kirk-Giannini (2024). Artificial in- telligence: Arguments for catastrophic risk.Philosophy Compass 19(2), e12964

  2. [2]

    Benson-Tilsen, T . and N. Soares (2016). Formalizing convergent instru- mental goals. InWorkshops at the Thirtieth AAAI Conference on Artificial Intelligence

  3. [3]

    Bostrom, N. (2003). Ethical issues in advanced artificial intelligence. In I. Smit, W . Wallach, and G. E. Lasker (Eds.),Cognitive, Emotive and Ethical Aspects of Decision Making in Humans and in Artificial Intelligence, Vol- ume 2, pp. 12–17. International Institute for Advanced Studies in Systems Research and Cybernetics

  4. [4]

    Bostrom, N. (2012). The superintelligent will: Motivation and instrumental rationality in advanced artificial agents.Minds and Machines 22(2), 71–85

  5. [5]

    (2014).Superintelligence: Paths, Dangers, Strategies

    Bostrom, N. (2014).Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press

  6. [6]

    Boyd, S. P . and L. Vandenberghe (2004).Convex Optimization. Cambridge University Press. 27

  7. [7]

    Carlsmith, J. (2022). Is power-seeking AI an existential risk? arXiv: 2206.13353v1[cs.CY]

  8. [8]

    Dung, L. (2024). The argument for near-term human disempowerment through AI.AI & Society, 1–14

Show all 24 references
  1. [9]

    Gabriel, I. (2020). Artificial intelligence, values, and alignment.Minds and Machines 30(3), 411–437

  2. [10]

    Gallow, J. D. (2024). Instrumental divergence.Philosophical Studies, 1–27

  3. [11]

    Goldstein, S. and C. D. Kirk-Giannini (2023). Language agents reduce the risk of existential catastrophe.AI & Society, 1–11

  4. [12]

    Good, I. J. (1966). On the principle of total evidence.British Journal for the Philosophy of Science 17(4), 319–321

  5. [13]

    Mazeika, and T

    Hendrycks, D., M. Mazeika, and T . Woodside (2023). An overview of catas- trophic AI risks. arXiv: 2306.12001v6[cs.CY]

  6. [14]

    Krakovna, V . and J. Kramar (2023). Power-seeking can be probable and predictive for trained agents. arXiv: 2304.06528[cs.AI]

  7. [15]

    Chan, and S

    Ngo, R., L. Chan, and S. Mindermann (2023). The alignment problem from a deep learning perspective. arXiv: 2209.00626v5[cs.AI]

  8. [16]

    Omohundro, S. M. (2008). The basic AI drives. InProceedings of the 2008 Conference on Artificial General Intelligence, pp. 483–492. IOS Press

  9. [17]

    Park, J. S., J. O’Brien, C. J. Cai, M. R. Morris, P . Liang, and M. S. Bernstein (2023). Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST ’23, New York, NY, USA. Associati...

  10. [18]

    (2019).Human Compatible: Artificial Intelligence and the Problem of Control

    Russell, S. (2019).Human Compatible: Artificial Intelligence and the Problem of Control. Penguin

  11. [19]

    Thornley, E. (2023). There are no coherence theorems. AI Alignment Forum. https://www.alignmentforum. org/posts/yCuzmCsE86BTu9PfA/there-are-no-coherence- theorems

  12. [20]

    Thorstad, D. (2024). What power-seeking theorems do not show.Global Priorities Institute Working Paper Series. GPI Working Paper No. 27-2024

  13. [21]

    Smith, R

    Turner, A., L. Smith, R. Shah, A. Critch, and P . Tadepalli (2021). Optimal policies tend to seek power. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P . Liang, and J. W . Vaughan (Eds.),Advances in Neural Information Pro- cessing Systems, Volume 34, pp. 23063–23074. Curran Asso...

  14. [22]

    Turner, A. and P . Tadepalli (2022). Parametrically retargetable decision- makers tend to seek power. In S. Koyejo, S. Mohamed, A. Agarwal, D. Bel- grave, K. Cho, and A. Oh (Eds.),Advances in Neural Information Processing

  15. [23]

    Wang, G., Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anand- kumar (2023). Voyager: An open-ended embodied agent with large lan- guage models. arXiv: 2305.16291v2[cs.AI]

  16. [24]

    Yudkowsky, E. (2015). Sufficiently optimized agents appear coherent. Arbital. https://arbital.com/p/optimized_agent_appears_coherent/. 29

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.