REVIEW 5 minor 24 references
Will artificial agents pursue power by default?
T0 review · 0 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Power is a convergent AI goal—but only an incomplete one.
desk verdict A careful formal treatment of power-seeking that delivers a new incompleteness result and honestly disclaims its own limits. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a finite decision tree whose root offers a set of lotteries over final outcomes, with an agent modeled as an expected utility maximizer whose utility function is drawn from a distribution that is symmetric under permutations of outcomes, regular in the sense of assigning nonzero probability to every open region of utility assignments, and almost surely non-uniform. Power is defined set-theoretically on the collections of lotteries available at subtrees: inclusion ($L(t_j) \subseteq L(t_i)$), interiority ($L(t_j)$ inside the interior of $L(t_i)$), and their permutation-invariant forms ($\pi(L(t_j)) \subseteq L(t_i)$ and $\pi(L(t_j))$ inside the interior of $L(t_i)$). Instrumental convergence is defined by whether a random agent is at least as likely, strictly more likely, or certain to choose the action leading to the higher-ranked subtree. The positive proofs use convexity and the hyperplane separation theorem, while Proposition 5 works by constructing a pair of subtrees whose relative attractiveness reverses under nearly binary utility distributions, so no distribution-independent total ranking can exist.
What would settle it
Exhibit a total reflexive relation on finite decision trees that is weakly instrumentally convergent and whose strict part is strictly instrumentally convergent while satisfying the paper's symmetry, regularity, and almost-sure non-uniformity assumptions; Proposition 5 asserts that no such relation exists.
Extended reading notes
Core claim
The paper's central claim is that power, formalized as the ability to achieve a wider set of probability distributions over outcomes, is a convergent instrumental goal in a nontrivial but limited sense. Four power relations are defined on decision trees: inclusion, where one tree's available lotteries contain another's; interiority, where one tree's lottery set contains another's in its interior; and their permutation-invariant analogues. The first is weakly instrumentally convergent, the second is absolutely instrumentally convergent, the third is weakly but not strictly convergent, and the fourth is strictly but not absolutely convergent. The negative result, Proposition 5, says that no total reflexive relation on decision trees can be weakly instrumentally convergent while its irreflexive part is strictly instrumentally convergent, so any complete ranking of options by power must either depend on substantive assumptions about the distribution of an agent's goals or fail to predict choice. The author concludes that the claim that power is a convergent instrumental goal contains an element of truth but has limited predictive utility, except for agents who can approach absolute or near-absolute power.
Load-bearing premise
The paper treats total ignorance of an agent's final goals as a probability distribution over utility functions that treats all outcomes symmetrically, never rules out any utility function, and almost never assigns the same utility to every outcome; if real AI training produces goal distributions with systematic biases or correlations, the probability claims can fail.
Editorial extensions
If this is right
- An agent with randomly generated final goals is more likely than not to choose an option that gives it a strictly larger set of available lotteries, and is almost certain to choose an option whose lottery set contains the alternative in its interior.
- Survival, information acquisition, and generic resource accumulation are vindicated as convergent instrumental goals in the special cases where they expand, or interior-expand, the set of available lotteries.
- No power ranking derived purely from ignorance of final goals can be complete: there are pairs of subtrees where which option a random agent prefers depends on features of the utility distribution beyond symmetry, regularity, and almost-sure non-uniformity.
- When an agent can approach absolute power, the incompleteness is less damaging, because near-absolute power is near-maximally attractive for almost all non-uniform expected utility maximizers.
- In a multipolar world where no agent can approach absolute control and power-seeking is costly, the fact of instrumental convergence gives relatively little basis for predicting that agents will seek power.
Reading between the lines
- Inference: the incompleteness result pushes the practical question onto the actual training distribution; quantitative predictions about power-seeking require modeling the specific process that generates an agent's utility function, not just modeling ignorance of its goals.
- Inference: the framework suggests a testable prediction: agents trained with broad goals in rich environments will show stronger power-seeking when they can achieve near-deterministic control, and weaker power-seeking when they face multipolar competition with costly failures.
- Inference: because the paper's notion of power is not inherently rivalrous, in zero-sum settings the convergent goal may translate into disempowering other agents; this connection, left open by the paper, could be formalized by adding conflicts among multiple agents.
- Inference: if real agents have non-consequentialist preferences, absolute instrumental convergence would fail, but weak and strict convergence might still hold under ignorance if intrinsic aversions to power-seeking are no more probable than intrinsic attractions to it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops a decision-theoretic formalization of instrumental convergence and power-seeking. The author defines a "random agent" as an expected-utility maximizer whose utility function is drawn from a symmetric, regular, and almost surely non-uniform distribution, and defines weak, strict, and absolute instrumental convergence as comparative likelihood properties of relations on decision trees. Four notions of power are considered: inclusion (≻p1), interiority (≻p2), permuted inclusion (≻p3), and permuted interiority (≻p4). Propositions 1–4 establish that these relations are convergent in the appropriate senses, with ≻p2 absolutely convergent and ≻p4 strictly but not absolutely convergent. Proposition 5 shows that no total reflexive relation on decision trees can be weakly instrumentally convergent with a strictly convergent irreflexive part, implying that power rankings are either incomplete or non-predictive without distributional information. Section 6 argues informally that absolute or near-absolute power is a more predictive target, and Section 7 explicitly disclaims any defense of the orthogonality thesis or alignment difficulty.
Significance. The paper makes a valuable and original contribution to the instrumental-convergence debate. It provides a general framework that avoids the specific RL assumptions of earlier power-seeking theorems, and it cleanly separates the logical relations among different notions of convergence. The positive results (Propositions 1–4) are elementary but nontrivial, and the negative result (Proposition 5) is a genuinely non-circular limitation result that addresses the predictive-utility question raised by critics such as Gallow and Thorstad. The paper is also commendably explicit about the limits of its framework, particularly the dependence on the "random agent" prior and the absence of any defense of the orthogonality thesis. These features make the paper suitable for a broad AI-safety or philosophy audience.
minor comments (5)
- [§5, Proposition 5] The proof of Proposition 5 needs an explicit construction of regular distributions that approximate the two point-mass distributions. The sentence 'they can be approximated arbitrarily closely be distributions that do' asserts the key step without demonstration; a concrete construction (e.g., a mixture of the symmetric orbit of the point mass with a small full-support continuous component) would make the proof complete.
- [Appendix, Proposition 4] The proof of Proposition 4 is compressed at the final step. To establish strict instrumental convergence, the author should explicitly state that any utility function satisfying u(π(t_j)) > u(t_j) also satisfies u(t_i) > u(t_j), because π(L(t_j)) ⊆ L(t_i), and that the positive-measure set of functions preferring t_j to π(t_j) and t_i to t_j breaks the symmetry equality between the two orderings.
- [§2] The definition of 'interior' for sets of lotteries should be clarified as relative interior with respect to the probability simplex, since a set of lotteries is a subset of a lower-dimensional affine space and may have empty Euclidean interior.
- [§4, Proposition 1 proof] In the proof of Proposition 1, the sentence 'if there is an expected-utility-maximizing lottery in L(t_j), that same lottery is in L(t_i), so the agent is equally likely to choose ai or aj' is potentially misleading because such a lottery always exists; the intended condition is that the maximal expected utility is equal in the two subtrees. Please rephrase.
- [Abstract] In the abstract, 'aconservative' should be 'a convergent'.
Circularity Check
No significant circularity: the formal results are valid consequences of stated definitions, the near-tautological positive results are explicitly acknowledged, and the negative result is independent.
full rationale
The paper's central claims are theorems derived from explicitly stated definitions rather than empirical predictions, so the derivation chain is self-contained. The only step that could be viewed as definitional is Power 1/Power 2, where 'more power' is defined as a larger set of available lotteries, making the preference for more power a near-immediate consequence of expected-utility maximization. However, the author transparently acknowledges this by quoting Gallow's tautology objection and noting that Propositions 1 and 2 are 'formally fairly elementary' rather than substantive empirical discoveries. The paper does not present these as predictions about real agents; instead, it emphasizes their limited predictive utility and proves an independent negative result (Proposition 5) showing that no total instrumental-convergence ranking exists under the stated ignorance assumptions. Section 7 explicitly disclaims any defense of the orthogonality thesis or alignment difficulty, so the 'random agent' prior is not smuggled in as a conclusion. There is no load-bearing self-citation, fitted parameter renamed as a prediction, or imported uniqueness theorem. The derivation is therefore not circular.
Assumptions & free parameters
assumptions (5)
- domain assumption Agents are expected utility maximizers with a utility function over final outcomes.
- domain assumption The set of outcomes O is finite with cardinality at least 3.
- domain assumption A random agent's utility function is drawn from a symmetric, regular, and almost surely non-uniform distribution.
- domain assumption Agents are certain their utility will not change over time and will continue to maximize expected utility at future choice nodes.
- standard math Standard results from convex analysis, including the hyperplane separation theorem, are valid.
Cite this review
Pith. "Pith review of Will artificial agents pursue power by default?." pith.science (2026). https://pith.science/paper/DIHDZZOK
@misc{pith2026250606352,
author = {Pith},
title = {Pith review of: Will artificial agents pursue power by default?},
year = {2026},
howpublished = {\url{https://pith.science/paper/DIHDZZOK}},
note = {Machine review of arXiv:2506.06352}
}
read the original abstract
Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that is useful for a wide range of final goals. Others have recently expressed skepticism of these claims. This paper aims to formalize the concepts of instrumental convergence and power-seeking in an abstract, decision-theoretic framework, and to assess the claim that power is a convergent instrumental goal. I conclude that this claim contains at least an element of truth, but might turn out to have limited predictive utility, since an agent's options cannot always be ranked in terms of power in the absence of substantive information about the agent's final goals. However, the fact of instrumental convergence is more predictive for agents who have a good shot at attaining absolute or near-absolute power.
Reference graph
Works this paper leans on
-
[1]
Bales, A., W . D’Alessandro, and C. D. Kirk-Giannini (2024). Artificial in- telligence: Arguments for catastrophic risk.Philosophy Compass 19(2), e12964
work page 2024
-
[2]
Benson-Tilsen, T . and N. Soares (2016). Formalizing convergent instru- mental goals. InWorkshops at the Thirtieth AAAI Conference on Artificial Intelligence
work page 2016
-
[3]
Bostrom, N. (2003). Ethical issues in advanced artificial intelligence. In I. Smit, W . Wallach, and G. E. Lasker (Eds.),Cognitive, Emotive and Ethical Aspects of Decision Making in Humans and in Artificial Intelligence, Vol- ume 2, pp. 12–17. International Institute for Advanced Studies in Systems Research and Cybernetics
work page 2003
-
[4]
Bostrom, N. (2012). The superintelligent will: Motivation and instrumental rationality in advanced artificial agents.Minds and Machines 22(2), 71–85
work page 2012
-
[5]
(2014).Superintelligence: Paths, Dangers, Strategies
Bostrom, N. (2014).Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press
work page 2014
-
[6]
Boyd, S. P . and L. Vandenberghe (2004).Convex Optimization. Cambridge University Press. 27
work page 2004
-
[7]
Carlsmith, J. (2022). Is power-seeking AI an existential risk? arXiv: 2206.13353v1[cs.CY]
arXiv 2022
-
[8]
Dung, L. (2024). The argument for near-term human disempowerment through AI.AI & Society, 1–14
work page 2024
Show all 24 references
-
[9]
Gabriel, I. (2020). Artificial intelligence, values, and alignment.Minds and Machines 30(3), 411–437
2020
-
[10]
Gallow, J. D. (2024). Instrumental divergence.Philosophical Studies, 1–27
2024
-
[11]
Goldstein, S. and C. D. Kirk-Giannini (2023). Language agents reduce the risk of existential catastrophe.AI & Society, 1–11
2023
-
[12]
Good, I. J. (1966). On the principle of total evidence.British Journal for the Philosophy of Science 17(4), 319–321
1966
-
[13]
Mazeika, and T
Hendrycks, D., M. Mazeika, and T . Woodside (2023). An overview of catas- trophic AI risks. arXiv: 2306.12001v6[cs.CY]
2023 arXiv
-
[14]
Krakovna, V . and J. Kramar (2023). Power-seeking can be probable and predictive for trained agents. arXiv: 2304.06528[cs.AI]
2023 arXiv
-
[15]
Chan, and S
Ngo, R., L. Chan, and S. Mindermann (2023). The alignment problem from a deep learning perspective. arXiv: 2209.00626v5[cs.AI]
2023 arXiv
-
[16]
Omohundro, S. M. (2008). The basic AI drives. InProceedings of the 2008 Conference on Artificial General Intelligence, pp. 483–492. IOS Press
2008
-
[17]
Park, J. S., J. O’Brien, C. J. Cai, M. R. Morris, P . Liang, and M. S. Bernstein (2023). Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST ’23, New York, NY, USA. Associati...
2023
-
[18]
(2019).Human Compatible: Artificial Intelligence and the Problem of Control
Russell, S. (2019).Human Compatible: Artificial Intelligence and the Problem of Control. Penguin
2019
-
[19]
Thornley, E. (2023). There are no coherence theorems. AI Alignment Forum. https://www.alignmentforum. org/posts/yCuzmCsE86BTu9PfA/there-are-no-coherence- theorems
2023
-
[20]
Thorstad, D. (2024). What power-seeking theorems do not show.Global Priorities Institute Working Paper Series. GPI Working Paper No. 27-2024
2024
-
[21]
Smith, R
Turner, A., L. Smith, R. Shah, A. Critch, and P . Tadepalli (2021). Optimal policies tend to seek power. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P . Liang, and J. W . Vaughan (Eds.),Advances in Neural Information Pro- cessing Systems, Volume 34, pp. 23063–23074. Curran Asso...
2021
-
[22]
Turner, A. and P . Tadepalli (2022). Parametrically retargetable decision- makers tend to seek power. In S. Koyejo, S. Mohamed, A. Agarwal, D. Bel- grave, K. Cho, and A. Oh (Eds.),Advances in Neural Information Processing
2022
-
[23]
Wang, G., Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anand- kumar (2023). Voyager: An open-ended embodied agent with large lan- guage models. arXiv: 2305.16291v2[cs.AI]
2023 arXiv
-
[24]
Yudkowsky, E. (2015). Sufficiently optimized agents appear coherent. Arbital. https://arbital.com/p/optimized_agent_appears_coherent/. 29
2015
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.