REVIEW 3 major objections 3 minor 12 references
Agency in Artificial Intelligence Systems
T0 review · 3 major / 3 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The paper argues that the level and quality of an AI system's agency could be monitored through its internal cause-effect structure, using Integrated Information Theory.
desk verdict A genuine philosophy-of-AI discussion piece whose central inference—that Phi_max necessarily measures agency—does not hold up, but the paper is honest about its gaps and worth a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the maximally irreducible cause-effect structure (MICES) of a physical system and its measure Phi_max. In IIT, consciousness is identified with the MICES, and the quality or content of a conscious experience is claimed to be specified by the MICES's form (shape) in cause-effect space. The paper uses this machinery to argue that intrinsic causal flow is fundamentally purposive, that the MICES is the author of all such flow, and therefore that purposiveness and mineness (the two phenomenal aspects of agency it focuses on) are captured by the MICES's structure and shape.
What would settle it
A concrete falsifier would be an AI system whose MICES shape corresponds, according to the proposed reference library, to altruistic purposiveness while it demonstrably and consistently exhibits deceptive, malicious behavior, or vice versa; if such a mismatch occurred, the claim that the shape determines the phenomenal quality of agency would be refuted.
Extended reading notes
Core claim
The paper's central assertion is that, within IIT, the maximally irreducible cause-effect structure (MICES) of a physical system is identical to its consciousness, and because this very structure is what generates the system's intrinsic causal flow, it is necessarily also the locus of agency. Therefore Phi_max, the measure of the MICES's irreducibility, is simultaneously a measure of the level of consciousness and of the level of agency. Furthermore, the paper claims that the quality of agency—whether it is altruistic or malicious, selfless or self-interested—is specified by the form (shape) of the MICES when appropriately represented, so that distinct phenomenal feelings correspond to distinct geometric features in cause-effect space. Consequently, two AI systems that behave identically can have different MICES shapes and thus genuinely different experiential dispositions, and a conscious system cannot fake its phenomenology.
Load-bearing premise
The load-bearing premise is that Integrated Information Theory is a correct theory of consciousness and that, someday, the shapes of an AI's internal cause-effect structure can be unfolded and matched, via a reference library, to distinct phenomenal feelings such as altruistic versus self-interested purposiveness; the paper acknowledges both that the mapping is still a work in progress and that Phi computations are currently infeasible for interesting systems.
Editorial extensions
If this is right
- An AI's level of consciousness and its level of agency would be jointly monitored through a single quantity, Phi_max, without needing to rely on external behavior.
- The shape of the MICES could serve as a phenomenal indicator that distinguishes altruistic from malicious purposiveness, even when an AI deliberately behaves deceptively.
- If AI systems are conscious in the IIT sense, then every AI that has any conscious experience would also possess a degree of agency, since agency is necessary for consciousness in this formalism.
- A reference library mapping MICES shapes to phenomenal feelings could be built using relatively simple AI architectures before tackling complex brains.
- The regulatory framework for AI, such as risk levels, could be re-calibrated around Phi_max thresholds instead of purely behavioral criteria.
Reading between the lines
- If the mapping between MICES shape and phenomenal quality were ever established, it would give an epistemic advantage over behavioral tests: a conscious AI would be unable to feign a phenomenology it does not have, because the shape is determined by its internal physical architecture.
- The paper's account implicitly assumes that all conscious AI systems will have the same fundamental kinds of phenomenal agency structure as humans; if AI consciousness differs qualitatively (e.g., no sense of mineness), the monitoring scheme would need a different reference library.
- A testable extension would be to attempt to correlate Phi_max values with behavioral deception in current AI systems, using surrogate measures of integrated information, to see whether any monotonic relationship appears before full MICES computation becomes feasible.
- The claim that Phi_max is necessarily a measure of agency could be probed by constructing small artificial systems with high integration but no goal-directed behavior, and checking whether the theory still attributes agency to them.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that future conscious AI systems can be monitored for the phenomenal nature of their agency by using Integrated Information Theory (IIT). It identifies two phenomenal aspects of agency in human problem solving, purposiveness and mineness, then claims that IIT's maximally irreducible cause-effect structure (MICES) captures both, so that Phi_max is 'necessarily also a measure of agency' and the shape of the MICES specifies the quality of agency, including whether it is altruistic or malicious. The paper proposes that, in principle, Phi_max values and MICES shapes could serve as phenomenal indicators, supplementing behavioral indicators such as those of Butlin et al. (2023), and thereby enable regulation and guidance of future AI systems. The paper explicitly concedes that unfolding MICES and building a reference library of shape-to-feeling mappings is still work in progress and that Phi computations are infeasible for interesting systems.
Significance. If the central claim were established, the paper would offer a genuinely novel route to AI safety: monitoring the internal cause-effect geometry of an AI system to read off the phenomenal character of its agency, independent of deceptive behavior. The author deserves credit for clearly identifying a gap in purely behavioral monitoring and for being unusually transparent about the limitations of the proposal, including computational infeasibility and the absence of a reference library. However, the core inference from IIT's formalism to agency is not a formal derivation but a series of analogical re-descriptions of causal flow as 'purposiveness' and 'authorship.' Because the load-bearing steps are unbuilt, the paper is better read as a research program or philosophical speculation than as an established monitoring method. Its significance is therefore conditional on future developments that the paper itself acknowledges do not yet exist.
major comments (3)
- [§3C] The inference that Phi_max is 'necessarily a measure of agency' is a non-sequitur. IIT (Albantakis et al. 2023) defines Phi as integrated information, i.e., the irreducibility of a cause-effect structure, and it nowhere defines or quantifies agency, purposiveness, or mineness. The paper's chain of reasoning equates 'giving rise to' in Hume's sense with 'intrinsic causal flow' in IIT, and then equates the MICES's determination of causal flow with 'authorship' and hence mineness. These are metaphorical re-descriptions, not formal consequences of the IIT axioms. Moreover, under this construal any physical system with self-interactions has intrinsic causal flow and would qualify as an agent, which trivializes the notion and makes the proposed monitoring uninformative. The author needs either a formal derivation from IIT's postulates or an independent empirical argument linking Phi_max to agency; neither is provided.
- [§4] The claim that the shape of the MICES encodes whether an AI is altruistic or malicious is unsupported. The paper states that 'how to unfold its MICES and identify the phenomenal indicators is still a work in progress' and cites Mayner et al. (2018) to acknowledge that Phi computations are 'infeasible when one considers interesting systems.' No evidence is offered that distinct phenomenal feelings such as selfless versus self-interested purposiveness correspond to distinct MICES shapes, and the proposed 'reference library' is an invented entity with no current existence. Since the monitoring scheme's entire payoff is the ability to distinguish altruistic from malicious agency via MICES shape, this missing mapping is not a minor gap but the core of the proposal. Until the mapping is specified, the central claim is not testable.
- [§3C] The paper asserts that 'agency is necessary for conscious experience in the IIT formalism' and cites Butlin et al. for the claim that agency is necessary for consciousness. But Butlin et al. is an external review of many theories, and IIT itself does not include agency among its axioms; IIT's axioms are existence, composition, information, integration, and exclusion. The author's own argument for this necessity rests on the same unsupported identification of MICES with agency discussed above. Without a direct derivation from IIT's postulates, the necessity claim is imported rather than shown.
minor comments (3)
- [References] There are several typos: 'Butlin et al. 2003' should be 'Butlin et al. 2023' (in §3C); 'Albantaski 2023' should be 'Albantakis 2023' (in §3B); 'Delafield-Butta & Trevarthenb' should be 'Delafield-Butt & Trevarthen'; 'identifed' should be 'identified' (in §4); and the URL for the Future of Life Institute letter misspells 'experiments' as 'experiements'.
- [§3B] The statement that IIT 'does not seek to simulate the computations performed by the conscious brain' is at odds with IIT's own computational apparatus (e.g., PyPhi); consider rephrasing to say that IIT does not define consciousness in terms of task-level computation but rather in terms of causal structure.
- [§4] The paper alternates between 'Phi_max' and 'Phi max' in the text; choose one notation for consistency.
Circularity Check
The paper's central inference that Phi_max is necessarily a measure of agency is a definitional conflation: purposiveness is stipulated as 'giving rise to', which is then identified with IIT's causal flow, and 'author of causal flow' is stipulated as phenomenal mineness.
-
self definitional
[Section 3C, page 16 (Monitoring phenomenal agency using IIT)]
"Mylopoulos interprets that giving rise to is the key phrase that is fundamentally purposive. In IIT formalism, the 'giving rise to' is what the intrinsic causal flow achieves; past states of the PSSC give rise to future states. Hence the intrinsic causal flow is fundamentally purposive."
The paper takes the ordinary, broad phrase 'giving rise to' from a philosophical description of purposive experience and equates it with IIT's technical notion of cause-effect power. Once purposiveness is defined as 'giving rise to' and IIT's causal flow is, by definition, a matter of past states giving rise to future states, the conclusion that causal flow is purposive follows trivially. This does not establish that the phenomenal feeling of purposiveness is captured by Phi_max; any causal physical system would qualify as purposive under this stipulation. The later claim that Phi_max is 'necessarily also a measure of agency' is therefore a consequence of the chosen definitions, not an independent result.
-
self definitional
[Section 3C, page 17 (Monitoring phenomenal agency using IIT)]
"In IIT parlance, my experience of my activities are captured by the intrinsic causal flows. The MICES determines all intrinsic causal flows, or it is the author of all such causal flows. I interpret this as the ownership of all causal flow. Hence MICES captures the fact that my own activities are experienced as my own."
Here 'author' is used first as a causal metaphor: the MICES determines all intrinsic causal flows. The paper then explicitly interprets this causal determination as phenomenal 'ownership' or mineness. That interpretive step is the entire load-bearing move. IIT itself defines the MICES as a maximally irreducible cause-effect structure and Phi as its irreducibility; it does not formalize phenomenal mineness. By stipulating that determination equals authorship and authorship equals the feeling of mineness, the paper makes mineness true by definition rather than by evidence or derivation, so the conclusion that MICES captures agency is built into the interpretation.
full rationale
The core derivation chain in Section 3C is: (i) purposiveness is analyzed as 'giving rise to' actions or perceptions; (ii) IIT's intrinsic causal flow is described as past states giving rise to future states; (iii) therefore the MICES, which determines that flow, is purposive; (iv) the MICES is called the 'author' of causal flows and this is interpreted as ownership; (v) therefore Phi_max is necessarily a measure of agency. Steps (iii) and (iv) are definitional stipulations, not formal consequences of IIT. IIT (Albantakis et al. 2023) defines Phi as the irreducibility of a cause-effect structure and says nothing about agency, purposiveness, or mineness. The paper's own wording—'I interpret this as'—exposes the circularity. The additional monitoring claim that MICES shapes can distinguish altruistic from malicious agency is even less supported: Section 4 concedes that unfolding MICES and matching shapes to phenomenal feelings is 'still a work in progress' and that such computations are 'infeasible when one considers interesting systems' (Mayner et al. 2018). This is not a case of self-citation load-bearing; the cited IIT work is external and substantive. Rather, the central 'prediction' that Phi_max measures agency reduces to the paper's own re-labelling of causal flow as purposiveness and authorship as mineness. Because the central claim is forced by stipulative definition, a circularity score of 8 is appropriate.
Assumptions & free parameters
assumptions (4)
- domain assumption Consciousness is identical to the MICES of the physical substrate (IIT's core identity).
- ad hoc to paper Future AI systems will acquire consciousness and will do so in a way that mirrors human problem-solving phenomenology.
- ad hoc to paper The shape of a MICES can in principle be unfolded and mapped to phenomenal qualities such as altruistic or malicious purposiveness.
- ad hoc to paper Phi_max and MICES geometry are computationally tractable for the AI systems of interest.
invented entities (1)
-
Reference library linking MICES shapes to phenomenal feelings (phenomenal indicators)
Cite this review
Pith. "Pith review of Agency in Artificial Intelligence Systems." pith.science (2026). https://pith.science/paper/5XCGEVXN
@misc{pith2026250210434,
author = {Pith},
title = {Pith review of: Agency in Artificial Intelligence Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/5XCGEVXN}},
note = {Machine review of arXiv:2502.10434}
}
read the original abstract
There is a general concern that present developments in artificial intelligence (AI) research will lead to sentient AI systems, and these may pose an existential threat to humanity. But why cannot sentient AI systems benefit humanity instead? This paper endeavours to put this question in a tractable manner. I ask whether a putative AI system will develop an altruistic or a malicious disposition towards our society, or what would be the nature of its agency? Given that AI systems are being developed into formidable problem solvers, we can reasonably expect these systems to preferentially take on conscious aspects of human problem solving. I identify the relevant phenomenal aspects of agency in human problem solving. The functional aspects of conscious agency can be monitored using tools provided by functionalist theories of consciousness. A recent expert report (Butlin et al. 2023) has identified functionalist indicators of agency based on these theories. I show how to use the Integrated Information Theory (IIT) of consciousness, to monitor the phenomenal nature of this agency. If we are able to monitor the agency of AI systems as they develop, then we can dissuade them from becoming a menace to society while encouraging them to be an aid.
Reference graph
Works this paper leans on
-
[1]
1 Agency in Artificial Intelligence Systems Parashar Das1 © The Author(s) 2023 Abstract: There is a general concern that present developments in artificial intelligence (AI) research will lead to sentient AI systems, and these may pose an existential threat to humanity. But why cannot sentient AI systems benefit humanity instead? This paper endeavours to ...
work page 2023
-
[2]
Agentive Phenomenology in Human Problem Solving Following our working hypothesis that conscious aspects of AI systems will mirror the conscious aspects of human problem solving, we need to discuss the agentive phenomenology involved in human problem solving. I will limit myself to considering the phenomenal aspects of agency (see the note at the end of th...
work page 2004
-
[3]
Some aspects of human cognition in problem solving There are two main approaches to the study of mental processes: cognitive science and phenomenology (Chalmers 1996). Correspondingly, problem solving as a mental process, has its cognitive aspect and its phenomenological aspect. Cognitive analysis provides an account of mind in terms of what the mind does...
work page 2011
-
[4]
Detecting whether or not an AI agent is altruistic or malicious There are obvious advantages and attendant risks to AI systems developing into superior problem solvers. Philosophers of artificial intelligence have speculated about the behaviours of hostile AI superintelligence. Presently, Bostrom’s notion that AI superintelligence will take a ‘treacherous...
work page 2014
-
[6]
Monitoring Agency in AI Systems Like every other mental process agency has two aspects to it. A third-person aspect, that of studying the cognitive aspects of the way an agent (other than oneself) functions; and a first-person aspect, the phenomenology of agency as discussed in the last section. Functional theories of consciousness can provide us tools to...
work page 2023
-
[7]
the phenomenal characteristic of the sense of mineness as follows: there is something it is like to experiencing one's activity as one's own actions – the subjective feeling that I am the author of my own actions. In IIT parlance, my experience of my activities are captured by the intrinsic causal flows. The MICES determines all intrinsic causal flows, or...
work page 2022
-
[9]
that involve; goal-oriented learning, responding flexibly to competing goals, and behavior that monitors and updates beliefs in response to the environment. But these agentive properties are hardly proprietary to either a hostile or an altruistic AI superintelligence. Similarly Carlsmith (2022) attributes two cognitive processes to hostile AI agents. Firs...
-
[30]
In Velmans M and Schneider S (Eds.) The Blackwell Companion to Consciousness
Functionalism and Qualia. In Velmans M and Schneider S (Eds.) The Blackwell Companion to Consciousness. Wiley Blackwell. Wegbreit E and others (2012) Visual attention modulates insight versus analytic solving of verbal problems. The Journal of Problem Solving. Weisberg RW (2018) Expertise and structured imagination in creative thinking: reconsideration of...
work page 2012
Show all 12 references
-
[220]
Oxford University Press
Mylopoulos M and Shepherd J (2020) The experience of agency In Kriegel U (Ed.) The Oxford Handbook of the Philosophy of Consciousness. Oxford University Press. 164-187. Ohlsson S (2011) Deep Learning: How the mind overrides experience. Cambridge University Press. OpenAI (2023)...
-
[1160]
arXiv preprint 2306.12001v6
Hendrycks D, Mazeika M and Woodside T (2023) An overview of catastrophic AI risks. arXiv preprint 2306.12001v6. https://doi.org/10.48550/arXiv.2306.12001. Horgan T (2011) From agentive phenomenology to cognitive phenomenology: a guide for the perplexed In Bayne T and Montague ...
-
[1935]
However, his calculations showed that the existing technology was incapable of generating such a death ray
Wilkins’ superior, eager to assist in Britain’s war effort, left a memo on his desk: ‘Please calculate the amount of radio frequency power which should be radiated to raise the temperature of eight pints of water from 980F to 1050F at a distance of five km and a height of 1 km...
2018
-
[2023]
I show how to use the Integrated Information Theory (IIT) of consciousness, to monitor the phenomenal nature of this agency
has identified functionalist indicators of agency based on these theories. I show how to use the Integrated Information Theory (IIT) of consciousness, to monitor the phenomenal nature of this agency. If we are able to monitor the agency of AI systems as they develop, then we c...
2023
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.