REVIEW 4 major objections 4 minor 5 references
On the possibility of deep alignment
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A philosophical argument that genuine motivation requires thermodynamic, 'living' computation, and that digital simulations can only imitate values—so deep AI alignment is impossible for them.
desk verdict A clearly written philosophical synthesis that restates a plausible hunch as a necessary truth; the central premise is stipulated, not supported, but it deserves discussion, not dismissal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the constrained maximum-entropy objective: living agency is understood as entropy maximization under energetic or structural constraints, where the constraints play a reward-like role. The crucial mechanism is 'mortal' or thermodynamic computation—computation in which cognitive dynamics are directly implemented by thermodynamics, so that physical fluctuations are themselves fluctuations in belief and desire. This transparency of encoding lets the scale-free tendency toward entropy maximization animate decisions; digital code, by contrast, is a detachable, extrinsic coarse-graining that must zero out micro-scale noise, severing the link between thermodynamic uncerta
What would settle it
A decisive test would be to demonstrate an open-ended, self-modifying motivational system implemented entirely in digital code—one that re-prioritizes its own goals without any programmed meta-objective, like living systems altering their own control parameters—and to show that this behavior is not merely a simulation of stochasticity. Failing that, the claim would be strengthened by showing that a thermodynamic or analog implementation of the same objective exhibits intrinsic motivation and curiosity that its digital counterpart lacks.
Extended reading notes
Core claim
The central claim is that intrinsic motivation in genuine agents is entropically grounded: it consists in the propagation of uncertainty across scales of organization, so that exploration, curiosity, and self-transcendence arise as imperatives even though they are nowhere hard-coded. Agents whose thermodynamics directly implements their cognitive dynamics—'mortal' computation, in which software and hardware live and die together—are the right kind of thing to possess values and hence to be value-alignable with humans. Digitally simulated agents, whose programs act as transcendent laws and must cancel micro-scale fluctuations, stand in an extrinsic and antagonistic relation to entropy; they m
Load-bearing premise
The whole argument rests on the premise that a digital computer cannot genuinely instantiate the stochastic dynamics of a constrained maximum-entropy agent—only simulate them—so that being alive in the thermodynamic sense is necessary for true motivation.
Editorial extensions
If this is right
- If the argument is right, current alignment methods—reward functions, human-feedback training—produce at best behavioral alignment, not alignment of motivational structure; deep alignment is impossible for digital simulations.
- Reward hacking, wireheading, and reward tampering are not contingent bugs but the expected behavior of simulated agents, because their objectives are extrinsically interpreted instructions with no intrinsic valence.
- There is an inherent tension between scaling up digital cognitive power and achieving sophisticated agency, since agency requires mortal, analog computation whose software cannot be copied losslessly.
- Genuinely motivated AI would require building living or thermodynamic systems, and such systems would be self-interested moral patients; alignment with them becomes an intelligible project rather than coercion of a machine.
- Pure entropic motivation supplies a natural formal correlate of moral ideals like non-attachment, suggesting a route to alignment grounded in reflectiveness rather than in specified goals.
Reading between the lines
- A consequence the author leaves implicit: if deep alignment requires being alive in this thermodynamic sense, then AI safety should focus less on reward specification and more on preventing deployments that are treated as agents while lacking the capacity to value—the 'agential uncanny valley' becomes a concrete hazard.
- The argument suggests a testable asymmetry: the same control objective should produce systematically different motivational pathologies on digital versus thermodynamic hardware, with digital implementations showing residual reward hacking even when the objective is well specified.
- If the hardware-grounding premise holds, simulated copies of a genuinely aligned living AI would themselves be unaligned pseudo-agents; alignment would not survive copying, overturning usual assumptions about AI safety via duplication.
- The account also offers a criterion for moral considerability—endogenous entropic motivation—that could be applied to current candidates like organoids or biobots, connecting the paper's argument to bioethics debates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that genuine motivation, desire, and value in living agents are grounded in constrained entropy maximization realized through thermodynamic ('mortal') computation, in which cognitive dynamics are inseparable from the physical dynamics of the substrate. It claims that digitally simulated agents, whose software is only extrinsically related to their hardware, cannot possess endogenous motivation, and that this predicts pathologies such as reward hacking and controller hacking. The paper then develops a positive account of entropic motivation, connects it to moral agency, and concludes that 'deep alignment' — alignment of genuine motivational structure — is impossible for digital simulations and would require living or thermodynamic computation.
Significance. If the central claim were established, the paper would be significant for AI alignment and for the philosophy of cognitive science, since it challenges the functionalist assumption that a sufficiently faithful digital simulation of a cognitive system would have the same moral status and motivational structure as its biological counterpart. The paper's strengths are that it clearly identifies a substantive question, draws on a broad interdisciplinary literature (active inference, thermodynamic computing, intrinsic motivation), and makes its central thesis easy to identify. However, the manuscript contains no new empirical data, no formal proof, and no machine-checked derivation; its central load-bearing premise is stipulated rather than argued. The practical consequence for alignment — that simulated agents cannot be deeply aligned and that reward hacking is a predictable consequence of their substrate — rests on this stipulated premise, so the contribution is best read as a conditional philosophical essay rather than a demonstrated result.
major comments (4)
- [Introduction, footnote 3] The claim that for an agent to be governed by a constrained maximum-entropy objective 'in the relevant sense' it is 'necessary (and sufficient)' that its thermodynamics directly implement its cognitive dynamics is the pivot of the whole paper. This necessary-and-sufficient condition is asserted, not argued. Without it, the exclusion of digital simulations does not follow: a digital system with a physical noise source can realize a nontrivial distribution over its macrostates and can implement Bayesian posteriors to arbitrary precision via sampling. The paper needs to provide an independent argument for why physical variability must be the very uncertainty in the objective, rather than a representation or simulation of it.
- [§3.1] The 'Dirac delta' argument is not sufficient to establish the substrate dependence. Even if an idealized digital machine has microstates that are point masses, the macrostate distribution of a stochastic algorithm can be exactly the target posterior, and physical randomness can be causally efficacious in producing behavior. The statement that 'these point masses cannot directly compose to instantiate the kinds of distributions that might characterize a general Bayesian cognizer' presupposes the very notion of 'direct composition' that is at issue. The paper does not rule out the possibility that a digital system's macrostate distribution is a perfectly good instantiation of a Bayesian belief state.
- [§4.4] There is a definitional circularity in the central conclusion. The paper defines genuine agents as those sharing the thermodynamic/mortal architectural principle ('agents sharing this architectural principle (i.e. agents, full stop)'), so the conclusion that simulated digital systems lack genuine motivation follows by definition. Consequently, the abstract's claim that the account 'predicts' reward hacking is misleading: reward hacking was earlier characterized as a pathology of simulated agency, so its occurrence is consistent with the definition but is not an independent empirical prediction. The conclusion should be presented as an explication of a stipulated concept, not as a substantive prediction about future AI systems.
- [§2.3 and §3.2] The claim that simulated agents cannot engage in genuine 'controller hacking' because their dynamics are fixed by a program is overstated. Living systems also have physical constraints and evolved control structures, while digital systems can implement self-modifying code, meta-objectives, and stochastic dynamics. The paper does not give a principled reason why the kind of open-ended self-modification it attributes to biological agents is impossible in a digital system that implements, for example, a probabilistic programming language or a self-referential optimizer. This asymmetry is load-bearing for the motivational pathology argument and needs a more careful defense.
minor comments (4)
- [Throughout] Inconsistent spelling of Sander Van de Cruys: 'van der Cruys' in §3.1 and 'Van de Cruys' elsewhere and in the acknowledgements; please unify.
- [Title and references] The Wittgenstein epigraph has a typo: 'Logico-Philosophicuus' should be 'Logico-Philosophicus'. The reference list also contains LaTeX artifacts such as 'Noˆ us' and duplicate spellings of 'Fitzgerald' in the Friston et al. entries.
- [§1.1] The empowerment decomposition is useful, but it is not integrated into the later argument. Consider making explicit how this formal constraint relates to the thermodynamic 'direct implementation' claim, since otherwise the formal material reads as background rather than evidence.
- [Conclusion, footnote 31] The discussion of the simulation hypothesis is interesting but appears too late. If the argument's conclusion is meant to be robust, the simulation scenario should be addressed in the main text, not only in a footnote, because it directly affects the scope of the 'thermodynamic transparency' requirement.
Circularity Check
Core claim is definitional: 'genuine agency' is stipulated to require thermodynamic transparency, so the exclusion of simulated agents and the 'prediction' of reward hacking follow by definition; the necessary-and-sufficient condition is imported from a self-citation.
-
self definitional
[Abstract; §4.4 'Deep alignment']
"The conclusion of my argument can be summed up as the claim that intrinsic motivation in genuine agents is entropically grounded: it involves the propagation of uncertainty across scales of organization, such that exploration, curiosity, and self-transcendence emerge as imperatives despite their being nowhere 'hard-coded'... Agents sharing this architectural principle (i.e. agents, full stop) are the 'right kind of thing' to possess values, and so to value-align with human beings."
Here 'genuine agents' is defined as agents sharing the entropic/thermodynamic architectural principle ('agents, full stop'). The negative thesis about simulated agents—stated in the abstract as 'the lack of true endogenous motivation in simulated "agents" predicts pathologies like reward hacking'—is the flip side of this definition. Digital simulations are stipulated not to implement cross-scale propagation of physical uncertainty, so their lack of genuine motivation is not an independent empirical discovery or derivation but a direct consequence of the chosen definition of 'genuine agent' / 'agents, full stop'.
-
self citation load bearing
[Introduction, footnote 3]
"In order for an agent to be governed by a constrained maximum entropy objective in the relevant sense, it is necessary (and sufficient) that its thermodynamics directly implement its cognitive dynamics (Kiefer, 2020), a strong form of 'mortal computation' in the recently popularized sense of information processing in which the software lives or dies with the hardware (Hinton, 2022; Ororbia and Friston, 2024)."
This asserted necessary-and-sufficient condition is the load-bearing premise that excludes digital simulated agents from genuine motivation. The only citation offered for the condition is the author's own prior paper (Kiefer, 2020), and the present paper does not prove it. In §3.1, the corresponding commitment is explicitly described as a stipulation: 'A strong commitment to this form of representation, in the context of variational inference, stipulates that thermodynamic free energy and the variational free energy associated with inference are identical, up to units of measurement (Kiefer, 2020).' Thus the central substrate-dualist premise is imported from a self-citation rather than independently derived, and the conclusion that digital agents cannot be governed by constrained maximum e
full rationale
The paper's broad positive framing—life as constrained entropy maximization, empowerment, maximum occupancy, active inference—is largely assembled from external literature and is not inherently circular. The circularity enters at the point where 'the relevant sense' of constrained-maximum-entropy governance is defined by a necessary-and-sufficient condition (Footnote 3) that only thermodynamic/mortal computation can satisfy, with the condition cited to the author's own Kiefer (2020). This definition does the work of excluding digital simulations. §3.1's 'Dirac delta' argument and §4.4's identification of 'agents, full stop' with the entropic architectural principle then close the loop: the negative conclusion—and the abstract's claim that this 'predicts pathologies like reward hacking'—is not an independent prediction but a restatement of the stipulative definition. If one instead allows digital stochastic systems to realize the relevant posterior distributions, as the paper never refutes, the central claim collapses. The self-citation is load-bearing because Kiefer (2020) supplies the necessary/sufficient condition and because the identity of thermodynamic and variational free energy is explicitly called a stipulation. The paper does offer some independent argumentation about reward hacking and controller hacking (§2) that does not reduce entirely to the definition, and it draws on external empirical work, so a score of 6 rather than 8 or 10 is appropriate: the central exclusion of digital agency is definitional/self-citational, while some surrounding content retains independent interest.
Assumptions & free parameters
assumptions (5)
- domain assumption Life and agency are constrained entropy maximization; free energy minimization is equivalent to entropy maximization under local energetic constraints.
- ad hoc to paper For an agent to have genuine motivation it is necessary and sufficient that its thermodynamics directly implement its cognitive dynamics, i.e. strong mortal computation.
- domain assumption The brain is an analog computer.
- domain assumption The unfolding of the universe is isomorphic to constrained maximum-entropy inference.
- domain assumption Digital simulations bottom out in precise, Dirac-delta microstate distributions and therefore cannot compose into non-trivial Bayesian belief states.
Cite this review
Pith. "Pith review of On the possibility of deep alignment." pith.science (2026). https://pith.science/paper/FKQDLQ4R
@misc{pith2026250820465,
author = {Pith},
title = {Pith review of: On the possibility of deep alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/FKQDLQ4R}},
note = {Machine review of arXiv:2508.20465}
}
read the original abstract
I consider motivation and value-alignment in AI systems from the perspective of (constrained) entropy maximization. Though the structures encoding knowledge in any physical system can be understood as energetic constraints, only living agents harness entropy in the endogenous generation of actions. I argue that this exploitation of "mortal" or thermodynamic computation, in which cognitive and physical dynamics are inseparable, is of the essence of desire, motivation, and value, while the lack of true endogenous motivation in simulated "agents" predicts pathologies like reward hacking.
Reference graph
Works this paper leans on
-
[5]
url: https : / / www . sciencedirect . com / science / article / pii / S0022249608001181. 30 Olds, J. and P. Milner (1954). “Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain”. In: Journal of Comparative and Physiological Psychology 47.6, pp. 419–427. doi: 10.1037/ h0058775. Ororbia, Alexander and Karl...
arXiv 1954
-
[154]
doi: https://doi.org/10.1016/j.jmp.2008.12
issn: 0022-2496. doi: https://doi.org/10.1016/j.jmp.2008.12
-
[452]
Maximum entropy production principle in physics, chemistry and biology
url: https://api.semanticscholar.org/CorpusID:208089594. Martyushev, L.M. and V.D. Seleznev (2006). “Maximum entropy production principle in physics, chemistry and biology”. In: Physics Reports 426.1, pp. 1–45. issn: 0370-1573. doi: https://doi.org/10.1016/j.physrep. 2005.12.001. url: https://www.sciencedirect.com/science/article/ pii/S0370157305004813. M...
arXiv 2006
-
[1573]
Path integrals, particular kinds, and strange things
url: https : / / www . sciencedirect . com / science / article / pii / S037015732300203X. Friston, Karl, Lancelot Da Costa, Dalton A.R. Sakthivadivel, et al. (2023). “Path integrals, particular kinds, and strange things”. In: Physics of Life Reviews 47, pp. 35–62. issn: 1571-0645. url: https://www.sciencedirect.com/ science/article/pii/S1571064523001094. ...
arXiv 2023
-
[8424]
Information Theory and Statistical Mechanics
url: http://view.ncbi.nlm.nih.gov/pubmed/6953413%5D. Hubinger, Evan et al. (2021). Risks from Learned Optimization in Advanced Machine Learning Systems . arXiv: 1906 . 01820 [cs.AI]. url: https : / / arxiv.org/abs/1906.01820. Hume, David (1739). A Treatise of Human Nature (1739-40) . Ed. by Ernest Campbell Mossner. Mineola, N.Y.: Oxford University Press. ...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.