REVIEW 3 major objections 5 minor 1 cited by
The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The same assumptions that motivate a US race to build superintelligent AI also imply that racing would be catastrophic, and that verified mutual restraint is the strategically sound choice.
desk verdict A valuable internal critique of the ASI-race logic, but the game-theoretic table is mis-specified and the verification premise is asserted rather than shown. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central device is the trust dilemma (distinct from a prisoner's dilemma): a game in which both players would rather cooperate than defect, but each would defect if the other defects, so there are two Nash equilibria and mutual restraint is self-reinforcing once trust is established. The paper couples this with a verification analysis from the arms-control literature: arms control succeeds when the controlled technology is highly distinguishable and non-integrated, and fails when it is dual-use and embedded. The argument claims ASI development is precisely the distinguishable, non-integrated case, so a far-from-perfect verification regime can support the cooperative equilibrium.
What would settle it
If a plausible ASI could be built through distributed training runs or algorithmic improvements small enough to hide inside ordinary civilian AI compute, then national technical means could not distinguish an ASI project, verification collapses, and the trust-dilemma path to cooperation no longer holds.
Extended reading notes
Core claim
Under the assumptions that motivate an ASI race—that the first developer gains a decisive military advantage and that states are rational actors prioritizing survival—racing to ASI is an existential threat to the racing states themselves. The paper shows that these assumptions imply three successive barriers: great-power conflict (adversaries rationally preempt a project that would eliminate their deterrent), loss of control (a system capable of overwhelming superpower militaries is by definition catastrophic if it goes wrong, and the rapid takeoff required for a decisive advantage is the scenario where control is hardest), and power concentration (the small group controlling ASI would hold unchecked domestic power, undermining liberal democracy). These dangers mean states would both prefer a world without ASI projects to a race, so the interaction is a trust dilemma with two equilibria, cooperate-cooperate and defect-defect, with the former preferable. Cooperation is achievable because an ASI project would be highly distinguishable from civilian AI and not integrated with the economy, making verification feasible and enabling a cooperative equilibrium.
Load-bearing premise
That an ASI project would be highly distinguishable from civilian AI work and not integrated into a state's economy, so a verification regime could actually detect it.
Editorial extensions
If this is right
- If the paper is right, a US race to ASI would undermine its own defensive purpose: it would invite preemptive attack, risk uncontrolled superintelligence, and erode the liberal democracy it claims to protect.
- Cooperation need not require perfect verification or military enforcement; mutual perception of rationality plus a demonstrably detectable ASI project can hold the cooperative equilibrium.
- General-purpose AI arms control is likely to fail because the technology is indistinguishable and integrated, but an ASI-specific treaty has favorable verification conditions.
- The analysis suggests the US and China should prefer a conditional-commitment arrangement: each credibly pledges not to pursue an ASI project if the other does the same.
Reading between the lines
- The trust-dilemma logic suggests a small number of states—perhaps just the US and China—could form a self-reinforcing restraint pact without needing a global treaty or airstrike enforcement.
- If algorithmic progress or decentralized training shrinks the physical footprint of an ASI run, the paper's window for verifiable cooperation may close; timing matters for any treaty effort.
- The paper's internal-critique structure means the conclusion survives even if one rejects loss-of-control as likely, since conflict and democratic-erosion risks alone would justify mutual restraint.
- A testable extension: analyze historical dual-use arms-control cases along the distinguishability/integration dimensions to estimate the minimum verification threshold needed for an ASI agreement.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper offers an internal critique of the case for a US-led race to artificial superintelligence (ASI). It argues that the two assumptions underlying the race—that ASI confers a decisive military advantage and that states are rational, survival-seeking actors—imply that racing is catastrophically dangerous via great-power conflict, loss of control, and domestic power concentration. It then models the strategic situation as a 'trust dilemma' rather than a prisoner's dilemma, concluding that mutual restraint is both preferable and achievable through a verification regime aimed at the distinguishability of ASI projects. The paper is a policy/international-relations analysis drawing on existing literature rather than an empirical or technical contribution.
Significance. If the argument were sound, the paper would be a timely and relevant contribution to AI governance debates, offering a clear rejoinder to Manhattan-Project-style calls for an ASI race. Its strengths are its internal-critique structure, its use of established IR concepts (defensive realism, the transparency-security tradeoff, and the dual-use typology), and its explicit acknowledgment of uncertainty about loss-of-control risk. It also makes a falsifiable empirical premise—that ASI training runs are detectable and distinguishable from civilian AI R&D—which is a useful target for future work. However, the formal game-theoretic support is currently mis-specified, and the verification premise is asserted rather than demonstrated, so the paper's central conclusion is not yet established.
major comments (3)
- [§4.1, Table 1] The payoff matrix in Table 1 does not implement the 'trust dilemma' described in the text. With payoffs (C,D)=(2,0) and (D,C)=(0,2), Cooperate is a strictly dominant strategy for both players: if the opponent cooperates, C yields 10 versus D's 0, and if the opponent defects, C yields 2 versus D's 1. Consequently (D,D) is not a Nash equilibrium, contradicting the claim that there are two Nash equilibria, and the condition 'if one player defects, it is better for the other to defect as well' fails. The intended stag-hunt payoffs would require (C,D) and (D,C) to be swapped (e.g., (0,2) and (2,0) respectively). This error is load-bearing because the paper's conclusion that cooperation is strategically sound rests on the trust-dilemma structure.
- [§5.1, incl. footnote 58] The claim that an ASI project would be 'highly distinguishable from civilian AI applications and not integrated with a state's economy' is the key premise for verification feasibility and hence for the conclusion that cooperation is achievable, but it is asserted rather than argued. Footnote 58 concedes that distributed training or large algorithmic efficiency gains would make ASI development control harder, and both are active research directions; the paper gives no reason to expect the current centralized-training paradigm to persist through the relevant window. Without this premise, the verification regime cannot provide the mutual confidence required to select the cooperate-cooperate equilibrium, so the paper's central conclusion is unsupported as written. The authors should either supply evidence or explicitly conditionalize the conclusion on the persistence of centralized training runs.
- [§5.1, paragraph beginning 'In fact, verification might play only a partial role'] The argument that rational states would understand that the other has 'no reason to defect' is circular in a trust dilemma: with two equilibria, rationality alone does not select cooperate-cooperate, and the entire problem is to establish the mutual expectation that prevents defection. The paper's suggestion that demonstrating rationality is sufficient effectively assumes the cooperation it is trying to establish. This does not invalidate the verification argument, but it should be removed or reframed as a comment about equilibrium selection rather than as an independent path to trust.
minor comments (5)
- [§6, Conclusion] The sentence 'Second it the heightened risk' should read 'Second is the heightened risk'; likewise, 'the scenario is which loss of control risk is greatest' should read 'the scenario in which loss of control risk is greatest'.
- [§4.1] The term 'trust dilemma' is attributed to Jervis 1978, but the cited article is primarily about the 'security dilemma'; the authors should clarify the provenance of the term and give a precise game-theoretic definition.
- [Table 1] The caption says 'A Trust Dilemma' but the payoffs are not those of a trust dilemma; even after correcting the entries, the caption should note the equilibrium-selection convention used.
- [§4.1] The discussion of commitment devices says that a device 'works' when it makes cooperation a dominant strategy, which describes a harmony game rather than a trust dilemma; this should be reconciled with the preceding definition, since a dominant-strategy cooperation equilibrium would eliminate the multiplicity that motivates verification.
- [§2.1 and §3.1] The paper uses 'decisive military advantage' at different levels of specificity (e.g., undermining nuclear deterrence versus broader military superiority); a single formal definition would strengthen the argument.
Circularity Check
No significant circularity: the paper is a conditional internal critique, and its cooperation conclusion depends on an asserted empirical premise rather than on a self-referential or fitted derivation.
full rationale
The derivation chain is self-contained and non-circular. The paper's central claim is a conditional reductio: from the racer's own premises (decisive military advantage plus rational, survival-prioritizing states) it derives three dangers, then argues that the same preference structure makes mutual restraint a preferred equilibrium. No parameter is fitted to data and then relabeled as a prediction; no equation defines its own output; no load-bearing result is imported from the authors' prior work. The two self-citations (footnotes 15 and 40) merely support background uncertainty about timelines and agency risk, and they are not used to justify the core strategic conclusion. The cooperation conclusion rests on external theory (Jervis; Coe and Vaynman; Vaynman and Volpe) plus an asserted empirical premise that an ASI project would be distinguishable from civilian AI and unintegrated with the economy. That premise is contestable: footnote 58 concedes that distributed training or algorithmic improvements would make verification harder, and Table 1's payoff matrix is internally inconsistent (cooperate strictly dominates, so defect-defect is not a Nash equilibrium). These are correctness and robustness weaknesses, not circularity. The paper does not assume cooperation is achievable in order to prove it; it offers conditional reasons why, given the racer's assumptions, cooperation would be preferred and verification could suffice.
Assumptions & free parameters
free parameters (2)
- Payoff matrix entries in Table 1 (10,10; 2,0; 0,2; 1,1) =
10,10; 2,0; 0,2; 1,1
- Illustrative probabilities in realist model example (90% control success, 20% adversary development chance) =
90% and 20%
assumptions (5)
- domain assumption States are rational actors prioritizing survival (defensive realism).
- domain assumption ASI would provide a decisive military advantage (DMA) to its first developer.
- domain assumption An ASI project would be highly distinguishable from civilian AI and not integrated into the economy, making verification feasible.
- standard math Game-theoretic concepts of Nash equilibrium and dominant strategies.
- domain assumption Loss of control is a real risk for ASI and is maximized by rapid capability gains.
Cite this review
Pith. "Pith review of The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating." pith.science (2026). https://pith.science/paper/VXVYBYYH
@misc{pith2026250114749,
author = {Pith},
title = {Pith review of: The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating},
year = {2026},
howpublished = {\url{https://pith.science/paper/VXVYBYYH}},
note = {Machine review of arXiv:2501.14749}
}
read the original abstract
This paper examines the strategic dynamics of international competition to develop Artificial Superintelligence (ASI). We argue that the same assumptions that might motivate the US to race to develop ASI also imply that such a race is extremely dangerous. These assumptions--that ASI would provide a decisive military advantage and that states are rational actors prioritizing survival--imply that a race would heighten three critical risks: great power conflict, loss of control of ASI systems, and the undermining of liberal democracy. Our analysis shows that ASI presents a trust dilemma rather than a prisoners dilemma, suggesting that international cooperation to control ASI development is both preferable and strategically sound. We conclude that cooperation is achievable.
Forward citations
Cited by 1 Pith paper
-
Flexible Hardware-Enabled Guarantees for AI Compute
The paper proposes flexHEGs, open-source tamper-proof hardware on AI accelerators that could enable privacy-preserving international verification and enforcement of AI compute usage.
Reference graph
Works this paper leans on
-
[1]
arXiv:2501.14749v1 [cs.CY] 22 Dec 2024 The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeatin g Corin Katzke ∗ Convergence Analysis Gideon Futerman University of Oxford December 2024 Abstract This paper examines the strategic dynamics of internationa l com- petition to develop Artificial Superintelligence (ASI). We argue that the sa...
arXiv 2024
-
[3]
, this pref- erence is similarly extreme—although our analysis still holds even if this is only a mild preference. The crucial implication is that if one state believ es that the other will not develop ASI, then that state will not develop ASI itself. The strategic situation can be described in the language of game the ory as a trust dilemma. 54 In a trus...
arXiv 2023
-
[1960]
Strategic Insights from Simulation Gaming of AI Race Dynamics
25George Bush: Soviet-United States Joint Statement on Future Ne - gotiations on Nuclear and Space Arms and Further Enhancing Strat e- gic Stability, June 1990, url: https://www.presidency.ucsb.edu/documents/ soviet-united-states-joint-statement-future-negotiations -nuclear-and-space-arms-and. 10 Advocates for an ASI project argument give the Chinese and ...
work page Pith review arXiv 1990
-
[1979]
18 mercy of an adversary is evaluated by such a utility function as a wor st-case scenario. Of course, an outcome in which all states lose—such as human extinc - tion—is also a worst-case scenario according to a realist utility funct ion. For example, Nathan Sears argues that: “Since there can be no states or nations without the continuation of humanity, ...
work page 2023
-
[1985]
We Need to Shut It All Down, in: TIME, Mar
63Eliezer Yudkowsky: Pausing AI Developments Isn’t Enough. We Need to Shut It All Down, in: TIME, Mar. 2023, url: https://time.com/6266923/ ai-eliezer-yudkowsky-open-letter-not-enough/. 25 If the analysis of ASI development as a trust dilemma is correct, and states can be modeled as rational actors, then the need for enfo rcement of an ASI treaty would be...
arXiv 2023
-
[2014]
Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI
2Cullen O’Keefe: Chips for Peace: How the U.S. and Its Allies Can Lead on Safe and Beneficial AI, in: Lawfare, July 2024, url: https://www.lawfaremedia.org/article/ chips-for-peace--how-the-u.s.-and-its-allies-can-lead-on-s afe-and-beneficial-ai. 3Leopold Aschenbrenner: SITUATIONAL A W ARENESS: The Decade A head, en-US, June 2024, url: https://situational-a...
work page Pith review arXiv doi:10.48550/arxiv.2310.09217 2024
-
[2024]
6 2 Assumptions Motivating an ASI Race The future of AI development is difficult to predict. The academic com munity remains divided on crucial uncertainties, or “strategic parameter s,” such as timelines to ASI and the risk of loss of control of ASI. 15 These significantly impact how states should approach the implications of AI developmen t on national sec...
work page 2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.