Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The same assumptions that motivate a US race to build superintelligent AI also imply that racing would be catastrophic, and that verified mutual restraint is the strategically sound choice.

desk verdict A valuable internal critique of the ASI-race logic, but the game-theoretic table is mis-specified and the verification premise is asserted rather than shown. read the letter →

arxiv 2501.14749 v1 pith:VXVYBYYH submitted 2024-12-22 cs.CY

classification cs.CY
keywords artificialsuperintelligenceASIracetrustdilemmainternationalcooperationverificationregimestrategicstabilitylossofcontrolliberaldemocracy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the strategic case for the United States to race China to artificial superintelligence is self-defeating under its own premises. Taking seriously the two assumptions that motivate the race—that ASI confers a decisive military advantage and that states are rational survival-seekers—the paper derives three escalating dangers: adversaries would rationally consider a rival ASI project an existential threat and may strike preemptively; the extremely rapid capability growth needed for a decisive edge maximizes the chance of losing control of the system; and even a controlled ASI would concentrate domestic power in ways incompatible with liberal democracy. Because both states would prefer mutual restraint to racing, the situation is a trust dilemma rather than a prisoner's dilemma, and a verification regime that makes an ASI project detectable can sustain cooperation. The paper concludes that international cooperation to avoid an ASI race is not only preferable but achievable.

What carries the argument

The central device is the trust dilemma (distinct from a prisoner's dilemma): a game in which both players would rather cooperate than defect, but each would defect if the other defects, so there are two Nash equilibria and mutual restraint is self-reinforcing once trust is established. The paper couples this with a verification analysis from the arms-control literature: arms control succeeds when the controlled technology is highly distinguishable and non-integrated, and fails when it is dual-use and embedded. The argument claims ASI development is precisely the distinguishable, non-integrated case, so a far-from-perfect verification regime can support the cooperative equilibrium.

What would settle it

If a plausible ASI could be built through distributed training runs or algorithmic improvements small enough to hide inside ordinary civilian AI compute, then national technical means could not distinguish an ASI project, verification collapses, and the trust-dilemma path to cooperation no longer holds.

Watch

Extended reading notes

Core claim

Under the assumptions that motivate an ASI race—that the first developer gains a decisive military advantage and that states are rational actors prioritizing survival—racing to ASI is an existential threat to the racing states themselves. The paper shows that these assumptions imply three successive barriers: great-power conflict (adversaries rationally preempt a project that would eliminate their deterrent), loss of control (a system capable of overwhelming superpower militaries is by definition catastrophic if it goes wrong, and the rapid takeoff required for a decisive advantage is the scenario where control is hardest), and power concentration (the small group controlling ASI would hold unchecked domestic power, undermining liberal democracy). These dangers mean states would both prefer a world without ASI projects to a race, so the interaction is a trust dilemma with two equilibria, cooperate-cooperate and defect-defect, with the former preferable. Cooperation is achievable because an ASI project would be highly distinguishable from civilian AI and not integrated with the economy, making verification feasible and enabling a cooperative equilibrium.

Load-bearing premise

That an ASI project would be highly distinguishable from civilian AI work and not integrated into a state's economy, so a verification regime could actually detect it.

Editorial extensions

If this is right

  • If the paper is right, a US race to ASI would undermine its own defensive purpose: it would invite preemptive attack, risk uncontrolled superintelligence, and erode the liberal democracy it claims to protect.
  • Cooperation need not require perfect verification or military enforcement; mutual perception of rationality plus a demonstrably detectable ASI project can hold the cooperative equilibrium.
  • General-purpose AI arms control is likely to fail because the technology is indistinguishable and integrated, but an ASI-specific treaty has favorable verification conditions.
  • The analysis suggests the US and China should prefer a conditional-commitment arrangement: each credibly pledges not to pursue an ASI project if the other does the same.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The trust-dilemma logic suggests a small number of states—perhaps just the US and China—could form a self-reinforcing restraint pact without needing a global treaty or airstrike enforcement.
  • If algorithmic progress or decentralized training shrinks the physical footprint of an ASI run, the paper's window for verifiable cooperation may close; timing matters for any treaty effort.
  • The paper's internal-critique structure means the conclusion survives even if one rejects loss-of-control as likely, since conflict and democratic-erosion risks alone would justify mutual restraint.
  • A testable extension: analyze historical dual-use arms-control cases along the distinguishability/integration dimensions to estimate the minimum verification threshold needed for an ASI agreement.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper offers an internal critique of the case for a US-led race to artificial superintelligence (ASI). It argues that the two assumptions underlying the race—that ASI confers a decisive military advantage and that states are rational, survival-seeking actors—imply that racing is catastrophically dangerous via great-power conflict, loss of control, and domestic power concentration. It then models the strategic situation as a 'trust dilemma' rather than a prisoner's dilemma, concluding that mutual restraint is both preferable and achievable through a verification regime aimed at the distinguishability of ASI projects. The paper is a policy/international-relations analysis drawing on existing literature rather than an empirical or technical contribution.

Significance. If the argument were sound, the paper would be a timely and relevant contribution to AI governance debates, offering a clear rejoinder to Manhattan-Project-style calls for an ASI race. Its strengths are its internal-critique structure, its use of established IR concepts (defensive realism, the transparency-security tradeoff, and the dual-use typology), and its explicit acknowledgment of uncertainty about loss-of-control risk. It also makes a falsifiable empirical premise—that ASI training runs are detectable and distinguishable from civilian AI R&D—which is a useful target for future work. However, the formal game-theoretic support is currently mis-specified, and the verification premise is asserted rather than demonstrated, so the paper's central conclusion is not yet established.

major comments (3)
  1. [§4.1, Table 1] The payoff matrix in Table 1 does not implement the 'trust dilemma' described in the text. With payoffs (C,D)=(2,0) and (D,C)=(0,2), Cooperate is a strictly dominant strategy for both players: if the opponent cooperates, C yields 10 versus D's 0, and if the opponent defects, C yields 2 versus D's 1. Consequently (D,D) is not a Nash equilibrium, contradicting the claim that there are two Nash equilibria, and the condition 'if one player defects, it is better for the other to defect as well' fails. The intended stag-hunt payoffs would require (C,D) and (D,C) to be swapped (e.g., (0,2) and (2,0) respectively). This error is load-bearing because the paper's conclusion that cooperation is strategically sound rests on the trust-dilemma structure.
  2. [§5.1, incl. footnote 58] The claim that an ASI project would be 'highly distinguishable from civilian AI applications and not integrated with a state's economy' is the key premise for verification feasibility and hence for the conclusion that cooperation is achievable, but it is asserted rather than argued. Footnote 58 concedes that distributed training or large algorithmic efficiency gains would make ASI development control harder, and both are active research directions; the paper gives no reason to expect the current centralized-training paradigm to persist through the relevant window. Without this premise, the verification regime cannot provide the mutual confidence required to select the cooperate-cooperate equilibrium, so the paper's central conclusion is unsupported as written. The authors should either supply evidence or explicitly conditionalize the conclusion on the persistence of centralized training runs.
  3. [§5.1, paragraph beginning 'In fact, verification might play only a partial role'] The argument that rational states would understand that the other has 'no reason to defect' is circular in a trust dilemma: with two equilibria, rationality alone does not select cooperate-cooperate, and the entire problem is to establish the mutual expectation that prevents defection. The paper's suggestion that demonstrating rationality is sufficient effectively assumes the cooperation it is trying to establish. This does not invalidate the verification argument, but it should be removed or reframed as a comment about equilibrium selection rather than as an independent path to trust.
minor comments (5)
  1. [§6, Conclusion] The sentence 'Second it the heightened risk' should read 'Second is the heightened risk'; likewise, 'the scenario is which loss of control risk is greatest' should read 'the scenario in which loss of control risk is greatest'.
  2. [§4.1] The term 'trust dilemma' is attributed to Jervis 1978, but the cited article is primarily about the 'security dilemma'; the authors should clarify the provenance of the term and give a precise game-theoretic definition.
  3. [Table 1] The caption says 'A Trust Dilemma' but the payoffs are not those of a trust dilemma; even after correcting the entries, the caption should note the equilibrium-selection convention used.
  4. [§4.1] The discussion of commitment devices says that a device 'works' when it makes cooperation a dominant strategy, which describes a harmony game rather than a trust dilemma; this should be reconciled with the preceding definition, since a dominant-strategy cooperation equilibrium would eliminate the multiplicity that motivates verification.
  5. [§2.1 and §3.1] The paper uses 'decisive military advantage' at different levels of specificity (e.g., undermining nuclear deterrence versus broader military superiority); a single formal definition would strengthen the argument.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a conditional internal critique, and its cooperation conclusion depends on an asserted empirical premise rather than on a self-referential or fitted derivation.

full rationale

The derivation chain is self-contained and non-circular. The paper's central claim is a conditional reductio: from the racer's own premises (decisive military advantage plus rational, survival-prioritizing states) it derives three dangers, then argues that the same preference structure makes mutual restraint a preferred equilibrium. No parameter is fitted to data and then relabeled as a prediction; no equation defines its own output; no load-bearing result is imported from the authors' prior work. The two self-citations (footnotes 15 and 40) merely support background uncertainty about timelines and agency risk, and they are not used to justify the core strategic conclusion. The cooperation conclusion rests on external theory (Jervis; Coe and Vaynman; Vaynman and Volpe) plus an asserted empirical premise that an ASI project would be distinguishable from civilian AI and unintegrated with the economy. That premise is contestable: footnote 58 concedes that distributed training or algorithmic improvements would make verification harder, and Table 1's payoff matrix is internally inconsistent (cooperate strictly dominates, so defect-defect is not a Nash equilibrium). These are correctness and robustness weaknesses, not circularity. The paper does not assume cooperation is achievable in order to prove it; it offers conditional reasons why, given the racer's assumptions, cooperation would be preferred and verification could suffice.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities or fitted quantities; the game model is illustrative rather than empirical. The load-bearing content comes from domain assumptions about state rationality, ASI's military impact, and verification observability.

free parameters (2)
  • Payoff matrix entries in Table 1 (10,10; 2,0; 0,2; 1,1) = 10,10; 2,0; 0,2; 1,1
    Arbitrarily chosen to represent a trust dilemma preference ordering, but the chosen values actually produce a harmony game with Cooperate as a dominant strategy; not derived from the risk analysis.
  • Illustrative probabilities in realist model example (90% control success, 20% adversary development chance) = 90% and 20%
    Used in the Section 4.1 example to show that a state may race despite existential risk; these are illustrative hand-set values, not estimated from evidence.
assumptions (5)
  • domain assumption States are rational actors prioritizing survival (defensive realism).
    Invoked in Section 2.2 as an assumption of the racing argument that the paper adopts for its internal critique.
  • domain assumption ASI would provide a decisive military advantage (DMA) to its first developer.
    Taken from Aschenbrenner and other racing advocates; the paper assumes this for the internal critique (Section 2.1).
  • domain assumption An ASI project would be highly distinguishable from civilian AI and not integrated into the economy, making verification feasible.
    This is the load-bearing assumption for the verification regime in Section 5.1; the paper asserts but does not prove it (see footnote 58 for caveats about distributed training).
  • standard math Game-theoretic concepts of Nash equilibrium and dominant strategies.
    Used in Section 4.1 to analyze the trust dilemma and prisoner's dilemma; standard background.
  • domain assumption Loss of control is a real risk for ASI and is maximized by rapid capability gains.
    Section 3.2 argues this using citations to Bostrom and others, but does not establish the empirical likelihood.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating." pith.science (2026). https://pith.science/paper/VXVYBYYH

@misc{pith2026250114749,
  author       = {Pith},
  title        = {Pith review of: The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeating},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VXVYBYYH}},
  note         = {Machine review of arXiv:2501.14749}
}
read the original abstract

This paper examines the strategic dynamics of international competition to develop Artificial Superintelligence (ASI). We argue that the same assumptions that might motivate the US to race to develop ASI also imply that such a race is extremely dangerous. These assumptions--that ASI would provide a decisive military advantage and that states are rational actors prioritizing survival--imply that a race would heighten three critical risks: great power conflict, loss of control of ASI systems, and the undermining of liberal democracy. Our analysis shows that ASI presents a trust dilemma rather than a prisoners dilemma, suggesting that international cooperation to control ASI development is both preferable and strategically sound. We conclude that cooperation is achievable.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flexible Hardware-Enabled Guarantees for AI Compute

    cs.CR 2025-06 conditional novelty 6.0 of 10

    The paper proposes flexHEGs, open-source tamper-proof hardware on AI accelerators that could enable privacy-preserving international verification and enforcement of AI compute usage.

Reference graph

Works this paper leans on

7 extracted references · 4 canonical work pages · cited by 1 Pith paper

  1. [1]

    We argue that the same assumptions that might motivate the US to race to dev elop ASI also imply that such a race is extremely dangerous

    arXiv:2501.14749v1 [cs.CY] 22 Dec 2024 The Manhattan Trap: Why a Race to Artificial Superintelligence is Self-Defeatin g Corin Katzke ∗ Convergence Analysis Gideon Futerman University of Oxford December 2024 Abstract This paper examines the strategic dynamics of internationa l com- petition to develop Artificial Superintelligence (ASI). We argue that the sa...

  2. [3]

    defense in depth,

    , this pref- erence is similarly extreme—although our analysis still holds even if this is only a mild preference. The crucial implication is that if one state believ es that the other will not develop ASI, then that state will not develop ASI itself. The strategic situation can be described in the language of game the ory as a trust dilemma. 54 In a trus...

  3. [1960]

    Strategic Insights from Simulation Gaming of AI Race Dynamics

    25George Bush: Soviet-United States Joint Statement on Future Ne - gotiations on Nuclear and Space Arms and Further Enhancing Strat e- gic Stability, June 1990, url: https://www.presidency.ucsb.edu/documents/ soviet-united-states-joint-statement-future-negotiations -nuclear-and-space-arms-and. 10 Advocates for an ASI project argument give the Chinese and ...

  4. [1979]

    Better to accept the slavery of the Nazis than to run the chance of drawing the final curtain on mankind!

    18 mercy of an adversary is evaluated by such a utility function as a wor st-case scenario. Of course, an outcome in which all states lose—such as human extinc - tion—is also a worst-case scenario according to a realist utility funct ion. For example, Nathan Sears argues that: “Since there can be no states or nations without the continuation of humanity, ...

  5. [1985]

    We Need to Shut It All Down, in: TIME, Mar

    63Eliezer Yudkowsky: Pausing AI Developments Isn’t Enough. We Need to Shut It All Down, in: TIME, Mar. 2023, url: https://time.com/6266923/ ai-eliezer-yudkowsky-open-letter-not-enough/. 25 If the analysis of ASI development as a trust dilemma is correct, and states can be modeled as rational actors, then the need for enfo rcement of an ASI treaty would be...

  6. [2014]

    Multinational AGI Consortium (MAGIC): A Proposal for International Coordination on AI

    2Cullen O’Keefe: Chips for Peace: How the U.S. and Its Allies Can Lead on Safe and Beneficial AI, in: Lawfare, July 2024, url: https://www.lawfaremedia.org/article/ chips-for-peace--how-the-u.s.-and-its-allies-can-lead-on-s afe-and-beneficial-ai. 3Leopold Aschenbrenner: SITUATIONAL A W ARENESS: The Decade A head, en-US, June 2024, url: https://situational-a...

  7. [2024]

    strategic parameter s,

    6 2 Assumptions Motivating an ASI Race The future of AI development is difficult to predict. The academic com munity remains divided on crucial uncertainties, or “strategic parameter s,” such as timelines to ASI and the risk of loss of control of ASI. 15 These significantly impact how states should approach the implications of AI developmen t on national sec...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.