Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

International Security Applications of Flexible Hardware-Enabled Guarantees

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Comprehensive flexHEG agreements could remain stable under reasonable assumptions about state preferences and catastrophic risks.

desk verdict A candid, well-hedged policy analysis; the stability math is conditional and correct, but the load-bearing assumption about detection speed is asserted, not demonstrated. read the letter →

arxiv 2506.15100 v1 pith:3H64HKR2 submitted 2025-06-18 cs.CR

classification cs.CR
keywords flexiblehardware-enabledguaranteesAIgovernanceinternationalsecuritygametheorychipsverification-basedagreementsruleset-basedstrategicstability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Flexible hardware-enabled guarantees (flexHEGs) put a small guarantee processor on AI chips that can make verifiable claims about the chip's past and future behavior. This report argues that such chips could form the technical backbone of international AI governance—ranging from export-control enforcement to safety commitments at training data centers—and, most centrally, that comprehensive flexHEG agreements could be stable rather than collapsing under competitive pressure. The stability argument is a two-player game: if both states cooperate they preserve the balance of power and avoid catastrophic risk, while a first defector's expected gain is $U(W)(1-P(\mathrm{doom}))P(W|D)$. Under the report's baseline parameters ($U(W)\le 1.5$, $P(\mathrm{doom})\ge 0.1$), cooperation is stable whenever a first defector's chance of ultimately winning the AI race is below about 0.74, and the report argues this is plausible if detection is fast. That matters because credible, self-enforcing agreements are the main prerequisite for using hardware-level control of compute to address AI-related security risks.

What carries the argument

The load-bearing object is the flexible hardware-enabled guarantee: an on-device guarantee processor that lets the chip make cryptographically verifiable claims about what it has done and what it will allow, so that a state can verify compliance without exposing private computations. The stability argument is carried by a two-player stag-hunt game whose central inequality is $U(W)(1-P(\mathrm{doom}))P(W|D)<1$, where $U(W)$ is the value of winning the AI race relative to cooperation, $P(\mathrm{doom})$ is the chance an uncoordinated race ends in catastrophe, and $P(W|D)$ is the first defector's probability of ultimately winning. The report combines this with an ecosystem design—trusted standardized designs, production oversight, chip registries, random inspections, and intelligence agencies looking for violations—to argue that the model's parameters are actually attainable.

What would settle it

Find empirically, through a red-team exercise or analysis of existing verification systems, the time from the start of a covert norm-violating training run to reliable detection; if that latency implies a first-defector win probability at or above about 0.74 (with $U(W)=1.5$, $P(\mathrm{doom})=0.1$), the stability result fails. A second concrete falsifier would be demonstrating a firmware or software vulnerability that allows rapid, scalable compromise of many flexHEG devices, since the report's ruleset-stability argument depends on ruling out such 'fast break' defections.

Watch

Extended reading notes

Core claim

The paper's central claim is that comprehensive flexHEG agreements—covering all or most AI-relevant data center chips, either through verification of compliance or through automatically enforced rulesets—could remain stable under reasonable assumptions about state preferences and catastrophic risks. In the report's model, stability is a stag hunt: both states prefer mutual cooperation to defection, and the agreement survives unless the first defector has both a high value of winning ($U(W)$), a low perceived chance of catastrophe ($P(\mathrm{doom})$), and a high conditional probability of winning after defecting ($P(W|D)$). The threshold calculation shows $U(W)(1-P(\mathrm{doom}))P(W|D)<1$ is the stability condition, and with $U(W)\le 1.5$ and $P(\mathrm{doom})\ge 0.1$ the regime holds when $P(W|D)<0.74$. The report further argues that verification-based agreements keep $P(W|D)$ closer to 0.5 because detection triggers a rapid response, whereas ruleset-based agreements bind persistently but carry a tail risk of a scalable circumvention that would make defection catastrophically attractive; entering the treaty can remain incentive-compatible even for a leading state, and recovery from small violations is possible if the defector's lead stays moderate.

Load-bearing premise

The load-bearing premise is that a flexHEG ecosystem can be built and defended so that a first mover's defection is detected before it gains a decisive lead, i.e. that the realized $P(W|D)$ stays below about 0.74; the report asserts this is plausible but provides no model or measurement of detection latency or hardware tamper-resistance failure rates.

Editorial extensions

If this is right

  • Comprehensive flexHEG agreements (especially verification-based ones) could be self-enforcing rather than requiring a supranational enforcer, as long as defection is detected quickly.
  • If states value winning at most 1.5 times the cooperative outcome and believe an uncoordinated race has at least a 10% chance of catastrophe, a first defector would need better than roughly 74% odds of winning to rationally defect, a bar the report argues is hard to meet.
  • A ruleset-based agreement gives states durable guarantees and prevents sudden unravelling, but the report warns it could be less stable if one state finds a scalable way to circumvent enforcement, because the defector's lead could grow large before others can respond.
  • The same hardware can support targeted applications—verifying training runs at frontier data centers, limiting chip proliferation, providing assurances about military AI, and underpinning balance-of-power or sovereignty guarantees—so the stability result is not limited to one governance form.
  • Managing the existing stock of non-flexHEG chips via registration and inspections matters: the stability logic assumes norm violations cannot be run at scale on unregistered compute.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The stability conclusion is conditional on the realized rate of detection; no measurement or model in the report quantifies how quickly a secret training run would be caught, so the practical question is empirical: measuring detection latency under adversarial conditions would directly test the claim.
  • A natural extension of the model is to intermediate defections—partial violations, delayed inspections, phased exits—where the binary cooperate/defect frame may understate the difficulty of sustaining cooperation; the report acknowledges this but does not model it.
  • The threshold logic suggests a design principle beyond flexHEGs: policies that make defection observable early, such as mandatory real-time attestation or continuous monitoring, are more stabilizing than policies that merely raise the cost of defection but leave it undetectable.
  • The same inequality could be applied to other compute-governance proposals (export controls, AI safety standards), since it identifies the three levers—value of winning, perceived catastrophe risk, and detection speed—that determine whether any international AI agreement is stable.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This report, the third part of a series on flexible hardware-enabled guarantees (flexHEGs), examines how hardware-based verification and enforcement mechanisms for AI-relevant chips could support international security governance. It surveys the prerequisites for an internationally trustworthy flexHEG ecosystem, including design standardization, production oversight, supply-chain tracking, and technical thresholds for AI-relevant chips. It then applies these mechanisms to four application areas: limiting proliferation, implementing AI safety norms, managing military AI risks, and supporting strategic stability. The central formal contribution is a two-player game-theoretic model in Appendix B, which derives a stability condition for comprehensive agreements: U(W)(1-P(doom))P(W|D)<1. With baseline assumptions U(W)≤1.5 and P(doom)≥0.1, cooperation is stable when the first defector's probability of winning the subsequent AI race satisfies P(W|D)<0.74. The paper asserts that a flexHEG verification regime could plausibly detect defection quickly enough to meet this bound, while documenting several model limitations, including the conflation of losing the race with catastrophic outcomes and the model's unrealistic behavior at very high P(doom).

Significance. If read as a conditional framework rather than an empirical demonstration, the paper has genuine value. It makes the strategic logic behind AI compute agreements explicit, identifies P(W|D) as the pivotal unknown, and provides a falsifiable threshold that could in principle be tested with better measurements of detection latency. The algebra in Appendix B is correct, the parameter choices and their uncertainties are disclosed, and the paper is notably candid about the limits of its own model. The main weakness is that the bridge from flexHEG detection mechanisms to an actual value or bound on P(W|D) is asserted with a single plausibility statement, so the central claim in the abstract overreaches the evidence provided. The report is best characterized as a transparent, well-contextualized policy analysis with a small formal core.

major comments (3)
  1. [Appendix B, 'Agreements Could Be Stable Given Reasonable Parameters' (p.55)] The central claim of the paper reduces to the inequality U(W)(1-P(doom))P(W|D)<1, and with U(W)≤1.5 and P(doom)≥0.1 the required bound is P(W|D)<0.74. The paper provides no model, measurement, or historical case connecting flexHEG detection mechanisms to P(W|D). The sentence 'it seems plausible that a flexHEG-based verification regime could detect defection quickly enough to make such high odds very unlikely' is an unsupported empirical assertion. Because P(W|D) is a function of detection latency, inspection coverage, and tamper-resistance failure rates, and these are the very properties the flexHEG ecosystem is supposed to provide, the demonstration is conditional on an unquantified engineering and policy claim. If detection is slower than the near-real-time scenario the paper gestures at, P(W|D) could exceed 0.74 and the cooperative equilibrium would not be stable even with reasonable U(W) and P(doom). Please either supply a substantive argument (for example, a timing analysis of the fastest plausible defection path, or a historical analogue from IAEA inspections or export-control verification) that bounds P(W|D), or explicitly reframe the result as 'stability holds provided P(W|D)<0.74' and temper the abstract's claim accordingly.
  2. [Appendix B, 'Payoff Structure' and 'Limitations and Extensions' (pp.54, 58)] The model sets the utility of losing the race to 0, identical to the utility of a catastrophic outcome. This normalization biases the model toward cooperation. The paper acknowledges the limitation but asserts that proper accounting would make defection 'somewhat more tempting' and is 'unlikely to overturn our key findings.' This robustness claim is not quantified. A simple calculation shows the bias is not negligible: if losing yields utility 0.5 instead of 0, the threshold P(W|D) at U(W)=1.5 and P(doom)=0.1 drops from 0.74 to about 0.61, a non-trivial reduction in the stability margin. The paper should either quantify the sensitivity of the threshold to the loser's utility or soften the claim that the qualitative conclusion is robust.
  3. [Appendix B, 'Impossible Conclusions' (p.59)] The appendix correctly notes that the model implies that for sufficiently high P(doom) no value of P(W|D) can make defection attractive, which is unrealistic because a defector could in principle trade off some of a dominant position to reduce the probability of catastrophe. The paper says this limitation is 'not particularly relevant' for the cases of interest, but no argument is given for why the parameter regime relevant to comprehensive flexHEG agreements avoids this artifact. Since the agreements analyzed here are aimed at high-risk scenarios, the boundary of this limitation should be clarified and its direction of bias stated explicitly in the main stability discussion.
minor comments (4)
  1. [Main text, 'How These Agreements Could be Stable' (p.39)] The phrase '(≥1.5 times)' is inconsistent with the analysis in Appendix B, which assumes U(W)≤1.5. Please correct the inequality sign to avoid confusing readers about the direction of the value assumption.
  2. [Appendix B, Figure 3 caption (p.56)] The caption states 'the agreement is stable for any values of U(W) and P(doom) above the blue line,' but the stability region should be described precisely as the set of (U(W),P(doom)) satisfying U(W)(1-P(doom)) < 1/P(W|D). Please clarify the axes and the region labeling.
  3. [Appendix B, 'Key Parameters' (p.53)] The definition of P(W|D) says only 'probability that a state would ultimately "win" the resulting AI race, conditional on being the first to defect'; the payoff formula multiplies by (1-P(doom)), indicating that the win probability is conditional on no catastrophic outcome. The definition should state this explicitly to avoid ambiguity.
  4. [Appendix B, 'Limitations and Extensions' (p.58)] Under 'Simplistic Utilities,' the paper says states would likely prefer being conquered by humans over destruction by AIs, but the follow-up sentence frames this only as making defection 'somewhat more tempting.' Given the concrete threshold sensitivity shown in the major comment above, the framing should acknowledge the quantitative impact on the P(W|D)| threshold.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the stability result in Appendix B is a conditional game-theoretic derivation, not a renamed fit or a self-citation chain.

full rationale

The paper's derivation chain is confined to Appendix B, where it defines a two-player cooperate/defect game with explicit parameters: U(C)=1, U(W), P(doom), and P(W|D). The expected utility of defecting is written as U(W)*(1-P(doom))*P(W|D), and stability is defined as this being less than 1. This is a standard expected-utility condition, not a hidden identity: the conclusion is the inequality, while P(W|D) is an exogenous parameter whose threshold is derived from the assumed U(W) and P(doom). The parameter values U(W) <= 1.5 and P(doom) >= 0.1 are justified by external arguments (linear resource valuation vs. status-quo preference, and the survey of ML researchers cited as [50]), not fitted to make the conclusion true. The statement that 'it seems plausible that a flexHEG-based verification regime could detect defection quickly enough to make such high odds very unlikely' is a qualitative plausibility judgment, not a fitted input renamed as a prediction. Citations to Parts I and II, and to the authors' earlier reports, provide background definitions and technical context, but the load-bearing stability inequality is derived within this paper from stated assumptions. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via citation, and no known result is merely renamed. The main substantive concerns are the unquantified detection-latency assumption and a small numerical slip in the main text (75% vs. the derived <0.74 threshold), but these are correctness and soundness risks, not circularity.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on four structural assumptions: the technical feasibility of flexHEG hardware, a high enough P(doom), a modest U(W), and the sufficiency of a simplified two-player game. The first is a domain assumption about an unbuilt technology; the second and third are parameter choices; the fourth is an acknowledged modeling simplification. No new physical entities are introduced in this part of the series.

free parameters (2)
  • U(W), utility of winning the AI race = assumed <= 1.5, with U(C) normalized to 1
    Chosen by hand in Appendix B as 'a reasonable balance' between linear resource valuation and indifference to additional gains; the stability condition is sensitive to this value.
  • P(doom), probability of catastrophic outcome in an uncoordinated race = assumed >= 0.1
    Justified via a 5% median extinction-risk survey plus sub-extinction catastrophes; selected to make the threshold analysis illustrative rather than measured.
assumptions (4)
  • domain assumption A trustworthy and tamper-resistant flexHEG ecosystem can be designed, produced, and defended against state-level adversaries.
    Assumed throughout 'Creating an Internationally Trustworthy FlexHEG Ecosystem'; the report proposes oversight measures but provides no implementation or security proof.
  • domain assumption Uncoordinated frontier AI development carries at least a 10% probability of catastrophic outcomes.
    Justified in Appendix B using a researcher survey and qualitative reasoning; contested and not independently measured.
  • ad hoc to paper States are expected-utility maximizers and value winning an AI race at no more than 1.5 times the value of preserving the current balance of power.
    Imposed in Appendix B to make cooperation stable; the authors acknowledge the results are sensitive to U(W).
  • ad hoc to paper A two-player, binary cooperate-or-defect abstraction captures the strategic essence of international AI agreements.
    Two-player stag-hunt simplification in Appendix B; binary utilities, outcomes, and actions are listed by the authors as limitations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of International Security Applications of Flexible Hardware-Enabled Guarantees." pith.science (2026). https://pith.science/paper/3H64HKR2

@misc{pith2026250615100,
  author       = {Pith},
  title        = {Pith review of: International Security Applications of Flexible Hardware-Enabled Guarantees},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3H64HKR2}},
  note         = {Machine review of arXiv:2506.15100}
}
read the original abstract

As AI capabilities advance rapidly, flexible hardware-enabled guarantees (flexHEGs) offer opportunities to address international security challenges through comprehensive governance frameworks. This report examines how flexHEGs could enable internationally trustworthy AI governance by establishing standardized designs, robust ecosystem defenses, and clear operational parameters for AI-relevant chips. We analyze four critical international security applications: limiting proliferation to address malicious use, implementing safety norms to prevent loss of control, managing risks from military AI systems, and supporting strategic stability through balance-of-power mechanisms while respecting national sovereignty. The report explores both targeted deployments for specific high-risk facilities and comprehensive deployments covering all AI-relevant compute. We examine two primary governance models: verification-based agreements that enable transparent compliance monitoring, and ruleset-based agreements that automatically enforce international rules through cryptographically-signed updates. Through game-theoretic analysis, we demonstrate that comprehensive flexHEG agreements could remain stable under reasonable assumptions about state preferences and catastrophic risks. The report addresses critical implementation challenges including technical thresholds for AI-relevant chips, management of existing non-flexHEG hardware, and safeguards against abuse of governance power. While requiring significant international coordination, flexHEGs could provide a technical foundation for managing AI risks at the scale and speed necessary to address emerging threats to international security and stability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements

    cs.CY 2026-06 conditional novelty 6.0 of 10

    Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.

Reference graph

Works this paper leans on

50 extracted references · 31 canonical work pages · cited by 1 Pith paper

  1. [1]

    Verifiable Compute White Paper,

    “Verifiable Compute White Paper,” EQTY Lab, 2024. [Online]. Available: https://www.eqtylab.io/verifiable-compute-white-paper

  2. [2]

    Secure Enclaves for AI Evaluation,

    A. Trask et al. , “Secure Enclaves for AI Evaluation,” OpenMined Blog. [Online]. Available: https://blog.openmined.org/secure-enclaves-for-ai-evaluation/

  3. [3]

    Anderson, Security engineering: a guide to building dependable distributed systems , 3rd edition

    R. Anderson, Security engineering: a guide to building dependable distributed systems , 3rd edition. John Wiley & Sons, 2020. [Online]. Available: https://www.cl.cam.ac.uk/~rja14/book.html

  4. [4]

    Security Verification of the OpenTitan Hardware Root of Trust,

    A. Meza, F. Restuccia, J. Oberg, D. Rizzo, and R. Kastner, “Security Verification of the OpenTitan Hardware Root of Trust,” IEEE Secur. Priv. , vol. 21, no. 3, pp. 27–36, May 2023, doi: 10.1109/MSEC.2023.3251954

  5. [5]

    The seL4 Microkernel – An Introduction v1.4,

    G. Heiser, “The seL4 Microkernel – An Introduction v1.4,” The seL4 Foundation, Jan. 2025. [6] CHIPS Alliance, “Caliptra: A Datacenter System on a Chip (SoC) Root of Trust (RoT),” GitHub. [Online]. Available: spec.caliptra.io

  6. [7]

    The Legacy of DES,

    B. Schneier, “The Legacy of DES,” Schneier on Security , Oct. 06, 2004. [Online]. Available: https://www.schneier.com/blog/archives/2004/10/the_legacy_of_d.html

  7. [8]

    How the NSA (may have) put a backdoor in RSA’s cryptography: A technical primer,

    “How the NSA (may have) put a backdoor in RSA’s cryptography: A technical primer,” The Cloudflare Blog , Jan. 06, 2014. [Online]. Available: https://blog.cloudflare.com/how-the-nsa-may-have-put-a-backdoor-in-rsas-cryptography- a-technical-primer/

  8. [9]

    Location Verification for AI Chips,

    A. Brass and O. Aarne, “Location Verification for AI Chips,” Institute for AI Policy and Strategy, Apr. 2024. [Online]. Available: https://www.iaps.ai/research/location-verification-for-ai-chips

Show all 50 references
  1. [10]

    Nuclear Arms Control Verification and Lessons for AI Treaties,

    M. Baker, “Nuclear Arms Control Verification and Lessons for AI Treaties,” Apr. 08, 2023, arXiv : arXiv:2304.04123. doi: 10.48550/arXiv.2304.04123

  2. [11]

    Statement by the Press Secretary

    “Statement by the Press Secretary.” Office of the Press Secretary, Apr. 24, 2008. [Online]. Available: https://georgewbush-whitehouse.archives.gov/news/releases/2008/04/20080424-14.html

  3. [12]

    International Verification and Intelligence,

    J. M. Acton, “International Verification and Intelligence,” Intell. Natl. Secur. , vol. 29, no. 3, pp. 341–356, May 2014, doi: 10.1080/02684527.2014.895592

  4. [13]

    Verification methods for international AI agreements,

    A. R. Wasil, T. Reed, J. W. Miller, and P. Barnett, “Verification methods for international AI agreements,” Aug. 28, 2024, arXiv : arXiv:2408.16074. [Online]. Available: http://arxiv.org/abs/2408.16074

  5. [14]

    Mechanisms to Verify International Agreements About AI Development,

    A. Scher and L. Thiergart, “Mechanisms to Verify International Agreements About AI Development,” Machine Intelligence Research Institute, Nov. 2024

  6. [15]

    Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use; Updates and Corrections,

    Bureau of Industry and Security, “Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use; Updates and Corrections,” Federal Register, 2023-23055 (88 FR 73458), Oct. 2023. [Online]. Available: https://www.federalr...

  7. [16]

    With Smugglers and Front Companies, China Is Skirting American A.I. Bans,

    A. Swanson and C. Fu, “With Smugglers and Front Companies, China Is Skirting American A.I. Bans,” The New York Times , Aug. 04, 2024. [Online]. Available: https://www.nytimes.com/2024/08/04/technology/china-ai-microchips.html

  8. [17]

    Nvidia AI Chip Smuggling to China Becomes an Industry,

    Q. Liu, “Nvidia AI Chip Smuggling to China Becomes an Industry,” The Information , Aug. 12, 2024. [Online]. Available: https://www.theinformation.com/articles/nvidia-ai-chip-smuggling-to-china-becomes-an- 43 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III in...

  9. [19]

    Training Compute Thresholds: Features and Functions in AI Regulation,

    L. Heim and L. Koessler, “Training Compute Thresholds: Features and Functions in AI Regulation,” Aug. 06, 2024, arXiv : arXiv:2405.10799. doi: 10.48550/arXiv.2405.10799

  10. [20]

    Compute Thresholds are Ineffective,

    D. W. Ball, “Compute Thresholds are Ineffective,” Hyperdimensional , Feb. 26, 2025. [Online]. Available: https://www.hyperdimensional.co/p/compute-thresholds-are-ineffective

  11. [21]

    Increased Compute Efficiency and the Diffusion of AI Capabilities,

    K. Pilz, L. Heim, and N. Brown, “Increased Compute Efficiency and the Diffusion of AI Capabilities,” Nov. 26, 2023, arXiv : arXiv:2311.15377. [Online]. Available: http://arxiv.org/abs/2311.15377

  12. [22]

    Societal Adaptation to Advanced AI,

    J. Bernardi, G. Mukobi, H. Greaves, L. Heim, and M. Anderljung, “Societal Adaptation to Advanced AI,” Jan. 23, 2025, arXiv : arXiv:2405.10295. doi: 10.48550/arXiv.2405.10295

  13. [23]

    Are Consumer GPUs a Problem for US Export Controls?,

    E. Grunewald, “Are Consumer GPUs a Problem for US Export Controls?,” Institute for AI Policy and Strategy, May 2024. [Online]. Available: https://www.iaps.ai/research/are-consumer-gpus-a-problem-for-us-export-controls

  14. [24]

    Framework for Artificial Intelligence Diffusion,

    Bureau of Industry and Security, “Framework for Artificial Intelligence Diffusion,” Federal Register, 2025–00636 (90 FR 4544), Jan. 2025. [Online]. Available: https://www.federalregister.gov/documents/2025/01/15/2025-00636/framework-for-artif icial-intelligence-diffusion

  15. [25]

    Secure, Governable Chips,

    O. Aarne, T. Fist, and C. Withers, “Secure, Governable Chips,” Center for a New American Security, Jan. 2024. [Online]. Available: https://www.cnas.org/publications/reports/secure-governable-chips

  16. [26]

    Chips for Peace: How the U.S. and Its Allies Can Lead on Safe and Beneficial AI,

    C. O’Keefe, “Chips for Peace: How the U.S. and Its Allies Can Lead on Safe and Beneficial AI,” Lawfare , Jul. 10, 2024. [Online]. Available: https://www.lawfaremedia.org/article/chips-for-peace--how-the-u.s.-and-its-allies-can-lea d-on-safe-and-beneficial-ai

  17. [27]

    International governance of advancing artificial intelligence,

    N. Emery-Xu, R. Jordan, and R. Trager, “International governance of advancing artificial intelligence,” AI Soc. , Sep. 2024, doi: 10.1007/s00146-024-02050-7

  18. [28]

    Statement on AI Risk

    “Statement on AI Risk.” Center for AI Safety, 2023. [Online]. Available: https://www.safe.ai/work/statement-on-ai-risk

  19. [29]

    Global AI Governance Initiative,

    “Global AI Governance Initiative,” Ministry of Foreign Affairs, the People’s Republic of China, Oct. 2023. [Online]. Available: https://www.mfa.gov.cn/eng/zy/gb/202405/t20240531_11367503.html

  20. [30]

    International AI Safety Report,

    Y. Bengio et al. , “International AI Safety Report,” DSIT 2025/001, 2025. [Online]. Available: https://www.gov.uk/government/publications/international-ai-safety-report-2025

  21. [31]

    IDAIS-Venice Statement,

    “IDAIS-Venice Statement,” Sep. 2024. [Online]. Available: https://idais.ai/dialogue/idais-venice/

  22. [32]

    AI Control: Improving Safety Despite Intentional Subversion,

    R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger, “AI Control: Improving Safety Despite Intentional Subversion,” Jul. 23, 2024, arXiv : arXiv:2312.06942. doi: 10.48550/arXiv.2312.06942

  23. [33]

    AI safety via debate,

    G. Irving, P. Christiano, and D. Amodei, “AI safety via debate,” Oct. 22, 2018, arXiv : arXiv:1805.00899. doi: 10.48550/arXiv.1805.00899

  24. [34]

    Factuality of Large Language Models: A Survey,

    Y. Wang et al. , “Factuality of Large Language Models: A Survey,” Oct. 31, 2024, arXiv : arXiv:2402.02420. doi: 10.48550/arXiv.2402.02420

  25. [35]

    Characterizing Manipulation from AI Systems,

    M. Carroll, A. Chan, H. Ashton, and D. Krueger, “Characterizing Manipulation from AI Systems,” Oct. 30, 2023, arXiv : arXiv:2303.09387. doi: 10.48550/arXiv.2303.09387

  26. [36]

    Discovering Latent Knowledge in Language Models Without Supervision,

    C. Burns, H. Ye, D. Klein, and J. Steinhardt, “Discovering Latent Knowledge in Language Models Without Supervision,” Mar. 02, 2024, arXiv : arXiv:2212.03827. doi: 10.48550/arXiv.2212.03827. 44 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III

  27. [37]

    Evaluating Frontier Models for Dangerous Capabilities,

    M. Phuong et al. , “Evaluating Frontier Models for Dangerous Capabilities,” Apr. 05, 2024, arXiv : arXiv:2403.13793. doi: 10.48550/arXiv.2403.13793

  28. [38]

    Evaluating Language-Model Agents on Realistic Autonomous Tasks,

    M. Kinniment et al. , “Evaluating Language-Model Agents on Realistic Autonomous Tasks,” Jan. 04, 2024, arXiv : arXiv:2312.11671. doi: 10.48550/arXiv.2312.11671

  29. [39]

    Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?,

    Y. Bengio et al. , “Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?,” Feb. 24, 2025, arXiv : arXiv:2502.15657. doi: 10.48550/arXiv.2502.15657

  30. [40]

    Tools for Verifying Neural Models’ Training Data,

    D. Choi, Y. Shavit, and D. Duvenaud, “Tools for Verifying Neural Models’ Training Data,” Jul. 02, 2023, arXiv : arXiv:2307.00682. [Online]. Available: http://arxiv.org/abs/2307.00682

  31. [41]

    Confidence-Building Measures for Artificial Intelligence: Workshop Proceedings,

    S. Shoker et al. , “Confidence-Building Measures for Artificial Intelligence: Workshop Proceedings,” Aug. 03, 2023, arXiv : arXiv:2308.00862. doi: 10.48550/arXiv.2308.00862

  32. [42]

    AI and International Stability: Risks and Confidence-Building Measures,

    M. Horowitz and P. Scharre, “AI and International Stability: Risks and Confidence-Building Measures,” Center for a New American Security, Jan. 2021. [Online]. Available: https://www.cnas.org/publications/reports/ai-and-international-stability-risks-and-confid ence-building-measures

  33. [43]

    Artificial General Intelligence’s Five Hard National Security Problems,

    J. Mitre and J. B. Predd, “Artificial General Intelligence’s Five Hard National Security Problems,” RAND Corporation, Feb. 2025. [Online]. Available: https://www.rand.org/pubs/perspectives/PEA3691-4.html

  34. [44]

    Situational Awareness: The Decade Ahead,

    L. Aschenbrenner, “Situational Awareness: The Decade Ahead,” Jun. 2024. [Online]. Available: https://situational-awareness.ai/

  35. [45]

    H. A. Kissinger, E. Schmidt, and C. Mundie, Genesis: Artificial Intelligence, Hope, and the Human Spirit . Little, Brown and Company, 2024

  36. [46]

    Superintelligence Strategy: Expert Version,

    D. Hendrycks, E. Schmidt, and A. Wang, “Superintelligence Strategy: Expert Version,” Mar. 07, 2025, arXiv : arXiv:2503.05628. doi: 10.48550/arXiv.2503.05628

  37. [47]

    Rationalist Explanations for War,

    J. D. Fearon, “Rationalist Explanations for War,” Int. Organ. , vol. 49, no. 3, pp. 379–414, 1995. [48] A. J. Coe and J. Vaynman, “Why Arms Control Is So Rare,” Am. Polit. Sci. Rev. , vol. 114, no. 2, pp. 342–355, May 2020, doi: 10.1017/S000305541900073X

  38. [48]

    Continuing to cooperate by adhering to the agreement's rules and verification requirements regarding AI development and deployment

  39. [49]

    The Anti-Ballistic Missile (ABM) Treaty at a Glance,

    D. Kimball and K. Reif, “The Anti-Ballistic Missile (ABM) Treaty at a Glance,” Arms Control Association, Dec. 2020. [Online]. Available: https://www.armscontrol.org/factsheets/anti-ballistic-missile-abm-treaty-glance

  40. [50]

    Thousands of AI Authors on the Future of AI,

    K. Grace, H. Stewart, J. F. Sandkühler, S. Thomas, B. Weinstein-Raun, and J. Brauner, “Thousands of AI Authors on the Future of AI,” Apr. 30, 2024, arXiv : arXiv:2401.02843. doi: 10.48550/arXiv.2401.02843. 45 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III A...

  41. [51]

    Defecting by covertly violating the agreement in an attempt to gain an advantage in AI capabilities over the other player, for example by secretly conducting norm-violating training runs or deploying powerful models without required safety measures We assume that once a defect...

  42. [52]

    compute overhang

    If P(doom) < 1 / U(W), stability requires: 54 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III Agreements Could Be Stable Given Reasonable Parameters As an example, we might assume that, for both players, U(W) ≤ 1.5, and P(doom) ≥ 0.1. U(W) ≤ 1.5 appears to b...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.