REVIEW 3 major objections 4 minor 1 cited by
International Security Applications of Flexible Hardware-Enabled Guarantees
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Comprehensive flexHEG agreements could remain stable under reasonable assumptions about state preferences and catastrophic risks.
desk verdict A candid, well-hedged policy analysis; the stability math is conditional and correct, but the load-bearing assumption about detection speed is asserted, not demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the flexible hardware-enabled guarantee: an on-device guarantee processor that lets the chip make cryptographically verifiable claims about what it has done and what it will allow, so that a state can verify compliance without exposing private computations. The stability argument is carried by a two-player stag-hunt game whose central inequality is $U(W)(1-P(\mathrm{doom}))P(W|D)<1$, where $U(W)$ is the value of winning the AI race relative to cooperation, $P(\mathrm{doom})$ is the chance an uncoordinated race ends in catastrophe, and $P(W|D)$ is the first defector's probability of ultimately winning. The report combines this with an ecosystem design—trusted standardized designs, production oversight, chip registries, random inspections, and intelligence agencies looking for violations—to argue that the model's parameters are actually attainable.
What would settle it
Find empirically, through a red-team exercise or analysis of existing verification systems, the time from the start of a covert norm-violating training run to reliable detection; if that latency implies a first-defector win probability at or above about 0.74 (with $U(W)=1.5$, $P(\mathrm{doom})=0.1$), the stability result fails. A second concrete falsifier would be demonstrating a firmware or software vulnerability that allows rapid, scalable compromise of many flexHEG devices, since the report's ruleset-stability argument depends on ruling out such 'fast break' defections.
Extended reading notes
Core claim
The paper's central claim is that comprehensive flexHEG agreements—covering all or most AI-relevant data center chips, either through verification of compliance or through automatically enforced rulesets—could remain stable under reasonable assumptions about state preferences and catastrophic risks. In the report's model, stability is a stag hunt: both states prefer mutual cooperation to defection, and the agreement survives unless the first defector has both a high value of winning ($U(W)$), a low perceived chance of catastrophe ($P(\mathrm{doom})$), and a high conditional probability of winning after defecting ($P(W|D)$). The threshold calculation shows $U(W)(1-P(\mathrm{doom}))P(W|D)<1$ is the stability condition, and with $U(W)\le 1.5$ and $P(\mathrm{doom})\ge 0.1$ the regime holds when $P(W|D)<0.74$. The report further argues that verification-based agreements keep $P(W|D)$ closer to 0.5 because detection triggers a rapid response, whereas ruleset-based agreements bind persistently but carry a tail risk of a scalable circumvention that would make defection catastrophically attractive; entering the treaty can remain incentive-compatible even for a leading state, and recovery from small violations is possible if the defector's lead stays moderate.
Load-bearing premise
The load-bearing premise is that a flexHEG ecosystem can be built and defended so that a first mover's defection is detected before it gains a decisive lead, i.e. that the realized $P(W|D)$ stays below about 0.74; the report asserts this is plausible but provides no model or measurement of detection latency or hardware tamper-resistance failure rates.
Editorial extensions
If this is right
- Comprehensive flexHEG agreements (especially verification-based ones) could be self-enforcing rather than requiring a supranational enforcer, as long as defection is detected quickly.
- If states value winning at most 1.5 times the cooperative outcome and believe an uncoordinated race has at least a 10% chance of catastrophe, a first defector would need better than roughly 74% odds of winning to rationally defect, a bar the report argues is hard to meet.
- A ruleset-based agreement gives states durable guarantees and prevents sudden unravelling, but the report warns it could be less stable if one state finds a scalable way to circumvent enforcement, because the defector's lead could grow large before others can respond.
- The same hardware can support targeted applications—verifying training runs at frontier data centers, limiting chip proliferation, providing assurances about military AI, and underpinning balance-of-power or sovereignty guarantees—so the stability result is not limited to one governance form.
- Managing the existing stock of non-flexHEG chips via registration and inspections matters: the stability logic assumes norm violations cannot be run at scale on unregistered compute.
Reading between the lines
- The stability conclusion is conditional on the realized rate of detection; no measurement or model in the report quantifies how quickly a secret training run would be caught, so the practical question is empirical: measuring detection latency under adversarial conditions would directly test the claim.
- A natural extension of the model is to intermediate defections—partial violations, delayed inspections, phased exits—where the binary cooperate/defect frame may understate the difficulty of sustaining cooperation; the report acknowledges this but does not model it.
- The threshold logic suggests a design principle beyond flexHEGs: policies that make defection observable early, such as mandatory real-time attestation or continuous monitoring, are more stabilizing than policies that merely raise the cost of defection but leave it undetectable.
- The same inequality could be applied to other compute-governance proposals (export controls, AI safety standards), since it identifies the three levers—value of winning, perceived catastrophe risk, and detection speed—that determine whether any international AI agreement is stable.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This report, the third part of a series on flexible hardware-enabled guarantees (flexHEGs), examines how hardware-based verification and enforcement mechanisms for AI-relevant chips could support international security governance. It surveys the prerequisites for an internationally trustworthy flexHEG ecosystem, including design standardization, production oversight, supply-chain tracking, and technical thresholds for AI-relevant chips. It then applies these mechanisms to four application areas: limiting proliferation, implementing AI safety norms, managing military AI risks, and supporting strategic stability. The central formal contribution is a two-player game-theoretic model in Appendix B, which derives a stability condition for comprehensive agreements: U(W)(1-P(doom))P(W|D)<1. With baseline assumptions U(W)≤1.5 and P(doom)≥0.1, cooperation is stable when the first defector's probability of winning the subsequent AI race satisfies P(W|D)<0.74. The paper asserts that a flexHEG verification regime could plausibly detect defection quickly enough to meet this bound, while documenting several model limitations, including the conflation of losing the race with catastrophic outcomes and the model's unrealistic behavior at very high P(doom).
Significance. If read as a conditional framework rather than an empirical demonstration, the paper has genuine value. It makes the strategic logic behind AI compute agreements explicit, identifies P(W|D) as the pivotal unknown, and provides a falsifiable threshold that could in principle be tested with better measurements of detection latency. The algebra in Appendix B is correct, the parameter choices and their uncertainties are disclosed, and the paper is notably candid about the limits of its own model. The main weakness is that the bridge from flexHEG detection mechanisms to an actual value or bound on P(W|D) is asserted with a single plausibility statement, so the central claim in the abstract overreaches the evidence provided. The report is best characterized as a transparent, well-contextualized policy analysis with a small formal core.
major comments (3)
- [Appendix B, 'Agreements Could Be Stable Given Reasonable Parameters' (p.55)] The central claim of the paper reduces to the inequality U(W)(1-P(doom))P(W|D)<1, and with U(W)≤1.5 and P(doom)≥0.1 the required bound is P(W|D)<0.74. The paper provides no model, measurement, or historical case connecting flexHEG detection mechanisms to P(W|D). The sentence 'it seems plausible that a flexHEG-based verification regime could detect defection quickly enough to make such high odds very unlikely' is an unsupported empirical assertion. Because P(W|D) is a function of detection latency, inspection coverage, and tamper-resistance failure rates, and these are the very properties the flexHEG ecosystem is supposed to provide, the demonstration is conditional on an unquantified engineering and policy claim. If detection is slower than the near-real-time scenario the paper gestures at, P(W|D) could exceed 0.74 and the cooperative equilibrium would not be stable even with reasonable U(W) and P(doom). Please either supply a substantive argument (for example, a timing analysis of the fastest plausible defection path, or a historical analogue from IAEA inspections or export-control verification) that bounds P(W|D), or explicitly reframe the result as 'stability holds provided P(W|D)<0.74' and temper the abstract's claim accordingly.
- [Appendix B, 'Payoff Structure' and 'Limitations and Extensions' (pp.54, 58)] The model sets the utility of losing the race to 0, identical to the utility of a catastrophic outcome. This normalization biases the model toward cooperation. The paper acknowledges the limitation but asserts that proper accounting would make defection 'somewhat more tempting' and is 'unlikely to overturn our key findings.' This robustness claim is not quantified. A simple calculation shows the bias is not negligible: if losing yields utility 0.5 instead of 0, the threshold P(W|D) at U(W)=1.5 and P(doom)=0.1 drops from 0.74 to about 0.61, a non-trivial reduction in the stability margin. The paper should either quantify the sensitivity of the threshold to the loser's utility or soften the claim that the qualitative conclusion is robust.
- [Appendix B, 'Impossible Conclusions' (p.59)] The appendix correctly notes that the model implies that for sufficiently high P(doom) no value of P(W|D) can make defection attractive, which is unrealistic because a defector could in principle trade off some of a dominant position to reduce the probability of catastrophe. The paper says this limitation is 'not particularly relevant' for the cases of interest, but no argument is given for why the parameter regime relevant to comprehensive flexHEG agreements avoids this artifact. Since the agreements analyzed here are aimed at high-risk scenarios, the boundary of this limitation should be clarified and its direction of bias stated explicitly in the main stability discussion.
minor comments (4)
- [Main text, 'How These Agreements Could be Stable' (p.39)] The phrase '(≥1.5 times)' is inconsistent with the analysis in Appendix B, which assumes U(W)≤1.5. Please correct the inequality sign to avoid confusing readers about the direction of the value assumption.
- [Appendix B, Figure 3 caption (p.56)] The caption states 'the agreement is stable for any values of U(W) and P(doom) above the blue line,' but the stability region should be described precisely as the set of (U(W),P(doom)) satisfying U(W)(1-P(doom)) < 1/P(W|D). Please clarify the axes and the region labeling.
- [Appendix B, 'Key Parameters' (p.53)] The definition of P(W|D) says only 'probability that a state would ultimately "win" the resulting AI race, conditional on being the first to defect'; the payoff formula multiplies by (1-P(doom)), indicating that the win probability is conditional on no catastrophic outcome. The definition should state this explicitly to avoid ambiguity.
- [Appendix B, 'Limitations and Extensions' (p.58)] Under 'Simplistic Utilities,' the paper says states would likely prefer being conquered by humans over destruction by AIs, but the follow-up sentence frames this only as making defection 'somewhat more tempting.' Given the concrete threshold sensitivity shown in the major comment above, the framing should acknowledge the quantitative impact on the P(W|D)| threshold.
Circularity Check
No circularity: the stability result in Appendix B is a conditional game-theoretic derivation, not a renamed fit or a self-citation chain.
full rationale
The paper's derivation chain is confined to Appendix B, where it defines a two-player cooperate/defect game with explicit parameters: U(C)=1, U(W), P(doom), and P(W|D). The expected utility of defecting is written as U(W)*(1-P(doom))*P(W|D), and stability is defined as this being less than 1. This is a standard expected-utility condition, not a hidden identity: the conclusion is the inequality, while P(W|D) is an exogenous parameter whose threshold is derived from the assumed U(W) and P(doom). The parameter values U(W) <= 1.5 and P(doom) >= 0.1 are justified by external arguments (linear resource valuation vs. status-quo preference, and the survey of ML researchers cited as [50]), not fitted to make the conclusion true. The statement that 'it seems plausible that a flexHEG-based verification regime could detect defection quickly enough to make such high odds very unlikely' is a qualitative plausibility judgment, not a fitted input renamed as a prediction. Citations to Parts I and II, and to the authors' earlier reports, provide background definitions and technical context, but the load-bearing stability inequality is derived within this paper from stated assumptions. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via citation, and no known result is merely renamed. The main substantive concerns are the unquantified detection-latency assumption and a small numerical slip in the main text (75% vs. the derived <0.74 threshold), but these are correctness and soundness risks, not circularity.
Assumptions & free parameters
free parameters (2)
- U(W), utility of winning the AI race =
assumed <= 1.5, with U(C) normalized to 1
- P(doom), probability of catastrophic outcome in an uncoordinated race =
assumed >= 0.1
assumptions (4)
- domain assumption A trustworthy and tamper-resistant flexHEG ecosystem can be designed, produced, and defended against state-level adversaries.
- domain assumption Uncoordinated frontier AI development carries at least a 10% probability of catastrophic outcomes.
- ad hoc to paper States are expected-utility maximizers and value winning an AI race at no more than 1.5 times the value of preserving the current balance of power.
- ad hoc to paper A two-player, binary cooperate-or-defect abstraction captures the strategic essence of international AI agreements.
Cite this review
Pith. "Pith review of International Security Applications of Flexible Hardware-Enabled Guarantees." pith.science (2026). https://pith.science/paper/3H64HKR2
@misc{pith2026250615100,
author = {Pith},
title = {Pith review of: International Security Applications of Flexible Hardware-Enabled Guarantees},
year = {2026},
howpublished = {\url{https://pith.science/paper/3H64HKR2}},
note = {Machine review of arXiv:2506.15100}
}
read the original abstract
As AI capabilities advance rapidly, flexible hardware-enabled guarantees (flexHEGs) offer opportunities to address international security challenges through comprehensive governance frameworks. This report examines how flexHEGs could enable internationally trustworthy AI governance by establishing standardized designs, robust ecosystem defenses, and clear operational parameters for AI-relevant chips. We analyze four critical international security applications: limiting proliferation to address malicious use, implementing safety norms to prevent loss of control, managing risks from military AI systems, and supporting strategic stability through balance-of-power mechanisms while respecting national sovereignty. The report explores both targeted deployments for specific high-risk facilities and comprehensive deployments covering all AI-relevant compute. We examine two primary governance models: verification-based agreements that enable transparent compliance monitoring, and ruleset-based agreements that automatically enforce international rules through cryptographically-signed updates. Through game-theoretic analysis, we demonstrate that comprehensive flexHEG agreements could remain stable under reasonable assumptions about state preferences and catastrophic risks. The report addresses critical implementation challenges including technical thresholds for AI-relevant chips, management of existing non-flexHEG hardware, and safeguards against abuse of governance power. While requiring significant international coordination, flexHEGs could provide a technical foundation for managing AI risks at the scale and speed necessary to address emerging threats to international security and stability.
Forward citations
Cited by 1 Pith paper
-
How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements
Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.
Reference graph
Works this paper leans on
-
[1]
Verifiable Compute White Paper,
“Verifiable Compute White Paper,” EQTY Lab, 2024. [Online]. Available: https://www.eqtylab.io/verifiable-compute-white-paper
work page 2024
-
[2]
Secure Enclaves for AI Evaluation,
A. Trask et al. , “Secure Enclaves for AI Evaluation,” OpenMined Blog. [Online]. Available: https://blog.openmined.org/secure-enclaves-for-ai-evaluation/
-
[3]
Anderson, Security engineering: a guide to building dependable distributed systems , 3rd edition
R. Anderson, Security engineering: a guide to building dependable distributed systems , 3rd edition. John Wiley & Sons, 2020. [Online]. Available: https://www.cl.cam.ac.uk/~rja14/book.html
work page 2020
-
[4]
Security Verification of the OpenTitan Hardware Root of Trust,
A. Meza, F. Restuccia, J. Oberg, D. Rizzo, and R. Kastner, “Security Verification of the OpenTitan Hardware Root of Trust,” IEEE Secur. Priv. , vol. 21, no. 3, pp. 27–36, May 2023, doi: 10.1109/MSEC.2023.3251954
arXiv 2023
-
[5]
The seL4 Microkernel – An Introduction v1.4,
G. Heiser, “The seL4 Microkernel – An Introduction v1.4,” The seL4 Foundation, Jan. 2025. [6] CHIPS Alliance, “Caliptra: A Datacenter System on a Chip (SoC) Root of Trust (RoT),” GitHub. [Online]. Available: spec.caliptra.io
work page 2025
-
[7]
B. Schneier, “The Legacy of DES,” Schneier on Security , Oct. 06, 2004. [Online]. Available: https://www.schneier.com/blog/archives/2004/10/the_legacy_of_d.html
work page 2004
-
[8]
How the NSA (may have) put a backdoor in RSA’s cryptography: A technical primer,
“How the NSA (may have) put a backdoor in RSA’s cryptography: A technical primer,” The Cloudflare Blog , Jan. 06, 2014. [Online]. Available: https://blog.cloudflare.com/how-the-nsa-may-have-put-a-backdoor-in-rsas-cryptography- a-technical-primer/
work page 2014
-
[9]
Location Verification for AI Chips,
A. Brass and O. Aarne, “Location Verification for AI Chips,” Institute for AI Policy and Strategy, Apr. 2024. [Online]. Available: https://www.iaps.ai/research/location-verification-for-ai-chips
work page 2024
Show all 50 references
- [10]
-
[11]
Statement by the Press Secretary
“Statement by the Press Secretary.” Office of the Press Secretary, Apr. 24, 2008. [Online]. Available: https://georgewbush-whitehouse.archives.gov/news/releases/2008/04/20080424-14.html
2008
-
[12]
International Verification and Intelligence,
J. M. Acton, “International Verification and Intelligence,” Intell. Natl. Secur. , vol. 29, no. 3, pp. 341–356, May 2014, doi: 10.1080/02684527.2014.895592
2014
-
[13]
Verification methods for international AI agreements,
A. R. Wasil, T. Reed, J. W. Miller, and P. Barnett, “Verification methods for international AI agreements,” Aug. 28, 2024, arXiv : arXiv:2408.16074. [Online]. Available: http://arxiv.org/abs/2408.16074
2024 arXiv
-
[14]
Mechanisms to Verify International Agreements About AI Development,
A. Scher and L. Thiergart, “Mechanisms to Verify International Agreements About AI Development,” Machine Intelligence Research Institute, Nov. 2024
2024
-
[15]
Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use; Updates and Corrections,
Bureau of Industry and Security, “Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use; Updates and Corrections,” Federal Register, 2023-23055 (88 FR 73458), Oct. 2023. [Online]. Available: https://www.federalr...
2023
-
[16]
With Smugglers and Front Companies, China Is Skirting American A.I. Bans,
A. Swanson and C. Fu, “With Smugglers and Front Companies, China Is Skirting American A.I. Bans,” The New York Times , Aug. 04, 2024. [Online]. Available: https://www.nytimes.com/2024/08/04/technology/china-ai-microchips.html
2024
-
[17]
Nvidia AI Chip Smuggling to China Becomes an Industry,
Q. Liu, “Nvidia AI Chip Smuggling to China Becomes an Industry,” The Information , Aug. 12, 2024. [Online]. Available: https://www.theinformation.com/articles/nvidia-ai-chip-smuggling-to-china-becomes-an- 43 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III in...
- [19]
-
[20]
Compute Thresholds are Ineffective,
D. W. Ball, “Compute Thresholds are Ineffective,” Hyperdimensional , Feb. 26, 2025. [Online]. Available: https://www.hyperdimensional.co/p/compute-thresholds-are-ineffective
2025
-
[21]
Increased Compute Efficiency and the Diffusion of AI Capabilities,
K. Pilz, L. Heim, and N. Brown, “Increased Compute Efficiency and the Diffusion of AI Capabilities,” Nov. 26, 2023, arXiv : arXiv:2311.15377. [Online]. Available: http://arxiv.org/abs/2311.15377
2023 arXiv
- [22]
-
[23]
Are Consumer GPUs a Problem for US Export Controls?,
E. Grunewald, “Are Consumer GPUs a Problem for US Export Controls?,” Institute for AI Policy and Strategy, May 2024. [Online]. Available: https://www.iaps.ai/research/are-consumer-gpus-a-problem-for-us-export-controls
2024
-
[24]
Framework for Artificial Intelligence Diffusion,
Bureau of Industry and Security, “Framework for Artificial Intelligence Diffusion,” Federal Register, 2025–00636 (90 FR 4544), Jan. 2025. [Online]. Available: https://www.federalregister.gov/documents/2025/01/15/2025-00636/framework-for-artif icial-intelligence-diffusion
2025
-
[25]
Secure, Governable Chips,
O. Aarne, T. Fist, and C. Withers, “Secure, Governable Chips,” Center for a New American Security, Jan. 2024. [Online]. Available: https://www.cnas.org/publications/reports/secure-governable-chips
2024
-
[26]
Chips for Peace: How the U.S. and Its Allies Can Lead on Safe and Beneficial AI,
C. O’Keefe, “Chips for Peace: How the U.S. and Its Allies Can Lead on Safe and Beneficial AI,” Lawfare , Jul. 10, 2024. [Online]. Available: https://www.lawfaremedia.org/article/chips-for-peace--how-the-u.s.-and-its-allies-can-lea d-on-safe-and-beneficial-ai
2024
-
[27]
International governance of advancing artificial intelligence,
N. Emery-Xu, R. Jordan, and R. Trager, “International governance of advancing artificial intelligence,” AI Soc. , Sep. 2024, doi: 10.1007/s00146-024-02050-7
2024 doi
-
[28]
Statement on AI Risk
“Statement on AI Risk.” Center for AI Safety, 2023. [Online]. Available: https://www.safe.ai/work/statement-on-ai-risk
2023
-
[29]
Global AI Governance Initiative,
“Global AI Governance Initiative,” Ministry of Foreign Affairs, the People’s Republic of China, Oct. 2023. [Online]. Available: https://www.mfa.gov.cn/eng/zy/gb/202405/t20240531_11367503.html
2023
-
[30]
International AI Safety Report,
Y. Bengio et al. , “International AI Safety Report,” DSIT 2025/001, 2025. [Online]. Available: https://www.gov.uk/government/publications/international-ai-safety-report-2025
2025
-
[31]
IDAIS-Venice Statement,
“IDAIS-Venice Statement,” Sep. 2024. [Online]. Available: https://idais.ai/dialogue/idais-venice/
2024
- [32]
- [33]
- [34]
- [35]
-
[36]
Discovering Latent Knowledge in Language Models Without Supervision,
C. Burns, H. Ye, D. Klein, and J. Steinhardt, “Discovering Latent Knowledge in Language Models Without Supervision,” Mar. 02, 2024, arXiv : arXiv:2212.03827. doi: 10.48550/arXiv.2212.03827. 44 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III
- [37]
- [38]
- [39]
-
[40]
Tools for Verifying Neural Models’ Training Data,
D. Choi, Y. Shavit, and D. Duvenaud, “Tools for Verifying Neural Models’ Training Data,” Jul. 02, 2023, arXiv : arXiv:2307.00682. [Online]. Available: http://arxiv.org/abs/2307.00682
2023 arXiv
- [41]
-
[42]
AI and International Stability: Risks and Confidence-Building Measures,
M. Horowitz and P. Scharre, “AI and International Stability: Risks and Confidence-Building Measures,” Center for a New American Security, Jan. 2021. [Online]. Available: https://www.cnas.org/publications/reports/ai-and-international-stability-risks-and-confid ence-building-measures
2021
-
[43]
Artificial General Intelligence’s Five Hard National Security Problems,
J. Mitre and J. B. Predd, “Artificial General Intelligence’s Five Hard National Security Problems,” RAND Corporation, Feb. 2025. [Online]. Available: https://www.rand.org/pubs/perspectives/PEA3691-4.html
2025
-
[44]
Situational Awareness: The Decade Ahead,
L. Aschenbrenner, “Situational Awareness: The Decade Ahead,” Jun. 2024. [Online]. Available: https://situational-awareness.ai/
2024
-
[45]
H. A. Kissinger, E. Schmidt, and C. Mundie, Genesis: Artificial Intelligence, Hope, and the Human Spirit . Little, Brown and Company, 2024
2024
- [46]
-
[47]
Rationalist Explanations for War,
J. D. Fearon, “Rationalist Explanations for War,” Int. Organ. , vol. 49, no. 3, pp. 379–414, 1995. [48] A. J. Coe and J. Vaynman, “Why Arms Control Is So Rare,” Am. Polit. Sci. Rev. , vol. 114, no. 2, pp. 342–355, May 2020, doi: 10.1017/S000305541900073X
1995 doi
-
[48]
Continuing to cooperate by adhering to the agreement's rules and verification requirements regarding AI development and deployment
-
[49]
The Anti-Ballistic Missile (ABM) Treaty at a Glance,
D. Kimball and K. Reif, “The Anti-Ballistic Missile (ABM) Treaty at a Glance,” Arms Control Association, Dec. 2020. [Online]. Available: https://www.armscontrol.org/factsheets/anti-ballistic-missile-abm-treaty-glance
2020
-
[50]
Thousands of AI Authors on the Future of AI,
K. Grace, H. Stewart, J. F. Sandkühler, S. Thomas, B. Weinstein-Raun, and J. Brauner, “Thousands of AI Authors on the Future of AI,” Apr. 30, 2024, arXiv : arXiv:2401.02843. doi: 10.48550/arXiv.2401.02843. 45 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III A...
2024 doi
-
[51]
Defecting by covertly violating the agreement in an attempt to gain an advantage in AI capabilities over the other player, for example by secretly conducting norm-violating training runs or deploying powerful models without required safety measures We assume that once a defect...
-
[52]
compute overhang
If P(doom) < 1 / U(W), stability requires: 54 Flexible Hardware-Enabled Guarantees | Part I | Part II | Part III Agreements Could Be Stable Given Reasonable Parameters As an example, we might assume that, for both players, U(W) ≤ 1.5, and P(doom) ≥ 0.1. U(W) ≤ 1.5 appears to b...
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.