Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds

T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper argues that AI technologists are engaging in 'safety revisionism' — redefining safety-engineering terms so that foundation models can enter military use without meeting established risk thresholds — and that this will weaken US…

desk verdict A pointed, well-sourced policy argument about 'safety revisionism' in AI defense work, though it overreaches by asserting without proof that traditional risk thresholds apply to foundation models. read the letter →

arxiv 2504.15088 v1 pith:UJWTMH75 submitted 2025-04-21 cs.CY

classification cs.CY
keywords safetyrevisionismAIriskthresholdsmilitaryfoundationmodelscasescapabilitiesevaluationtestverificationvalidationnationalsecurity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that because no democratic body has set agreed risk thresholds for AI systems, the people who build and sell them have ended up setting their own tolerances for acceptable harm — and they have set them low in order to speed foundation models into military use. The authors call this process 'safety revisionism': familiar terms from safety engineering, such as 'safety,' 'safety cases,' and 'red-teaming,' are redefined in ways that strip them of their original assurance power. The paper connects this redefinition to the historical risk thresholds developed for nuclear systems and defense standards, and claims that AI evaluations based on 'capabilities' rather than quantified risk cannot show that a system is safe enough for use in targeting or decision support. If the argument is right, the accelerated military adoption of AI will expose personnel and civilians to documented failure modes, create vulnerabilities adversaries can exploit, and in the end weaken rather than strengthen US national security.

What carries the argument

The load-bearing mechanism is the risk threshold — a quantified level of risk exposure above which action must be taken — together with the linguistic drift that detaches it from assurance practice. On one side stands the historical framework in which thresholds are fixed by societal deliberation and then substantiated by safety cases: structured arguments, backed by evidence, that a system is acceptably safe for a defined application in a defined environment. On the other side stand the contemporary 'frontier safety frameworks' that convert safety cases into capability targets and substitute red-teaming for boundary testing, with no threshold to test against. The comparison between those two sides is the gear that turns missing AI governance into lowered real-world safety for defense and, through precedent, for civilian critical infrastructure.

What would settle it

The central claim would fail if a military AI system that passed a current capabilities-evaluation framework was then assessed under a standard military system-safety analysis and met the same quantified risk thresholds required of the conventional system it replaces, such as a probability of a fatal hazardous event per deployment at or below the legacy-system threshold. It would also be weakened by a documented case in which an AI 'red-teaming' exercise uncovered a previously unknown hazard and measurably reduced estimated risk, showing that the redefined practice can carry real assurance weight.

Watch

Extended reading notes

Core claim

The core discovery is a mechanism with a name: 'safety revisionism.' When societally accepted risk tolerances for AI do not exist, technologists become the de facto arbiters of how much harm is acceptable, and they exercise that power by quietly replacing the vocabulary of safety engineering. Risk thresholds that were once set through democratic deliberation — such as the quantified fatal-risk criteria derived for nuclear power or the Safety Integrity Levels used in defense system safety — are superseded by 'frontier safety frameworks' that state unverifiable goals about model intent, such as not causing 'catastrophe,' without any numeric threshold a system can be tested against. The paper shows that the methodologies meant to substantiate safety, particularly safety cases, have been redefined into 'affirmative cases' and red-teaming exercises that cannot demonstrate risk reduction. The conclusion is that this trajectory is a self-fulfilling prophecy: the weakened thresholds adopted to preserve a supposed AI advantage are exactly what will compromise the safety and security of the military systems that depend on them.

Load-bearing premise

The argument assumes that the risk thresholds developed for earlier safety-critical technologies — nuclear power plants, military hardware — apply to AI-based military systems and that foundation models are 'no exception'; if AI failure modes are so different that those thresholds cannot meaningfully be transferred, the charge that AI firms are 'revising' safety loses its force.

Editorial extensions

If this is right

  • Military evaluation frameworks built on 'capabilities evaluations' cannot demonstrate that a foundation model is acceptably safe, because they never commit to a quantified risk threshold.
  • Foundation models in targeting and decision-support roles will carry documented failure modes — poor performance on non-English data, supply-chain attack surfaces, brittle safeguards — into settings where those failures can kill civilians and violate international humanitarian law.
  • If defense adoption sets the precedent, civilian safety-critical uses of AI will be judged against the same lowered bar, since military and civilian assurance standards influence each other.
  • Democratic bodies, not technologists, need to set explicit AI risk tolerances, including the number of societally accepted fatalities a given deployment assumes.
  • The national-security justification for accelerated adoption is self-defeating: a brittle AI system is an asset adversaries can exploit, so weakening thresholds to win an AI race undermines the very advantage the race is meant to secure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: its logic yields a testable prediction — among militaries that field AI-enabled targeting and decision support, those with weaker assurance thresholds will show more civilian casualties and friendly-fire incidents, then loosen thresholds further to keep pace.
  • A practical audit rule follows: any AI safety framework that cannot state a numeric threshold (for example, a maximum probability of a fatal hazardous event per deployment) is not a safety framework in the engineering sense, whatever its name; applying that test to current policy documents would give governance bodies a concrete checklist.
  • The same capabilities-evaluation logic is already visible outside defense, in proposals to run government administration with foundation models, so the missing-threshold problem is broader than the military case the paper examines.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper is an argumentative policy analysis arguing that, because no societally deliberated risk thresholds have been set for AI, AI technologists—mainly industry labs and self-styled 'AI safety' organizations—have been able to redefine safety terminology and substitute 'capabilities' and 'alignment' evaluations for traditional safety-engineering assurance. The authors reconstruct the history of risk thresholds from Chauncey Starr's nuclear-era work through MIL-STD-882 and Safety Integrity Levels, then argue that current AI evaluation frameworks (Google DeepMind's Frontier Safety Framework, Anthropic's Responsible Scaling Policy, Task Force Lima, the UK AI Security Institute) hollow out the meaning of 'safety', 'safety case', and 'red-teaming'. They use the Gaza deployments and supply-chain vulnerabilities to argue that the resulting absence of enforceable thresholds enables deployment of foundation models in military contexts at lowered safety and security thresholds, ultimately undermining US national security and international humanitarian law. The paper concludes by calling for democratic deliberation to set AI risk thresholds and for preserving traditional TEVV standards in military AI evaluation.

Significance. If the argument holds, the paper makes a valuable contribution by connecting AI safety discourse to the history and vocabulary of safety engineering and by identifying a concrete mechanism—terminological revisionism—through which existing assurance standards can be bypassed. Its strengths are the historically grounded account of risk thresholds, the concrete documentation of military AI deployment failures, and a clear normative proposal: preserve established assurance frameworks and set AI risk thresholds through democratic deliberation rather than industry self-assessment. The paper is not an empirical study and does not ship machine-checked proofs or reproducible code; its contribution is conceptual and evidentiary. The central claim is defensible but requires an important qualification: the transferability of quantitative risk thresholds to foundation models is asserted rather than demonstrated, and the national-security conclusion is hedged in some places but stated categorically in others.

major comments (3)
  1. [Section 2 / Table 1] The paper's central charge of 'lowered' risk thresholds presupposes that established quantitative thresholds (Starr's 10^-4 deaths per person per year, MIL-STD-882E categories, SIL levels) can be populated for foundation models. The paper asserts this ('their risk categories fall firmly within the scope of traditional risk frameworks') but never shows how a probability of a hazardous event can be assigned to a context-dependent, nondeterministic model at assurance time. Without such a demonstration, the 'lowered threshold' claim is not established; the situation may instead be that no applicable threshold has been operationalized. This distinction is load-bearing for the 'revisionism' charge, because mislabeling an absence of a standard as a weakening of a standard changes the normative force of the argument. Please either provide a worked example for a concrete use case (e.g., target nomination or intelligence analysis) or reframe the argument as 'no applicable risk threshold has been operationalized'.
  2. [Section 4, final paragraph] The claim that AI systems are 'deployed at levels far below the risk thresholds that would be deemed acceptable through standardized safety processes' is not supported by a quantitative comparison. The cited evidence—Arabic mistranslation, cell-tower-based civilian casualty estimates, and evasion failures—demonstrates qualitative failures, but the paper provides no probability-of-failure estimates, exposure analysis, or comparison against any specific threshold. As stated, the claim is not falsifiable. It should either be quantified for a specific deployment scenario or explicitly presented as a qualitative judgment about the absence of demonstrated compliance.
  3. [Sections 1 and 5] The title's 'self-fulfilling prophecy' and the national-security conclusion are causal claims that the evidence does not fully support. The paper shows that some military AI uses have failed and that some frameworks use nonstandard terminology, but it does not systematically consider alternative explanations for the observed outcomes, such as bureaucratic incentives, technical immaturity, or good-faith disagreements about how to adapt assurance methods to AI. The hedging in Section 1 ('may be precisely what disadvantages') is honest, but the categorical conclusion in Section 5 ('will imperil' and 'will result in a significant civilian death toll') goes beyond the evidence presented. Please separate the well-supported claim that current frameworks do not satisfy traditional assurance standards from the more speculative claim that this trajectory will compromise US national security.
minor comments (6)
  1. [Section 2] The paper defines risk tolerance and risk threshold as distinct concepts but later uses them interchangeably (e.g., Section 3: 'risk thresholds are derived from risk tolerances'). Please maintain the distinction consistently.
  2. [Section 3] The statement that a safety case 'is not intended to produce any 'targets' or thresholds' is too categorical; some safety-case methodologies incorporate quantitative safety requirements as part of the argument. The substantive criticism of Google DeepMind's framework does not depend on this categorical claim, so it can be softened without harming the argument.
  3. [Section 3.1] The characterization of Task Force Lima as taking a 'capabilities evaluation' approach relies on the task force's executive summary [55]. Please verify the citation and, if possible, quote the relevant language, since the executive summary may not be publicly accessible.
  4. [Section 4] The claim that foundation models have 'poor accuracy with non-English languages, especially for Arabic' is supported by general LLM cultural-bias studies [61, 62]; the connection to the specific IDF systems described in [9] should be made more explicit, since those systems may include non-LLM components.
  5. [Section 5] The phrase 'existential risks that are very real and present' sits awkwardly with the paper's earlier rejection of 'speculative existential risks' in Sections 1 and 2. Please clarify whether this refers to concrete military harms rather than the speculative AI-existential-risk literature.
  6. [Throughout] Please standardize spelling and formatting inconsistencies, including 'DOD' vs 'DoD' and 'LAWs' vs 'LAWS', and check the column headings in Table 1, which appear truncated.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a standards-based policy critique whose argument is externally anchored, with no fitted parameters or self-referential derivation.

full rationale

The paper does not present a formal derivation, a fitted model, or a predictive claim that reduces to its own inputs by construction. Its central charge—that AI technologists have engaged in 'safety revisionism' by substituting ill-defined safety terminology for established frameworks—is argued by comparison to external standards (MIL-STD-882E, Safety Integrity Levels, Starr's risk analysis, TEVV) and to documented deployment failures (e.g., IDF systems, red-teaming brittleness). The authors' self-citations, such as [50], [51], [52], and [53], are used as supporting background references on AI risk and military proliferation rather than as load-bearing uniqueness theorems or as definitions that presuppose the conclusion. The closest thing to a potentially circular premise is the assertion that foundation models' failure modes 'fall firmly within the scope of traditional risk frameworks' (Section 2), but this is an unproven assumption about applicability, not a self-referential derivation; it could be challenged as a correctness or evidence concern, not as circularity. No equation, threshold, or evaluation result is shown to be equivalent to its own input, and no fitted parameter is renamed as a prediction. The paper is self-contained as a policy argument and scores 0 on circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on normative and domain assumptions about who should set risk thresholds and whether traditional safety frameworks transfer to AI. No free parameters or invented physical entities are introduced.

assumptions (3)
  • domain assumption Societal risk thresholds for technology should be determined through democratic deliberation rather than by technologists.
    The entire argument depends on the premise that technologists setting risk thresholds is illegitimate; this is asserted via Starr's work but not independently defended.
  • domain assumption Established safety-engineering frameworks (e.g., MIL-STD-882, SIL, safety cases) are applicable to AI-based military systems and should be preserved.
    The paper assumes foundation models must comply with the same assurance frameworks as other defense systems; it does not prove that AI systems share the relevant properties.
  • domain assumption The cited examples of AI military failure (e.g., IDF targeting in Gaza) are representative and accurately reported.
    The argument's empirical weight rests on these cases; they are drawn from journalism and advocacy rather than systematic measurement, but are treated as evidence of systemic risk.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds." pith.science (2026). https://pith.science/paper/UJWTMH75

@misc{pith2026250415088,
  author       = {Pith},
  title        = {Pith review of: Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UJWTMH75}},
  note         = {Machine review of arXiv:2504.15088}
}
read the original abstract

Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first risk analyses, such as those devised for nuclear systems, cemented societally accepted risk thresholds against which safety-critical and defense systems are now evaluated. But today, the appropriate risk tolerances for AI systems have yet to be agreed on by global governing efforts, despite the need for democratic deliberation regarding the acceptable levels of harm to human life. Absent such AI risk thresholds, AI technologists-primarily industry labs, as well as "AI safety" focused organizations-have instead advocated for risk tolerances skewed by a purported AI arms race and speculative "existential" risks, taking over the arbitration of risk determinations with life-or-death consequences, subverting democratic processes. In this paper, we demonstrate how such approaches have allowed AI technologists to engage in "safety revisionism," substituting traditional safety methods and terminology with ill-defined alternatives that vie for the accelerated adoption of military AI uses at the cost of lowered safety and security thresholds. We explore how the current trajectory for AI risk determination and evaluation for foundation model use within national security is poised for a race to the bottom, to the detriment of the US's national security interests. Safety-critical and defense systems must comply with assurance frameworks that are aligned with established risk thresholds, and foundation models are no exception. As such, development of evaluation frameworks for AI-based military systems must preserve the safety and security of US critical and defense infrastructure, and remain in alignment with international humanitarian law.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harmonizing AI Safety Thresholds

    cs.AI 2026-07 conditional novelty 5.0 of 10

    The authors propose harmonized AI capability floors: non-zero full-chain TLO cyber completion triggers safeguards, and AI progress at 5× trend for 3 months triggers safeguards, with biorisk left as a diagnostic.

Reference graph

Works this paper leans on

106 extracted references · 55 canonical work pages · cited by 1 Pith paper

  1. [1]

    ‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza, April 2024

    Yuval Abraham. ‘Lavender’: The AI machine directing Israel’s bombing spree in Gaza, April 2024. URL https://www.972mag.com/ lavender-ai-israeli-army-gaza/

  2. [2]

    Canada launches first AI strategy for federal public service - Global Government Forum, April 2025

    Jack Aldane. Canada launches first AI strategy for federal public service - Global Government Forum, April 2025. URL https: //www.globalgovernmentforum.com/canada-launches-first-ai-strategy-for-federal-public-service/

  3. [3]

    Trump Can Keep America’s AI Advantage

    Dario Amodei and Matt Pottinger. Trump Can Keep America’s AI Advantage. Wall Street Journal , January 2025. URL https: //www.wsj.com/opinion/trump-can-keep-americas-ai-advantage-china-chips-data-eccdce91

  4. [4]

    Frontier AI Regulation: Managing Emerging Risks to Public Safety

    Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf. Fr...

  5. [5]

    Anduril Partners with OpenAI to Advance U.S

    Anduril. Anduril Partners with OpenAI to Advance U.S. Artificial Intelligence Leadership and Protect U.S. and Allied Forces, December

  6. [6]

    Anthropic’s Responsible Scaling Policy, September 2023

    Anthropic. Anthropic’s Responsible Scaling Policy, September 2023. URL https://www.anthropic.com/news/anthropics-responsible- scaling-policy

  7. [7]

    Security-informed safety, January 2023

    National Protective Security Authority. Security-informed safety, January 2023. URL https://www.npsa.gov.uk/security-informed- safety

  8. [8]

    Why policy makers should beware claims of new ‘arms races’

    Haydn Belfield and Christian Ruhl. Why policy makers should beware claims of new ‘arms races’. Bulletin of the Atomic Scientists , July

Show all 106 references
  1. [9]

    As Israel uses US-made AI models in war, concerns arise about tech’s role in who lives and who dies

    Michael Biesecker, Sam Mednick, and Garance Burke. As Israel uses US-made AI models in war, concerns arise about tech’s role in who lives and who dies. AP News , February 2025. URL https://apnews.com/article/israel-palestinians-ai-technology- 737bc17af7b03e98c29cec4e15d0f108

  2. [10]

    Assurance of AI Systems From a Dependability Perspective

    Robin Bloomfield and John Rushby. Assurance of AI Systems From a Dependability Perspective. arXiv, August 2024. doi: 10.48550/ arXiv.2407.13948. URL http://arxiv.org/abs/2407.13948

  3. [11]

    Defeaters and eliminative argumentation in assurance 2.0

    Robin Bloomfield, Kate Netkachova, and John Rushby. Defeaters and eliminative argumentation in assurance 2.0. arXiv, 2024. doi: 10.48550/arXiv.2405.15800. URL https://arxiv.org/abs/2405.15800

  4. [12]

    How a billionaire-backed network of AI advisers took over Washington

    Brendan Bordelon. How a billionaire-backed network of AI advisers took over Washington. POLITICO, October 2023. URL https://www.politico.com/news/2023/10/13/open-philanthropy-funding-ai-policy-00121362

  5. [13]

    Key Congress staffers in AI debate are funded by tech giants like Google and Microsoft

    Brendan Bordelon. Key Congress staffers in AI debate are funded by tech giants like Google and Microsoft. POLITICO, March 2023. URL https://www.politico.com/news/2023/12/03/congress-ai-fellows-tech-companies-00129701

  6. [14]

    The Risks of Integrating Generative AI into Weapon Systems

    Vincent Boulanin. The Risks of Integrating Generative AI into Weapon Systems. Technical report, GC REAIM Expert Policy Note Series, April 2025

  7. [15]

    Michael Brenes and William D. Hartung. Private Finance and the Quest to Remake Modern Warfare, June 2024. URL https: //quincyinst.org/research/private-finance-and-the-quest-to-remake-modern-warfare/

  8. [16]

    Britain dances to JD Vance’s tune as it renames AI institute

    Tom Bristow. Britain dances to JD Vance’s tune as it renames AI institute. POLITICO, February 2025. URL https://www.politico.eu/ article/jd-vance-britain-ai-safety-institute-aisi-security/

  9. [17]

    Brown, Jordan Schneider, Anca D

    Daniel S. Brown, Jordan Schneider, Anca D. Dragan, and Scott Niekum. Value Alignment Verification. arXiv, June 2021. doi: 10.48550/arXiv.2012.01557. URL http://arxiv.org/abs/2012.01557. arXiv:2012.01557

  10. [18]

    Extracting training data from diffusion models

    Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, and Eric Wallace. Extracting training data from diffusion models. In Proceedings of the 32nd USENIX Conference on Security Symposium (SEC ’23). USENIX Ass...

  11. [19]

    Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr

    Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. Poisoning Web-Scale Training Datasets is Practical. In 2024 IEEE Symposium on Security and Privacy (SP), pages 407–4...

  12. [20]

    Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr

    Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Itay Yona, Eric Wallace, David Rolnick, and Florian Tramèr. Stealing Part of a Production Language Model. ...

  13. [21]

    ÓhÉigeartaigh

    Stephen Cave and Seán S. ÓhÉigeartaigh. An AI Race for Strategic Advantage: Rhetoric and Risks. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society , pages 36–40, New Orleans LA USA, December 2018. ACM. ISBN 9781450360128. doi: 10.1145/3278721.3278780. UR...

  14. [22]

    Brian J. Chen. Dispelling Myths of AI and Efficiency. Data & Society, 2025. URL https://datasociety.net/library/dispelling-myths-of-ai- and-efficiency/?trk=feed_main-feed-card_feed-article-content

  15. [23]

    Safety Cases: How to Justify the Safety of Advanced AI Systems

    Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen. Safety Cases: How to Justify the Safety of Advanced AI Systems. arXiv, March 2024. doi: 10.48550/arXiv.2403.10462. URL http://arxiv.org/abs/2403.10462

  16. [24]

    Nuclear Regulatory Commission

    U.S. Nuclear Regulatory Commission. Safety Goals for Nuclear Power Plant Operation. U.S. Nuclear Regulatory Commission, May

  17. [25]

    Marine Corps Warfighting Laboratory: Experiment Division

    The United States Marine Corps. Marine Corps Warfighting Laboratory: Experiment Division. URL https://www.mcwl.marines.mil/ Divisions/Experiment/

  18. [26]

    The AI that Wasn’t There: Global Order and the (Mis)Perception of Powerful AI

    Mary (Missy) Cummings. The AI that Wasn’t There: Global Order and the (Mis)Perception of Powerful AI. Policy Roundtable: Artificial Intelligence and International Security, Texas National Security Review , June 2020. URL https://tnsr.org/roundtable/policy-roundtable- artificia...

  19. [27]

    Cummings and Ben Bauchwitz

    M.L. Cummings and Ben Bauchwitz. Identifying Research Gaps through Self-Driving Car Data Analysis.IEEE Transactions on Intelligent Vehicles, pages 1–10, 2024. ISSN 2379-8904. doi: 10.1109/TIV.2024.3506936. URL https://ieeexplore.ieee.org/document/10778107/ ?arnumber=10778107

  20. [28]

    Introducing the Frontier Safety Framework, April 2025

    Google DeepMind. Introducing the Frontier Safety Framework, April 2025. URL https://deepmind.google/discover/blog/introducing- the-frontier-safety-framework/

  21. [29]

    Behind China’s Plans to Build AI for the World

    Bill Drexel and Hannah Kelley. Behind China’s Plans to Build AI for the World. POLITICO, November 2023. URL https://www.politico. com/news/magazine/2023/11/30/china-global-ai-plans-00129160

  22. [30]

    Israel built an ‘AI factory’ for war

    Elizabeth Dwoskin. Israel built an ‘AI factory’ for war. It unleashed it in Gaza. The Washington Post , December 2024. URL https://www.washingtonpost.com/technology/2024/12/29/ai-israel-war-gaza-idf/

  23. [31]

    Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

    Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfi...

  24. [32]

    Paul M. Grant. Chauncey Starr (1912–2007). Nature, 447(7146):789–789, June 2007. ISSN 1476-4687. doi: 10.1038/447789a. URL https://www.nature.com/articles/447789a

  25. [33]

    AI Safety: Navigating the Expanding Landscape of Potential Harms

    Ibrahim Habli and John Alexander McDermid. AI Safety: Navigating the Expanding Landscape of Potential Harms. Safety-Critical Systems Club Newsletter, May 2024. URL https://eprints.whiterose.ac.uk/id/eprint/213035/

  26. [34]

    The BIG Argument for AI Safety Cases

    Ibrahim Habli, Richard Hawkins, Colin Paterson, Philippa Ryan, Yan Jia, Mark Sujan, and John McDermid. The BIG Argument for AI Safety Cases. arXiv, March 2025. doi: 10.48550/arXiv.2503.11705. URL http://arxiv.org/abs/2503.11705

  27. [35]

    The Artificial Intelligence Arms Race: Trends and World Leaders in Autonomous Weapons Development

    Justin Haner and Denise Garcia. The Artificial Intelligence Arms Race: Trends and World Leaders in Autonomous Weapons Development. Global Policy, 10(3):331–337, September 2019. ISSN 1758-5880, 1758-5899. doi: 10.1111/1758-5899.12713. URL https://onlinelibrary. wiley.com/doi/10...

  28. [36]

    The Nuclear-Level Risk of Superintelligent AI

    Dan Hendrycks and Eric Schmidt. The Nuclear-Level Risk of Superintelligent AI. TIME, March 2025. URL https://time.com/7265056/ nuclear-level-risk-of-superintelligent-ai/

  29. [37]

    Read: JD Vance’s full speech on AI and the EU.The Spectator, February 2025

    Coffee House. Read: JD Vance’s full speech on AI and the EU.The Spectator, February 2025. URL https://www.spectator.co.uk/article/read- jd-vances-full-speech-on-ai-and-the-eu/

  30. [38]

    Tackling AI Security Risks to Unleash Growth and Deliver Plan for Change

    AI Security Institute. Tackling AI Security Risks to Unleash Growth and Deliver Plan for Change. Department for Science, Innovation and Technology, February 2025. URL https://www.gov.uk/government/news/tackling-ai-security-risks-to-unleash-growth-and-deliver- plan-for-change

  31. [39]

    Safety Cases at AISI

    Geoffrey Irving. Safety Cases at AISI. AI Security Institute, August 2024. URL https://www.aisi.gov.uk/work/safety-cases-at-aisi

  32. [40]

    The Fifth Branch

    Sheila Jasanoff. The Fifth Branch. Harvard University Press, 1994. URL https://www.hup.harvard.edu/books/9780674300620

  33. [41]

    Containing the Atom: Sociotechnical Imaginaries and Nuclear Power in the United States and South Korea

    Sheila Jasanoff and Sang-Hyun Kim. Containing the Atom: Sociotechnical Imaginaries and Nuclear Power in the United States and South Korea. Minerva, 47(2):119–146, June 2009. ISSN 0026-4695, 1573-1871. doi: 10.1007/s11024-009-9124-4. URL http: //link.springer.com/10.1007/s11024...

  34. [42]

    Under the radar? examining the evaluation of foundation models

    Elliot Jones, Mahi Hardalupas, and William Agnew. Under the radar? examining the evaluation of foundation models. Ada Lovelace Institute, July 2024. URL https://www.adalovelaceinstitute.org/wp-content/uploads/2024/09/Ada-Lovelace-Institute-Under-the-radar- 230924.pdf

  35. [43]

    2023 Landscape: Confronting Tech Power

    Amba Kak and Sarah Myers West. 2023 Landscape: Confronting Tech Power. AI Now Institute, April 2023. URL https://ainowinstitute. org/wp-content/uploads/2023/04/AI-Now-2023-Landscape-Report-FINAL.pdf

  36. [44]

    Nidhi Kalra and Susan M. Paddock. Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability? Technical report, RAND Corporation, April 2016. URL https://www.rand.org/pubs/research_reports/RR1478.html

  37. [45]

    Measurement challenges in AI catastrophic risk governance and safety frameworks

    Atoosa Kasirzadeh. Measurement challenges in AI catastrophic risk governance and safety frameworks. Tech Policy Press, Sep 2024. URL https://www.techpolicy.press/measurement-challenges-in-ai-catastrophic-risk-governance-and-safety-frameworks/

  38. [46]

    Dependability analysis of safety critical systems: Issues and challenges

    Raj kamal Kaur, Babita Pandey, and Lalit Kumar Singh. Dependability analysis of safety critical systems: Issues and challenges. Annals of Nuclear Energy, 120:127–154, October 2018. ISSN 0306-4549. doi: 10.1016/j.anucene.2018.05.027. URL https://www.sciencedirect. com/science/a...

  39. [47]

    Doomed to Repeat History? Lessons from the Crypto Wars of the 1990s

    Danielle Kehl, Andi Wilson, and Kevin Bankston. Doomed to Repeat History? Lessons from the Crypto Wars of the 1990s. URL http: //newamerica.org/cybersecurity-initiative/policy-papers/doomed-to-repeat-history-lessons-from-the-crypto-wars-of-the-1990s/

  40. [48]

    Elon Musk Ally Tells Staff ’AI-First’ Is the Future of Key Government Agency

    Makena Kelly. Elon Musk Ally Tells Staff ’AI-First’ Is the Future of Key Government Agency. Wired, 2025. ISSN 1059-1028. URL https://www.wired.com/story/elon-musk-lieutenant-gsa-ai-agency/

  41. [49]

    A Systematic Approach to Safety Case Management

    Tim Kelly. A Systematic Approach to Safety Case Management. SAE Transactions, 113:257–266, 2004. ISSN 0096-736X. URL https://www.jstor.org/stable/44699541

  42. [50]

    How AI Can Be Regulated Like Nuclear Energy

    Heidy Khlaaf. How AI Can Be Regulated Like Nuclear Energy. TIME, October 2023. URL https://time.com/6327635/ai-needs-to-be- regulated-like-nuclear-weapons/

  43. [51]

    Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems

    Heidy Khlaaf. Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems. Trail of Bits , March 2023. URL https://www.trailofbits.com/documents/Toward_comprehensive_risk_assessments.pdf

  44. [52]

    A Hazard Analysis Framework for Code Synthesis Large Language Models

    Heidy Khlaaf, Pamela Mishkin, Joshua Achiam, Gretchen Krueger, and Miles Brundage. A Hazard Analysis Framework for Code Synthesis Large Language Models. arXiv, July 2022. URL https://arxiv.org/abs/2207.14157v1

  45. [53]

    Mind the Gap: Foundation Models and the Covert Proliferation of Military Intelligence, Surveillance, and Targeting

    Heidy Khlaaf, Sarah Myers West, and Meredith Whittaker. Mind the Gap: Foundation Models and the Covert Proliferation of Military Intelligence, Surveillance, and Targeting. arXiv, October 2024. doi: 10.48550/arXiv.2410.14831. URL http://arxiv.org/abs/2410.14831

  46. [54]

    Nato use of civil standards, October 2018

    Steven Lapsley. Nato use of civil standards, October 2018. URL https://www.dsp.dla.mil/Portals/26/Documents/Publications/ Conferences/2018/2018%20International%20Standardization%20Workshop/20181031-Item4-UK_NATO_UseOfCivilStandards- IntlStdznWorkshop_Lapsely.pdf?ver=2018-11-06...

  47. [55]

    Task Force Lima Executive Summary

    Task Force Lima. Task Force Lima Executive Summary. Technical report, US Department of Defense, December 2024

  48. [56]

    MacKenzie

    Donald A. MacKenzie. Inventing Accuracy: A Historical Sociology of Nuclear Missile Guidance . MIT Press, 1990. ISBN 9780262132589

  49. [57]

    Pasquale

    Gianclaudio Malgieri and Frank A. Pasquale. From transparency to justification: Toward ex ante accountability for ai. SSRN Electronic Journal, 2022. doi: 10.2139/ssrn.4099657

  50. [58]

    AI should replace some work of civil servants, Starmer to announce

    Rowena Mason and Rowena Mason Whitehall editor. AI should replace some work of civil servants, Starmer to announce. The Guardian, March 2025. ISSN 0261-3077. URL https://www.theguardian.com/technology/2025/mar/12/ai-should-replace-some-work- of-civil-servants-under-new-rules-k...

  51. [59]

    Upstream and Downstream AI Safety: Both on the Same River? arXiv, December 2024

    John McDermid, Yan Jia, and Ibrahim Habli. Upstream and Downstream AI Safety: Both on the Same River? arXiv, December 2024. doi: 10.48550/arXiv.2501.05455. URL http://arxiv.org/abs/2501.05455

  52. [60]

    Anthropic’s Dario Amodei: Democracies must maintain the lead in AI

    Madhumita Murgia. Anthropic’s Dario Amodei: Democracies must maintain the lead in AI. Financial Times, December 2024. URL https://www.ft.com/content/e75e3388-4700-413d-ab67-778410c2d977

  53. [61]

    On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena

    Tarek Naous and Wei Xu. On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena. arXiv, January 2025. doi: 10.48550/arXiv.2501.04662. URL http://arxiv.org/abs/2501.04662

  54. [62]

    Having Beer after Prayer? Measuring Cultural Bias in Large Language Models

    Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. Having Beer after Prayer? Measuring Cultural Bias in Large Language Models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (...

  55. [63]

    AI as Normal Technology

    Arvind Narayanan and Sayash Kapoor. AI as Normal Technology. Columbia University, The Knight First Amendment Institute , 2025. URL https://www.aisnakeoil.com/p/ai-as-normal-technology

  56. [64]

    Feder Cooper, Daphne Ippolito, Christopher A

    Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee. Scalable Extraction of Training Data from (Production) Language Models, 2023. URL 18 • Khlaaf and...

  57. [65]

    NIST. U.S. AI Safety Institute Establishes New U.S. Government Taskforce to Collaborate on Research and Testing of AI Models to Manage National Security Capabilities & Risks, November 2024. URL https://www.nist.gov/news-events/news/2024/11/us-ai-safety- institute-establishes-n...

  58. [66]

    Implementation of improved Air System Safety Case Regulation - RA 1205, 2019

    Ministry of Defence. Implementation of improved Air System Safety Case Regulation - RA 1205, 2019. URL https://www.gov.uk/ government/news/implementation-of-improved-air-system-safety-case-regulation-ra-1205

  59. [67]

    Department of Defense

    The U.S. Department of Defense. Standard Practice: System Safety. Technical Report MIL-STD-882E, The U.S. Department of Defense, May 2012

  60. [68]

    DoD Modeling and Simulation (M&S) Verification, Validation, and Accreditation (VV&A),

    U.S. Department of Defense. DoD Instruction 5000.61, “DoD Modeling and Simulation (M&S) Verification, Validation, and Accreditation (VV&A), ”. U.S. Department of Defense, September 2024. URL https://www.esd.whs.mil/Portals/54/Documents/DD/issuances/dodi/ 500061p.pdf

  61. [69]

    Artificial Intelligence Safety and Security Board, 2025

    Department of Homeland Security. Artificial Intelligence Safety and Security Board, 2025. URL https://www.dhs.gov/artificial- intelligence-safety-and-security-board

  62. [70]

    2024 ICRC IHL Challenges Report | ICRC, September 2024

    International Committee of the Red Cross. 2024 ICRC IHL Challenges Report | ICRC, September 2024. URL https://www.icrc.org/en/ report/2024-icrc-report-ihl-challenges

  63. [71]

    The Safety Goals of the U.S

    David Okrent. The Safety Goals of the U.S. Nuclear Regulatory Commission. Science, 236(4799):296–300, April 1987. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.3563510. URL https://www.science.org/doi/10.1126/science.3563510

  64. [72]

    Nathan Matias

    Amy Orben and J. Nathan Matias. Fixing the science of digital technology harms. Science, 388(6743):152–155, April 2025. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.adt6807. URL https://www.science.org/doi/10.1126/science.adt6807

  65. [73]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...

  66. [74]

    Palantir IR, 2024

    Palantir. Palantir IR, 2024. URL https://investors.palantir.com/news-details/2024/Anthropic-and-Palantir-Partner-to-Bring-Claude-AI- Models-to-AWS-for-U.S.-Government-Intelligence-and-Defense-Operations/

  67. [75]

    Red Teaming Language Models with Language Models

    Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. Red Teaming Language Models with Language Models. arXiv, February 2022. doi: 10.48550/arXiv.2202.03286. URL http: //arxiv.org/abs/2202.03286

  68. [76]

    Chauncey Starr: A personal memoir, June 2007

    POWER. Chauncey Starr: A personal memoir, June 2007. URL https://www.powermag.com/chauncey-starr-a-personal-memoir/

  69. [77]

    Introducing Primer Delta: Our next-gen AI-native platform to transform information overload into decision advantage, April

    Primer. Introducing Primer Delta: Our next-gen AI-native platform to transform information overload into decision advantage, April

  70. [78]

    Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! arXiv, October 2023

    Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! arXiv, October 2023. URL https://arxiv.org/abs/2310.03693v1

  71. [79]

    Concrete Problems in AI Safety, Revisited.arXiv, December 2023

    Inioluwa Deborah Raji and Roel Dobbe. Concrete Problems in AI Safety, Revisited.arXiv, December 2023. doi: 10.48550/arXiv.2401.10899. URL http://arxiv.org/abs/2401.10899

  72. [80]

    Jane O. Rathbun. Don guidance on the use of generative artificial intelligence and large language models. Department of Navy Chief Information Officer, 2023. URL https://www.doncio.navy.mil/ContentView.aspx?id=16442

  73. [81]

    arms race

    Heather M. Roff. The frame problem: The AI “arms race” isn’t one. Bulletin of the Atomic Scientists , 75(3):95–98, May 2019. ISSN 0096-3402, 1938-3282. doi: 10.1080/00963402.2019.1604836. URL https://www.tandfonline.com/doi/full/10.1080/00963402.2019.1604836

  74. [82]

    Fintech Founder Charged with Fraud after “AI” Shopping App Found to be Powered by Humans in the Philippines

    Charles Rollet. Fintech Founder Charged with Fraud after “AI” Shopping App Found to be Powered by Humans in the Philippines. TechCrunch, Apr 2025. URL https://techcrunch.com/2025/04/10/fintech-founder-charged-with-fraud-after-ai-shopping-app-found-to- be-powered-by-humans-in-t...

  75. [83]

    Anatomy of an AI Coup

    Eryk Salvaggio. Anatomy of an AI Coup. Tech Policy Press, February 2025. URL https://techpolicy.press/anatomy-of-an-ai-coup

  76. [84]

    Silicon Valley Goes to War: Artificial Intelligence, Weapons Systems, and Moral Agency

    Elke Schwarz and DePaul University. Silicon Valley Goes to War: Artificial Intelligence, Weapons Systems, and Moral Agency. Philosophy Today, 65(3):549–569, 2021. ISSN 0031-8256. doi: 10.5840/philtoday2021519407. URL http://www.pdcnet.org/oom/service? url_ver=Z39.88-2004&rft_v...

  77. [85]

    AI Red Teaming: Applying Software TEVV for AI Evaluations | CISA, November 2024

    Jonathan Spring and Divjot Singh Bawa. AI Red Teaming: Applying Software TEVV for AI Evaluations | CISA, November 2024. URL https://www.cisa.gov/news-events/news/ai-red-teaming-applying-software-tevv-ai-evaluations. Safety Co-Option and Compromised National Security: The Self-...

  78. [86]

    Risk Criteria for Nuclear Power Plants: A Pragmatic Proposal

    Chauncey Starr. Risk Criteria for Nuclear Power Plants: A Pragmatic Proposal. Risk Analysis, 1(2):113–120, June 1981. ISSN 0272-4332, 1539-6924. doi: 10.1111/j.1539-6924.1981.tb01406.x. URL https://onlinelibrary.wiley.com/doi/10.1111/j.1539-6924.1981.tb01406.x

  79. [87]

    Risks of Risk Decisions

    Chauncey Starr and Chris Whipple. Risks of Risk Decisions. In Risk In The Technological Society . Routledge, 1982. ISBN 9780429304873

  80. [88]

    Philosophical Basis for Risk Analysis

    Chauncey Starr, Richard Rudman, and Chris Whipple. Philosophical Basis for Risk Analysis. Annual Review of Energy, 1(1):629–662, November 1976. ISSN 0362-1626. doi: 10.1146/annurev.eg.01.110176.003213. URL https://www.annualreviews.org/doi/10.1146/annurev. eg.01.110176.003213

  81. [89]

    How can we know a self-driving car is safe? Ethics and Information Technology, 23(4):635–647, December 2021

    Jack Stilgoe. How can we know a self-driving car is safe? Ethics and Information Technology, 23(4):635–647, December 2021. ISSN 1572-8439. doi: 10.1007/s10676-021-09602-1. URL https://doi.org/10.1007/s10676-021-09602-1

  82. [90]

    The Illusion of China’s AI Prowess

    Helen Toner, Jenny Xiao, and Jeffrey Ding. The Illusion of China’s AI Prowess. Foreign Affairs, June 2023. URL https://www. foreignaffairs.com/china/illusion-chinas-ai-prowess-regulation-helen-toner

  83. [91]

    Thunderforge Project: Integrating Commercial AI-Powered Decision-Making, 2025

    Defense Innovation Unit. Thunderforge Project: Integrating Commercial AI-Powered Decision-Making, 2025. URL https://www.diu. mil/latest/dius-thunderforge-project-to-integrate-commercial-ai-powered-decision-making

  84. [92]

    ‘Breeder’ Reactor Plan Facing Delays

    McElheny Victor K. ‘Breeder’ Reactor Plan Facing Delays. The New York Times, December 1973

  85. [93]

    Inside Task Force Lima’s exploration of 180-plus generative AI use cases for DOD

    Brandi Vincent. Inside Task Force Lima’s exploration of 180-plus generative AI use cases for DOD. DefenseScoop, November 2023. URL https://defensescoop.com/2023/11/06/inside-task-force-limas-exploration-of-180-plus-generative-ai-use-cases-for-dod/

  86. [94]

    The risks and inefficacies of AI systems in military targeting support

    Jimena Sofía Viveros Álvarez. The risks and inefficacies of AI systems in military targeting support. ICRC Humanitarian Law & Policy Blog, September 2024. URL https://blogs.icrc.org/law-and-policy/2024/09/04/the-risks-and-inefficacies-of-ai-systems-in-military- targeting-support/

  87. [95]

    Gaza: Israeli military’s digital tools risk civilian harm, 2024

    Human Rights Watch. Gaza: Israeli military’s digital tools risk civilian harm, 2024. URL https://www.hrw.org/news/2024/09/10/gaza- israeli-militarys-digital-tools-risk-civilian-harm

  88. [96]

    Ethical and social risks of harm from Language Models

    Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks...

  89. [97]

    The China Trap: U.S

    Jessica Chen Weiss. The China Trap: U.S. Foreign Policy and the Perilous Logic of Zero-Sum Competition. Foreign Affairs, August

  90. [98]

    DOGE’s Plans to Replace Humans With AI Are Already Under Way, March 2025

    Matteo Wong. DOGE’s Plans to Replace Humans With AI Are Already Under Way, March 2025. URL https://www.theatlantic.com/ technology/archive/2025/03/gsa-chat-doge-ai/681987/

  91. [99]

    Work and Greg Grant

    Robert O. Work and Greg Grant. Beating the Americans at their Own Game. An Offset Strategy with Chinese Characteristics. Center for a New American Security , 2019. ISSN 2510-2648, 2510-263X. doi: 10.1515/sirius-2019-4022

  92. [100]

    Persistent Pre-Training Poisoning of LLMs

    Yiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi, Eric Michael Smith, Nicholas Carlini, Florian Tramèr, and Daphne Ippolito. Persistent Pre-Training Poisoning of LLMs. arXiv, October 2024. doi: 10.48550/arXiv.2410.13722. URL http://arxiv.org/abs/2410.13722

  93. [101]

    Zico Kolter, and Matt Fredrikson

    Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv, December 2023. doi: 10.48550/arXiv.2307.15043. URL http://arxiv.org/abs/2307.15043

  94. [102]

    URL https://www.foreignaffairs.com/china/china-trap-us-foreign-policy-zero-sum-competition

  95. [1983]

    URL https://www.nrc.gov/docs/ML0717/ML071770230.pdf

  96. [2022]

    URL https://thebulletin.org/2022/07/why-policy-makers-should-beware-claims-of-new-arms-races/

  97. [2023]

    URL https://primer.ai/business-solutions/introducing-primer-delta-our-next-gen-ai-native-platform-to-transform-information- overload-into-decision-advantage/

  98. [2024]

    URL https://www.anduril.com/anduril-partners-with-openai-to-advance-u-s-artificial-intelligence-leadership-and-protect-u-s/

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.