{"id":"b0a68103-1cdc-47bd-a2af-cad4d0e04aad","arxiv_id":"2504.15088","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI technologists are co-opting safety engineering terminology to weaken risk thresholds for military AI, a move the authors argue will backfire on US national security.","lead":"This paper argues that AI companies and 'AI safety' groups have quietly redefined safety engineering terms, replacing proven risk thresholds with vague 'capabilities' metrics, which could weaken US military AI oversight. It urges democratic institutions to set AI risk tolerances before military deployment accelerates.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central charge depends on applying Cold War-era risk thresholds to foundation-model military systems; if AI failure modes cannot populate MIL-STD-882/SIL probability categories, the 'lowered threshold' claim conflates weakened standards with no applicable standard.","rationale":"The reader's weakest-assumption analysis identified the transferability of established safety-engineering risk thresholds to AI-based military systems, and this stress-test concurs that this is the most load-bearing point. The paper's central claim is not merely that AI systems are unsafe in military contexts—that is well documented—but that AI technologists have actively 'substituted' and 'weakened' pre-existing risk thresholds. For that substitution claim to be meaningful, the pre-existing thresholds must be applicable to the systems in question. If, for example, a foundation model's probability of a hazardous failure cannot be estimated in the operational environment because its behavior is distribution-dependent and deployment-context-sensitive, then MIL-STD-882E's probability categories and SIL's probability-of-failure-on-demand levels cannot be populated. In that case, the situation is better described as the absence of an operationalized threshold rather than the lowering of one. The paper itself supplies evidence that AI failure modes are hard to characterize in this way (Section 4's discussion of nondeterminism and 'intractable' risk), which creates a tension with its assertion in Section 2 that AI risk categories 'fall firmly within the scope of traditional risk frameworks.' The paper does not resolve this tension. This does not vitiate the paper's policy recommendations or its documentation of specific harms; it does mean that the specific accusation of 'safety revisionism' needs an additional argument showing how the numerical thresholds from nuclear and defense practice can be instantiated for LLM-based systems. The proposed concrete test—actually populating a MIL-STD-882E matrix for one of the paper's own examples—would settle whether the transferability assumption holds. The reader's CONDITIONAL verdict already accounts for this uncertainty, so no change to the verdict is recommended.","tokens_in":19834,"tokens_out":5062,"duration_ms":51704,"concrete_test":"Construct a worked example using a military use case from the paper: LLM-assisted target identification via intercepted Arabic communications (refs [9], [61], [62]). Using the published accuracy and failure data cited, attempt to populate the MIL-STD-882E hazard risk assessment matrix (severity categories I–IV by probability levels A–E) for this system. If the probability level cannot be assigned because no stable failure rate exists for the model in that operational context—or if the resulting risk index falls in an 'acceptable' range despite the documented failure modes—then the premise that traditional thresholds apply to foundation models is unsupported, and the 'lowered threshold' claim must be re-specified as 'no threshold was operationalized.' If, instead, the matrix can be populated and shows unacceptable risk, the premise survives and the paper's argument is strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central accusation—that AI technologists are engaged in 'safety revisionism' that weakens established thresholds—rests on the premise that quantitative risk thresholds inherited from nuclear and defense safety engineering (Starr's 10^-4 deaths per person per year, MIL-STD-882E severity/probability categories, SIL levels) can be meaningfully applied to foundation-model-based military systems. The paper asserts this in Section 2 ('their risk categories fall firmly within the scope of traditional risk frameworks') and repeats 'foundation models are no exception' in the abstract, but it never demonstrates that the statistical assumptions of those thresholds are satisfiable for AI. MIL-STD-882E and SIL frameworks require hazards to be enumerated and assigned a probability of occurrence in the intended operating context; current LLM failure modes are context-dependent, nondeterministic, and in many documented cases (e.g., the Arabic mistranslation and civilian-casualty estimation failures cited in Section 4) only manifest in deployment, making 'probability of hazardous event' ill-defined at assurance time. If probabilities cannot be assigned, then one cannot establish that a threshold has been lowered—one can only say that no applicable threshold was operationalized. The paper conflates these two distinct situations, and the normative charge of 'revisionism' loses its force unless the transferability of the thresholds is shown, not just asserted. The weakness is not that the documented safety failures are irrelevant; it is that the argumentative bridge from 'these systems are unsafe' to 'the established thresholds were intentionally weakened' is unbuilt.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is an argumentative policy analysis arguing that, because no societally deliberated risk thresholds have been set for AI, AI technologists—mainly industry labs and self-styled 'AI safety' organizations—have been able to redefine safety terminology and substitute 'capabilities' and 'alignment' evaluations for traditional safety-engineering assurance. The authors reconstruct the history of risk thresholds from Chauncey Starr's nuclear-era work through MIL-STD-882 and Safety Integrity Levels, then argue that current AI evaluation frameworks (Google DeepMind's Frontier Safety Framework, Anthropic's Responsible Scaling Policy, Task Force Lima, the UK AI Security Institute) hollow out the meaning of 'safety', 'safety case', and 'red-teaming'. They use the Gaza deployments and supply-chain vulnerabilities to argue that the resulting absence of enforceable thresholds enables deployment of foundation models in military contexts at lowered safety and security thresholds, ultimately undermining US national security and international humanitarian law. The paper concludes by calling for democratic deliberation to set AI risk thresholds and for preserving traditional TEVV standards in military AI evaluation.","tokens_in":20063,"tokens_out":8376,"duration_ms":80884,"significance":"If the argument holds, the paper makes a valuable contribution by connecting AI safety discourse to the history and vocabulary of safety engineering and by identifying a concrete mechanism—terminological revisionism—through which existing assurance standards can be bypassed. Its strengths are the historically grounded account of risk thresholds, the concrete documentation of military AI deployment failures, and a clear normative proposal: preserve established assurance frameworks and set AI risk thresholds through democratic deliberation rather than industry self-assessment. The paper is not an empirical study and does not ship machine-checked proofs or reproducible code; its contribution is conceptual and evidentiary. The central claim is defensible but requires an important qualification: the transferability of quantitative risk thresholds to foundation models is asserted rather than demonstrated, and the national-security conclusion is hedged in some places but stated categorically in others.","major_comments":[{"comment":"The paper's central charge of 'lowered' risk thresholds presupposes that established quantitative thresholds (Starr's 10^-4 deaths per person per year, MIL-STD-882E categories, SIL levels) can be populated for foundation models. The paper asserts this ('their risk categories fall firmly within the scope of traditional risk frameworks') but never shows how a probability of a hazardous event can be assigned to a context-dependent, nondeterministic model at assurance time. Without such a demonstration, the 'lowered threshold' claim is not established; the situation may instead be that no applicable threshold has been operationalized. This distinction is load-bearing for the 'revisionism' charge, because mislabeling an absence of a standard as a weakening of a standard changes the normative force of the argument. Please either provide a worked example for a concrete use case (e.g., target nomination or intelligence analysis) or reframe the argument as 'no applicable risk threshold has been operationalized'.","section":"Section 2 / Table 1"},{"comment":"The claim that AI systems are 'deployed at levels far below the risk thresholds that would be deemed acceptable through standardized safety processes' is not supported by a quantitative comparison. The cited evidence—Arabic mistranslation, cell-tower-based civilian casualty estimates, and evasion failures—demonstrates qualitative failures, but the paper provides no probability-of-failure estimates, exposure analysis, or comparison against any specific threshold. As stated, the claim is not falsifiable. It should either be quantified for a specific deployment scenario or explicitly presented as a qualitative judgment about the absence of demonstrated compliance.","section":"Section 4, final paragraph"},{"comment":"The title's 'self-fulfilling prophecy' and the national-security conclusion are causal claims that the evidence does not fully support. The paper shows that some military AI uses have failed and that some frameworks use nonstandard terminology, but it does not systematically consider alternative explanations for the observed outcomes, such as bureaucratic incentives, technical immaturity, or good-faith disagreements about how to adapt assurance methods to AI. The hedging in Section 1 ('may be precisely what disadvantages') is honest, but the categorical conclusion in Section 5 ('will imperil' and 'will result in a significant civilian death toll') goes beyond the evidence presented. Please separate the well-supported claim that current frameworks do not satisfy traditional assurance standards from the more speculative claim that this trajectory will compromise US national security.","section":"Sections 1 and 5"}],"minor_comments":[{"comment":"The paper defines risk tolerance and risk threshold as distinct concepts but later uses them interchangeably (e.g., Section 3: 'risk thresholds are derived from risk tolerances'). Please maintain the distinction consistently.","section":"Section 2"},{"comment":"The statement that a safety case 'is not intended to produce any 'targets' or thresholds' is too categorical; some safety-case methodologies incorporate quantitative safety requirements as part of the argument. The substantive criticism of Google DeepMind's framework does not depend on this categorical claim, so it can be softened without harming the argument.","section":"Section 3"},{"comment":"The characterization of Task Force Lima as taking a 'capabilities evaluation' approach relies on the task force's executive summary [55]. Please verify the citation and, if possible, quote the relevant language, since the executive summary may not be publicly accessible.","section":"Section 3.1"},{"comment":"The claim that foundation models have 'poor accuracy with non-English languages, especially for Arabic' is supported by general LLM cultural-bias studies [61, 62]; the connection to the specific IDF systems described in [9] should be made more explicit, since those systems may include non-LLM components.","section":"Section 4"},{"comment":"The phrase 'existential risks that are very real and present' sits awkwardly with the paper's earlier rejection of 'speculative existential risks' in Sections 1 and 2. Please clarify whether this refers to concrete military harms rather than the speculative AI-existential-risk literature.","section":"Section 5"},{"comment":"Please standardize spelling and formatting inconsistencies, including 'DOD' vs 'DoD' and 'LAWs' vs 'LAWS', and check the column headings in Table 1, which appear truncated.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a timely and provocative policy argument that fits the journal's scope if the authors are willing to substantially qualify the threshold-transferability claim. I would not reject it: the conceptual critique of 'AI safety' terminology is valuable and well-sourced. The main revision burden is to either operationalize traditional risk thresholds for a concrete AI use case or reframe the argument around the absence of operationalized thresholds. The authors should also distinguish more carefully between industry actors, government evaluators, and 'AI safety' nonprofits, since the paper sometimes lumps them together as 'AI technologists'. I did not find evidence of circularity or parameter fitting; this is not a technical derivation, so the standard for evidence should be appropriate to an argumentative policy paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is the 'safety revisionism' framing: AI labs and 'AI safety' organizations have redefined safety, safety cases, and red-teaming in ways that sever them from established assurance practice, and this is happening in military contexts with little democratic scrutiny. That framing is genuinely new and the paper supports it well, drawing on Chauncey Starr's risk-analysis history and on concrete contemporary examples like Task Force Lima, the UK AI Security Institute, and IDF targeting in Gaza. The authors know the safety-engineering literature and they connect it to AI governance more carefully than most work in this area.\n\nThe main soft spot is the transferability claim. The abstract insists foundation models are 'no exception' to MIL-STD-882 and SIL-style thresholds, but the paper never shows that the probability categories in those frameworks can actually be populated for a nondeterministic, context-dependent LLM. If a hazard's probability of occurrence cannot be estimated at assurance time, then the situation is not 'threshold lowered' but 'no applicable threshold defined.' That conflation weakens the normative force of the accusation. It is a real gap, but not a fatal one: the deeper argument—that safety language is being laundered into a veneer of compliance—stands even if the numerical threshold question remains unresolved.\n\nThe causal claim about compromised national security is speculative and hedged ('may be precisely what disadvantages...'), and the paper does not seriously engage the counterargument that AI decision support is not a weapon in the traditional sense and therefore requires different assurance approaches. The evidence is selected case studies, not systematic analysis. Still, this is a serious piece of policy analysis, not a throwaway preprint. It deserves a careful referee who will push the authors to defend the applicability of legacy thresholds, or to reframe their critique as an absence-of-standards problem rather than a lowering-of-thresholds problem. I would send this to a policy venue or a standards body for discussion, and I would not cite it in my own technical work, but I would flag it for people working on AI governance and military evaluation.","headline":"A pointed, well-sourced policy argument about 'safety revisionism' in AI defense work, though it overreaches by asserting without proof that traditional risk thresholds apply to foundation models.","tokens_in":20613,"tokens_out":3281,"would_cite":false,"duration_ms":32927,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that AI technologists are engaging in 'safety revisionism' — redefining safety-engineering terms so that foundation models can enter military use without meeting established risk thresholds — and that this will weaken US…","keywords":["safety revisionism","AI risk thresholds","military AI","foundation models","safety cases","capabilities evaluation","test evaluation verification validation","national security"],"falsifier":"The central claim would fail if a military AI system that passed a current capabilities-evaluation framework was then assessed under a standard military system-safety analysis and met the same quantified risk thresholds required of the conventional system it replaces, such as a probability of a fatal hazardous event per deployment at or below the legacy-system threshold. It would also be weakened by a documented case in which an AI 'red-teaming' exercise uncovered a previously unknown hazard and measurably reduced estimated risk, showing that the redefined practice can carry real assurance weight.","tokens_in":19621,"feed_emoji":"⚠️","tokens_out":9861,"duration_ms":85069,"temperature":0.7,"pith_summary":"This paper argues that because no democratic body has set agreed risk thresholds for AI systems, the people who build and sell them have ended up setting their own tolerances for acceptable harm — and they have set them low in order to speed foundation models into military use. The authors call this process 'safety revisionism': familiar terms from safety engineering, such as 'safety,' 'safety cases,' and 'red-teaming,' are redefined in ways that strip them of their original assurance power. The paper connects this redefinition to the historical risk thresholds developed for nuclear systems and defense standards, and claims that AI evaluations based on 'capabilities' rather than quantified risk cannot show that a system is safe enough for use in targeting or decision support. If the argument is right, the accelerated military adoption of AI will expose personnel and civilians to documented failure modes, create vulnerabilities adversaries can exploit, and in the end weaken rather than strengthen US national security.","feed_headline":"AI labs are lowering the safety bar for military systems","feed_subtitle":"With no agreed risk thresholds, technologists set the bar themselves — and it is dropping.","key_machinery":"The load-bearing mechanism is the risk threshold — a quantified level of risk exposure above which action must be taken — together with the linguistic drift that detaches it from assurance practice. On one side stands the historical framework in which thresholds are fixed by societal deliberation and then substantiated by safety cases: structured arguments, backed by evidence, that a system is acceptably safe for a defined application in a defined environment. On the other side stand the contemporary 'frontier safety frameworks' that convert safety cases into capability targets and substitute red-teaming for boundary testing, with no threshold to test against. The comparison between those two sides is the gear that turns missing AI governance into lowered real-world safety for defense and, through precedent, for civilian critical infrastructure.","core_discovery":"The core discovery is a mechanism with a name: 'safety revisionism.' When societally accepted risk tolerances for AI do not exist, technologists become the de facto arbiters of how much harm is acceptable, and they exercise that power by quietly replacing the vocabulary of safety engineering. Risk thresholds that were once set through democratic deliberation — such as the quantified fatal-risk criteria derived for nuclear power or the Safety Integrity Levels used in defense system safety — are superseded by 'frontier safety frameworks' that state unverifiable goals about model intent, such as not causing 'catastrophe,' without any numeric threshold a system can be tested against. The paper shows that the methodologies meant to substantiate safety, particularly safety cases, have been redefined into 'affirmative cases' and red-teaming exercises that cannot demonstrate risk reduction. The conclusion is that this trajectory is a self-fulfilling prophecy: the weakened thresholds adopted to preserve a supposed AI advantage are exactly what will compromise the safety and security of the military systems that depend on them.","pith_inferences":["Going beyond the paper: its logic yields a testable prediction — among militaries that field AI-enabled targeting and decision support, those with weaker assurance thresholds will show more civilian casualties and friendly-fire incidents, then loosen thresholds further to keep pace.","A practical audit rule follows: any AI safety framework that cannot state a numeric threshold (for example, a maximum probability of a fatal hazardous event per deployment) is not a safety framework in the engineering sense, whatever its name; applying that test to current policy documents would give governance bodies a concrete checklist.","The same capabilities-evaluation logic is already visible outside defense, in proposals to run government administration with foundation models, so the missing-threshold problem is broader than the military case the paper examines."],"forward_implications":["Military evaluation frameworks built on 'capabilities evaluations' cannot demonstrate that a foundation model is acceptably safe, because they never commit to a quantified risk threshold.","Foundation models in targeting and decision-support roles will carry documented failure modes — poor performance on non-English data, supply-chain attack surfaces, brittle safeguards — into settings where those failures can kill civilians and violate international humanitarian law.","If defense adoption sets the precedent, civilian safety-critical uses of AI will be judged against the same lowered bar, since military and civilian assurance standards influence each other.","Democratic bodies, not technologists, need to set explicit AI risk tolerances, including the number of societally accepted fatalities a given deployment assumes.","The national-security justification for accelerated adoption is self-defeating: a brittle AI system is an asset adversaries can exploit, so weakening thresholds to win an AI race undermines the very advantage the race is meant to secure."],"supporting_citations":[{"why":"Supplies the foundational definition of risk tolerance versus risk threshold, and the distinction between societal and individual risk that underpins the democratic-versus-autocratic critique.","marker":"[88]"},{"why":"Provides the historical model of a quantified risk threshold for a new technology, including an upper bound of roughly one death per ten thousand people per year for nuclear power.","marker":"[86]"},{"why":"Defines the military system-safety standard and the risk-acceptance categories the paper insists AI-based defense systems must still satisfy.","marker":"[67]"},{"why":"Is the contemporary frontier-safety framework that redefines 'safety cases' as capability targets, the paper's main example of revisionism.","marker":"[28]"},{"why":"Supplies the 'affirmative cases' language that the paper treats as an unfalsifiable replacement for traditional safety cases.","marker":"[6]"},{"why":"Documents the capabilities-evaluation approach for defense foundation models that the paper argues breaks with test, evaluation, verification, and validation norms.","marker":"[55]"},{"why":"Establishes the context-specific definition of safety and the argument that 'general' safety evaluation is intractable, which underlies the paper's rejection of broad benchmarking.","marker":"[51]"},{"why":"Grounds the paper's claim that assurance of AI systems must be application-specific and tied to dependability properties, not general capability claims.","marker":"[10]"}],"fun_headline_variants":["Safety revisionism: AI labs quietly lower military risk thresholds","Without agreed AI risk limits, techies set the bar — and it's dropping","How 'safety revisionism' endangers US national security and defense AI","The self-fulfilling prophecy of weakened AI risk thresholds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the risk thresholds developed for earlier safety-critical technologies — nuclear power plants, military hardware — apply to AI-based military systems and that foundation models are 'no exception'; if AI failure modes are so different that those thresholds cannot meaningfully be transferred, the charge that AI firms are 'revising' safety loses its force.","fun_headline_variants_meta":{"raw":{"variants":["Safety revisionism: AI labs quietly lower military risk thresholds","Without agreed AI risk limits, techies set the bar — and it's dropping","How 'safety revisionism' endangers US national security and defense AI","The self-fulfilling prophecy of weakened AI risk thresholds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1599,"prompt_tokens":1057,"completion_tokens":542,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":673,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":673,"tokens_out":542,"duration_ms":4971,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:32:53.948272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The central claim would fail if a military AI system that passed a current capabilities-evaluation framework was then assessed under a standard military system-safety analysis and met the same quantified risk thresholds required of the conventional system it replaces, such as a probability of a fatal hazardous event per deployment at or below the legacy-system threshold. It would also be weakened by a documented case in which an AI 'red-teaming' exercise uncovered a previously unknown hazard and measurably reduced estimated risk, showing that the redefined practice can carry real assurance weight.","supporting_citations":[{"cited_title":"Philosophical Basis for Risk Analysis","cited_arxiv_id":null,"evidence_quote":"Supplies the foundational definition of risk tolerance versus risk threshold, and the distinction between societal and individual risk that underpins the democratic-versus-autocratic critique."},{"cited_title":"Department of Defense","cited_arxiv_id":null,"evidence_quote":"Defines the military system-safety standard and the risk-acceptance categories the paper insists AI-based defense systems must still satisfy."},{"cited_title":"Task Force Lima Executive Summary","cited_arxiv_id":null,"evidence_quote":"Documents the capabilities-evaluation approach for defense foundation models that the paper argues breaks with test, evaluation, verification, and validation norms."},{"cited_title":"Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems","cited_arxiv_id":null,"evidence_quote":"Establishes the context-specific definition of safety and the argument that 'general' safety evaluation is intractable, which underlies the paper's rejection of broad benchmarking."}],"review_version":1}