{"id":"4535206a-8318-46c3-91fb-281bfff89bc2","arxiv_id":"2505.17084","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A qualitative framework that transfers more than 100 non-probabilistic risk-management strategies from engineering fields to securing and safely deploying LLM-powered systems.","lead":"The paper takes over 100 risk-management strategies developed for nuclear power, aviation, medicine, and other high-risk fields and maps them onto security and safety problems in LLM systems. It argues these qualitative methods are more practical than traditional probability-based risk analysis when AI risks are new, fast-moving, and shaped by adaptive adversaries.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The security-specific transfer claim rests on a hallucination pilot that did not address adversarial settings; the mapping in §3/Table A1 is unsupported by evidence.","rationale":"The reader's verdict is CONDITIONAL and the weakest assumption identified is exactly the transferability of the same-group hallucination pilot to adversarial LLM security. I agree with that assessment. The paper is a clearly written position/framework paper, and its conceptual proposal is plausible, but the evidence supporting its most novel security-oriented claim is entirely self-referential and non-adversarial. The Section 3 examples are illustrative, not validation, and the LLM-assisted workflow is asserted without reproducible details. A focused expert review on an adversarial security task would directly test whether RDOT adds value beyond existing security practice. Since my concern does not alter the reader's reasoning or verdict (CONDITIONAL is appropriate), I mark the verdict as UNCHANGED. I would emphasize that the burden is on the authors to supply this external validation before the security-transfer claim is treated as more than a hypothesis.","tokens_in":10974,"tokens_out":4197,"duration_ms":42558,"concrete_test":"Run the ESREL expert-review methodology from Shi and Gutfraind (2025) on a security-focused task: e.g., defending a RAG chatbot against deliberately crafted prompt-injection and jailbreak attacks. Recruit independent LLM security practitioners who are blinded to the origin of the strategy list; give them the full RDOT catalogue plus a matched control list of standard software-security practices (e.g., OWASP LLM Top 10 controls). Ask them to rate each strategy on applicability, likely effectiveness against adaptive adversaries, and whether it was already known/used in their security work. Compare the proportion of RDOT strategies rated 'highly promising and not previously considered' for the security task against the same proportion for the control list.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that 100+ non-probabilistic strategies from fields like nuclear engineering can mitigate risks in LLM-powered systems, 'including risks from adaptive adversaries.' The only direct evidence cited for the value of the RDOT catalogue is a same-group ESREL pilot (Shi and Gutfraind, 2025) that applied RDOT to reduce hallucinations in a RAG system, reporting 15% of strategies as highly promising and 31% as already considered. That task is a safety/accuracy problem, not a security problem involving adaptive adversaries. Section 2 explicitly acknowledges the pilot 'did not consider the potential of RDOT to improve security in adversarial settings.' Section 3 then hand-waves the transfer with examples (multi-layer defense, canary detection, red teaming) but provides no systematic mapping or empirical test of whether these strategies provide new, effective mitigations against prompt injection, jailbreaks, or data exfiltration. Thus the load-bearing premise—that the hallucination pilot generalizes to adversarial LLM security—is unsubstantiated. If that premise fails, the security-specific contribution of the paper is reduced to a list of generic engineering heuristics already known in software security.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that qualitative, non-probabilistic risk-management strategies drawn from engineering and other fields, collected in the RDOT catalogue of 100+ strategies, can be applied to improve the security and safety of LLM-powered systems, particularly against adaptive adversaries. It positions the argument against probabilistic risk analysis, describes five categories of strategies, gives illustrative examples from nuclear and other safety-critical domains, and presents an LLM-assisted workflow for selecting strategies. The main evidence offered for the catalogue's value is a prior same-group pilot that applied RDOT to reduce hallucinations in a RAG system, with 15% of strategies judged highly promising and 31% already considered.","tokens_in":11164,"tokens_out":4356,"duration_ms":42333,"significance":"If the transfer claim were established, the paper would provide a practically useful complement to probabilistic risk analysis for AI risk, especially in open-ended settings where event spaces cannot be enumerated. The paper's strength is its cross-disciplinary synthesis: the RDOT catalogue is a valuable resource, and the discussion of safety culture, FMECA, HAZOP, and similar methods is a useful reminder that qualitative risk management in high-stakes engineering deserves attention in AI security. However, the current manuscript does not validate the transfer to adversarial settings: the evidence is a hallucination pilot, and the promised mapping to LLM security is not actually presented. The paper is a position/conceptual piece, not an empirical study, and its contribution is accordingly limited to a plausible but unsubstantiated proposal.","major_comments":[{"comment":"The only quantitative evidence cited for RDOT is a same-group pilot on reducing hallucinations in a RAG system (Shi and Gutfraind 2025), with 15% of strategies judged highly promising and 31% already considered. Section 2 explicitly notes that this pilot 'did not consider the potential of RDOT to improve security in adversarial settings.' Section 3 then generalizes to 'LLM security' and the abstract claims coverage 'including risks from adaptive adversaries.' This is a load-bearing leap: the paper needs either evidence or a structured argument that hallucination-reduction findings transfer to security threats such as prompt injection, jailbreaks, or data exfiltration, or it must substantially weaken the claim to non-adversarial safety.","section":"§2 and §3"},{"comment":"The abstract states that the strategies 'are mapped to LLM security,' but no such mapping is presented. Table A1 is an abbreviated catalogue of strategies with definitions and categories; it does not map strategies to specific LLM security risks, threat models, or attack stages. Section 3's examples (multi-layer defense, canary detection, red teaming) are already standard in software security and are not a systematic mapping. The paper should either include the promised mapping (for example, a table listing each strategy and the applicable LLM threats or system components) or explicitly reframe the contribution as an illustrative selection rather than a complete mapping.","section":"§3 and Table A1"},{"comment":"The proposed LLM-powered workflow for matching strategies to projects is described with 'We have also had success' but no evaluation methodology, no criteria for success, and no comparison against a baseline or alternative selection process are provided. Since this workflow is listed as one of the paper's contributions, it needs at least a case study with a qualitative assessment, or the claim must be reduced to a suggestion for future work.","section":"§3.6"},{"comment":"The paper relies on the RDOT catalogue (Gutfraind 2023) as a comprehensive enumeration of risk-reducing strategies, but it does not summarize the catalogue's construction, inclusion criteria, or validation. Without this, it is difficult to assess whether the catalogue is complete or whether omissions could undermine the proposed transfer. Additionally, the five categories overlap (e.g., 'Adversarial' and 'harnessing' are combined in some rows), and some entries such as 'Wait and see' and 'Use the default action' may be trivial or unhelpful for LLM security. The paper should at minimum describe Gutfraind 2023's methodology and note any known limitations of the catalogue.","section":"§2 and Table A1"}],"minor_comments":[{"comment":"The paper repeatedly refers to '100+' strategies, but Table A1 contains fewer than 100 entries and is explicitly abbreviated; the full dataset is only available online. The reader cannot verify the count or the completeness from the manuscript, so a note describing the full dataset's size and composition would be helpful.","section":"Abstract and §3"},{"comment":"In the reference list, 'Anazon Web Services' should be 'Amazon Web Services.'","section":"References"},{"comment":"There are several typos and inconsistencies in the table: 'escaling' should be 'escalating'; one row uses 'format - workflow' where others use 'formal - workflow'; and category labels are inconsistent (e.g., 'adversarial' alone vs. 'adversarial; harnessing'). These should be corrected for readability.","section":"Table A1"},{"comment":"The sentence 'Another important class of techniques for risk reduction is technical standards and regulation' is grammatically incomplete and should be revised.","section":"§3.1"},{"comment":"The pointers 'see Table A1 in the Appendix' appear in sections that discuss applications, but Table A1 is only a catalogue, not a mapping to LLM security. These pointers are misleading and should be removed or replaced with a reference to the actual mapping, if one is added.","section":"§3.3 and §3.4"},{"comment":"The Limitations section is brief and does not explicitly acknowledge the lack of empirical validation of the transfer to adversarial security. Adding a sentence noting that the security claims are proposals supported only by anecdotal examples and an unrelated pilot would improve the paper's honesty.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is essentially a position paper, and its central claim is not empirically tested. The only quantitative evidence is a same-group pilot on hallucinations, which is self-referential and does not address adversarial settings. For a security-focused journal, I would expect at least a demonstration that the proposed strategies yield new or improved mitigations against concrete LLM attacks, or a clear reframing as a proposal with no evidence. The paper might fit better in a risk-analysis or AI-safety venue than in cs.CR. The editor may also wish to consider whether the promised mapping, which is central to the abstract, should be present in the manuscript itself rather than deferred to online material."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, clearly scoped position paper. The surprise if you skim the abstract is that it's not an empirical claim at all—the 100+ strategies are a transfer proposal, and the authors are upfront that the only pilot evidence comes from a hallucination study, not an adversarial security setting. So don't read it as a result; read it as a professional checklist with a plausible argument.\n\nWhat's genuinely new here is the mapping of Gutfraind's earlier RDOT catalogue onto adversarial LLM security. That catalogue, which draws on nuclear engineering, medicine, military science, and elsewhere, is a real artifact—Table A1 gives an abbreviated but readable sample, and the full version is on Zenodo. The examples (multi-layer defense, canary detection, red teaming, FMECA, checkrides) are well chosen and largely unfamiliar to mainstream software engineers, which gives the paper practical value as a provocation to think beyond the usual guardrails. The LLM-assisted strategy-matching workflow is a small but genuinely new addition; it's described at the level of a prompt and a judgment call by the engineer, which is fine for a position paper.\n\nThe soft spots are real but not fatal. The paper's one piece of quantitative evidence for RDOT's value—the Shi and Gutfraind hallucination pilot—reported 15% 'highly promising' and 31% 'already considered' strategies, and the paper itself notes it didn't consider adversarial settings. So the security-specific transfer is supported by argument and example, not data. That's a limitation, but it's named by the authors, and the argument doesn't hinge on the pilot's numbers. The LLM workflow anecdote is unreproducible, and the paper doesn't offer a systematic method for choosing among the 100+ strategies beyond three quick dimensions. Those are minor in a paper that's trying to open up a design space, not close it.\n\nI'd also push back on the stress-test note that the load-bearing premise fails. The premise is not that the hallucination pilot generalizes to security; it's that cross-field generic strategies are worth considering, and that's plausible on its face. The paper would be stronger with a threat-model-driven mapping or a small case study, but that's a revision request, not a rejection.\n\nAudience: practitioners and solution architects who want a broader risk-management menu, and researchers in AI risk who might expand the mapping. It deserves a serious referee; I'd send it out.","headline":"A useful, honest position paper that maps a 100+ engineering risk strategy catalogue into LLM security; the transfer is a proposal backed by examples, not by security-specific data, and the paper is appropriately modest about that.","tokens_in":11692,"tokens_out":2518,"would_cite":true,"duration_ms":18956,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that 100+ qualitative, non-probabilistic risk-management strategies developed in nuclear engineering, medicine, and other fields can be adapted to secure LLM-powered systems, especially where event spaces are open-ended…","keywords":["large language model security","non-probabilistic risk management","RDOT catalogue","adaptive adversaries","qualitative risk analysis","AI safety","retrieval-augmented generation","risk strategies transfer"],"falsifier":"Give the full RDOT catalogue to one set of LLM security engineers and a standard LLM security checklist to another, apply both to the same set of realistic attack scenarios, and count how many successful mitigations each group produces; the transfer claim would fail if the catalogue group finds no new effective strategies beyond the checklist, or if a review of documented LLM security incidents shows that most successful mitigations correspond to strategies already standard in the field rather than to novel RDOT entries.","tokens_in":10707,"feed_emoji":"🛡️","tokens_out":6035,"duration_ms":52865,"temperature":0.7,"pith_summary":"Large language models fail in ways that standard probabilistic risk analysis cannot easily track, because the space of possible failures is open-ended and adversaries adapt. This paper argues that a catalogue of 100+ non-probabilistic, field-agnostic risk-reducing strategies, developed in nuclear engineering, medicine, and other high-risk fields, can be carried over to LLM-powered systems. The strategies are grouped into five categories—structural, reactive, formal, counter-adversarial, and multi-stage—and mapped to concrete LLM security measures, including defenses against adaptive adversaries. A pilot study on retrieval-augmented generation hallucinations reported that 15% of the strategies were highly promising and previously unconsidered, which the paper takes as evidence the catalogue can also help in adversarial security settings. If the transfer holds, managers and engineers would gain a practical risk-management toolkit that works before enough incident data exists for probabilistic estimates.","feed_headline":"Nuclear-safety playbook offers 100+ ways to harden LLM systems","feed_subtitle":"Qualitative, non-probabilistic strategies from engineering and medicine are mapped to security against adaptive adversaries.","key_machinery":"The central object is RDOT, the Risk-reducing design and operation toolkit: a catalogue of 100+ generic strategies, each defined in one line and grouped into structural, reactive, formal, counter-adversarial, and multi-stage categories. The catalogue carries the argument by providing a common vocabulary for risks that resist enumeration, and the paper's Appendix A1 condenses it. The second mechanism is a matching workflow: a reasoning-capable LLM is given the strategy list plus a project description and asked to pair components with candidate strategies, then an engineer filters the suggestions for feasibility. A third element is the project triage criteria (impact, resources, adversarial relevance), which reduce the 100+ options to a few dozen.","core_discovery":"The paper's central claim is that qualitative risk management can be generalized across fields, so the Risk-reducing design and operation toolkit (RDOT) is a useful starting point for LLM security. It demonstrates this by reviewing each of five RDOT categories and showing familiar and unfamiliar examples: robust and fail-safe design, multi-layer defense, intrusion detection, FMECA and HAZOP-style hazard identification, game-theoretic and deception strategies against adaptive adversaries, and long-term steps like safety culture, sequential prototyping, and knowledge dissemination. The paper also gives selection criteria for projects (impact, resources, adversarial relevance) and an LLM-assisted workflow in which a reasoning model proposes candidate strategies from the catalogue. The evidence offered is a generalization from a single pilot in which 15% of strategies were judged highly promising and not previously considered and 31% were already considered, and the paper frames the result as a complement to, not replacement for, probabilistic risk analysis as data accumulates.","pith_inferences":["If the catalogue's transferability holds, an organization's safety culture may matter more than its AI tooling: many of the strategies are organizational practices such as training, checkrides, and independent audits, and teams in safety-critical industries already have the management processes to absorb them.","The same mapping exercise could be run for other responsible-AI dimensions, such as fairness or privacy, by pairing the RDOT catalogue with a domain-specific list of harms and a few real incidents; the paper does not do this.","A controlled comparison of incident-response teams, one using the full catalogue and one using standard LLM security checklists, would separate the value of the catalogue itself from the value of having any structured prompt; the paper reports only a single-project pilot.","The counter-adversarial strategies such as deception, randomization, and secrecy are described as protections against adversaries, but the same tools could plausibly be used defensively against LLM-powered attackers in ways the paper only gestures at."],"forward_implications":["LLM security teams can mine the full RDOT catalogue for mitigations that are uncommon in software engineering, such as FMECA, operator checkrides, and sacrificial subsystems, rather than relying only on known LLM defenses.","Architects can use the three triage dimensions to narrow the catalogue to a manageable set for a specific system, making the approach usable in ordinary product development.","The LLM-assisted matching workflow gives a repeatable way to surface unfamiliar strategies, though the raw output needs engineer vetting before adoption.","As incident data accumulates, teams should gradually shift from qualitative strategies to probabilistic methods, which can then refine or replace the non-probabilistic choices.","Because overlapping safeguards can introduce new failure modes, iterative prototyping and progressive refinement are expected to be part of any serious adoption of the toolkit."],"supporting_citations":[{"why":"Defines probabilistic risk assessment and its assumptions, the baseline the paper argues is impractical for open-ended AI risk.","marker":"(Kumamoto and Henley, 1996)"},{"why":"Introduces the small-world setting where probabilistic frameworks apply, providing the conceptual limit the paper uses to motivate non-probabilistic strategies.","marker":"(Savage, 1951)"},{"why":"Revisits Savage and supports the claim that probabilities are difficult to quantify in open-ended settings.","marker":"(Shafer, 1986)"},{"why":"Catalogued generic risk-reduction principles within engineering, supplying the cross-field strategy idea that RDOT extends.","marker":"(Todinov, 2007)"},{"why":"Provides specific structural strategies such as separation and segmentation that the paper maps to LLM system design.","marker":"(Todinov, 2015)"},{"why":"Publishes the RDOT catalogue, the central object whose completeness and validity the paper's argument depends on.","marker":"(Gutfraind, 2023)"},{"why":"Reports the pilot study on RAG hallucinations that gives the empirical evidence of 15% highly promising and 31% already considered strategies.","marker":"(Shi and Gutfraind, 2025)"},{"why":"Supplies the adversarial risk analysis perspective underlying the strategies designed for adaptive adversaries.","marker":"(Rios Insua et al., 2009)"},{"why":"Provides a system-theoretic accident model that the paper places in its formal-methods category for proactive risk discovery.","marker":"(Leveson, 2012)"}],"fun_headline_variants":["Nuclear-safety tactics offer 100+ LLM security strategies","Qualitative risk management from nuclear engineering for LLM security","From reactor safety to LLM security: a 100+ strategy playbook","Non-probabilistic safety strategies for LLM systems from nuclear engineering"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the RDOT catalogue is a complete and valid enumeration of generic risk-reducing strategies, and that the single-project pilot finding on retrieval-augmented generation hallucinations transfers to adversarial LLM security settings.","fun_headline_variants_meta":{"raw":{"variants":["Nuclear-safety tactics offer 100+ LLM security strategies","Qualitative risk management from nuclear engineering for LLM security","From reactor safety to LLM security: a 100+ strategy playbook","Non-probabilistic safety strategies for LLM systems from nuclear engineering"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1750,"prompt_tokens":922,"completion_tokens":828,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":753}},"tokens_in":538,"tokens_out":828,"duration_ms":12135,"temperature":1.0,"reasoning_tokens":753,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:32:22.084499+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the full RDOT catalogue to one set of LLM security engineers and a standard LLM security checklist to another, apply both to the same set of realistic attack scenarios, and count how many successful mitigations each group produces; the transfer claim would fail if the catalogue group finds no new effective strategies beyond the checklist, or if a review of documented LLM security incidents shows that most successful mitigations correspond to strategies already standard in the field rather than to novel RDOT entries.","supporting_citations":[],"review_version":1}