{"id":"54f571df-8573-4da0-a70d-cc62cb805573","arxiv_id":"2502.03467","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper mapping traditional critical systems engineering to AI safety frameworks and advocating Assurance 2.0-style cases, broader boundaries, and explicit risk tolerability.","lead":"This paper argues that AI safety frameworks should borrow rigor from aviation and nuclear safety engineering, including broader system boundaries and formal assurance cases. It draws on decades of critical systems experience to help prevent AI assurance from addressing the wrong risks, using weak techniques, or failing to communicate its claims.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The no-accidents-due-to-Step-7 premise in Section 2 is an unsourced, potentially definitional claim; it is the empirical anchor for the transferability argument and should be tested directly.","rationale":"Agreement with the reader is partial. The reader correctly locates the soft spot in the transferability of critical-systems experience to frontier AI. I sharpen the concern: the most load-bearing part of that transfer is not the abstract analogy but the specific empirical claim in Section 2 that verification failures have never caused a modern aircraft accident. That claim is doing real work, because it supports the paper's ranking of hazards: requirements and hazard analysis over verification, and hence the critique of red-teaming and the call for theories that allow behaviours to be assured with adequate confidence. The paper presents this as extensive historical experience and much data, but provides no dataset, no citation, and no definition of modern aircraft or failures of Step 7. I am not arguing the claim is false; I am arguing it is unsecured in a way that matters. If a structured review confirms the claim, the current ACCEPT verdict stands. If the review cannot confirm it, or shows that the Step 5/Step 7 split is an artifact of causal attribution, the paper should be accepted only conditionally, with the Section 2 base-rate claim either evidenced or explicitly downgraded to the authors' professional experience. The paper deserves credit for being honest about the difficulty of transferring these methods, for including the Assurance 2.0 framework transparently in an appendix, and for prescribing coherent, actionable topics such as broader system boundaries, explicit risk tolerability, guards, and defence in depth. The concern raised here does not warrant rejection; it warrants one specific evidentiary check before the central analogy is used as a strong empirical foundation.","tokens_in":17155,"tokens_out":6769,"duration_ms":72143,"concrete_test":"Compile a structured root-cause review of software-related commercial aviation accidents and incidents since roughly 1980 using NTSB, FAA, EASA, AAIB, and BEA reports plus peer-reviewed investigations, including cases such as 737 MAX MCAS, QF72, AF447, and the A400M crash. For each event, independently classify the dominant causal step using the paper's 8-step definitions, with two coders blind to the authors' claim. If any type-certified accident has an implementation-specification mismatch as a necessary cause, the no-Step-7 premise is falsified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 2 anchors the paper's lesson on the empirical assertion that there have been no accidents of modern aircraft due to failures of Step 7 (Verification), and that all such failures are attributable to Step 5, Requirements Validation. This base rate is load-bearing: it justifies the claim that red-teaming and testing are inadequate for AI and that hazard analysis, risk tolerability, and theory-backed assurance cases are the real priorities. The assertion is stated without citation and has two structural problems. First, the Step 5/Step 7 boundary is partly definitional: if an accident is caused by a dangerous behavior that was not captured in the safety requirements, it is classified as Step 5 regardless of whether a stronger verification activity, such as integrated system testing, flight testing, or analysis of the learned behavior in context, would have caught it in service. Second, aviation accident reports are not written in the 8-step taxonomy, so no-Step-7-failures is an interpretive coding of selected accidents, not a measured frequency. The modern software-intensive fleet may also differ from the historical base from which the lesson is drawn. The paper's recommendations may still be right, but this empirical anchor should not be accepted on authority.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper distills the authors' experience in critical-systems assurance into an eight-step safety-engineering process, claims that traditional aviation accidents are attributable to requirements validation (Step 5) rather than verification failures (Step 7), and uses this base rate to argue that AI Safety Frameworks are likely to go wrong by addressing the wrong risks, relying on techniques inadequate for the risks they address, and communicating claims poorly. The paper recommends broadening system boundaries, making risk tolerability explicit, using design basis events and threats, adopting architectural guards and defence in depth, and building assurance cases based on the authors' Assurance 2.0 framework. It maps these recommendations to two questions from the FAISC call and closes with suggestions for how AI safety frameworks should evolve.","tokens_in":17374,"tokens_out":11018,"duration_ms":107945,"significance":"If the paper's lessons are accepted, its main contribution is to redirect AI-safety attention from verification-centric activities such as red-teaming and testing toward hazard analysis, requirements validation, and explicit risk-tolerability reasoning, and to give concrete architectural and assurance-case vocabulary for doing so. The paper is transparent about being selective and experience-based, and it offers several testable empirical claims (the aviation accident base rate; the absence of these topics in current frameworks), which is a strength because it allows readers to check the evidence behind the recommendations. Its significance is limited by the qualitative, analogy-based nature of the argument and by the fact that the positive recommendation (Assurance 2.0) is the authors' own framework, presented without independent evaluation; nevertheless, as an initial position paper it is a useful and readable contribution to the FAISC dialogue.","major_comments":[{"comment":"The assertion that 'there have been no accidents of modern aircraft due to failures of Step 7 (Verification)' and that 'all modern aircraft failures have been attributed to Step 5' is load-bearing: it anchors the later argument that testing and red-teaming are not the bottleneck for AI assurance, and that hazard analysis and requirements validation should be prioritized. The assertion is given without citation and is partly definitional. Because the Step 5/Step 7 boundary is drawn in terms of whether the safety requirements captured the hazard, an accident caused by a dangerous behavior that a more contextual verification activity (e.g., integrated system testing, flight testing, or in-service analysis of learned behavior) would have caught will still be coded as Step 5 if the requirement was incomplete. The paper should either cite systematic accident-taxonomy studies supporting the base rate, or weaken the claim into an explicitly coded statement and discuss how classification may differ for machine-learning systems. The 737 MAX example supports the need for better hazard analysis, but it does not by itself establish the universal 'no Step 7 accidents' prior.","section":"Section 2, paragraph beginning 'There is extensive historical experience.'"},{"comment":"The universal negative that 'none of the corporate or national frameworks that we have examined make any mention of these topics' is central to the paper's critique of current AI Safety Frameworks, but the examined frameworks are not identified, the review period is not given, and the criterion for 'mention' is not defined. A reader cannot verify whether this is a measured absence or a selective sample. Please provide a list (or at least a table) of the frameworks reviewed with dates and the specific topics searched for, or replace the universal claim with a bounded statement such as 'in the frameworks we reviewed, these topics were absent or barely developed.'","section":"Section 3, paragraph beginning 'None of the corporate or national frameworks.'"},{"comment":"The paper recommends Design Basis Events and Design Basis Threats as a way to bound the open-ended risk space for AI, while acknowledging two paragraphs later that 'for AI applications it may be difficult to enumerate a set with adequate coverage.' This is a real tension in the transfer argument: design-basis reasoning is only useful if a justifiable worst-case set can be constructed, and the paper does not say how such a set would be built or validated for frontier AI. Since this recommendation underlies the call for broader system boundaries and risk tolerability analysis, the authors should either provide a worked example of a design basis for one AI application, or state more explicitly that the concept is being offered as an open question rather than a ready-made solution.","section":"Sections 3.2.1 and 4"}],"minor_comments":[{"comment":"'e.q.' should be 'e.g.' and 'Requirements V alidation' contains a stray space.","section":"Section 2, items 3 and 5"},{"comment":"The phrase 'All modern aircraft failures' is ambiguous; please specify whether it means accidents, fatal accidents, or all system failures, and clarify the date range or aircraft generation covered by 'modern.'","section":"Section 2"},{"comment":"These two entries appear to be duplicates of the same IAEA report; please consolidate them.","section":"References [20] and [22]"},{"comment":"There is a typo: 'Stationaery Office' should be 'Stationery Office.'","section":"Reference [19]"},{"comment":"The statement that 'the only things certified by the FAA are airplanes and engines (and propellers)' is an oversimplification of FAA certification (which also includes type designs and Technical Standard Order authorizations); consider rewording to avoid an unnecessary quibble.","section":"Section 3.1"},{"comment":"'CrowdStrike crash' should be 'CrowdStrike outage' or 'CrowdStrike incident,' since the event was not a crash in the technical sense.","section":"Section 3.1.1"},{"comment":"The term 'indefeasible assurance' is used before it is defined; please define it at first use or add a pointer to the Appendix.","section":"Section 2, Step 8 and Section 3.4.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a workshop-style position paper with a strong authorial voice. It relies heavily on the authors' own Assurance 2.0 framework and cites it extensively, which is transparent but means the positive recommendation is not independently evaluated; this should be weighed in editorial decision. The recommendation of major_revision is based on two unsupported empirical anchors (the aviation accident base rate in Section 2 and the universal negative about frameworks in Section 3); if the authors are willing to support or carefully bound those claims, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: two veterans of safety-critical systems assurance map the traditional 8-step engineering process onto AI Safety Frameworks and warn that system boundaries are drawn too narrow, risk tolerability is under-elaborated, and assurance methods lack theory. It is a position piece, not a technical result, but it is honest, well-referenced, and worth reading before drafting or critiquing an AI safety framework. The genuinely new bit is the AFGI concept (Artificial Fairly General Intelligence) and the detailed mapping to the FAISC questions. The promotion of the authors' own Assurance 2.0 is consistent and transparent, and the appendix gives a compact summary.\n\nWhat the paper does well: the 8-step summary is accurate and accessible; the critique of red-teaming as 'the fox guarding the henhouse' is pointed but fair; and the insistence on hazard analysis, requirements validation, and explicit tolerability judgments targets exactly what policy documents tend to skip. The paper is also refreshingly clear that frontier AI is a component in a system, not a system itself.\n\nThe soft spots are real but not fatal. The load-bearing empirical claim in Section 2 — that there have been no accidents of modern aircraft due to failures of Step 7 (Verification) — is stated without citation. That is a concern because the Step 5/Step 7 boundary is partly definitional: an accident caused by a dangerous behavior not captured in the safety requirements gets classified as Step 5, even when stronger verification in context might have caught it. I would treat that sentence as an experienced engineer's assertion, not a measured statistic. The transferability of 10^-9-style failure rates and assurance cases to frontier AI is argued by analogy; the authors acknowledge this, but the acknowledgment does not close the gap. Also, the paper leans heavily on the authors' own Assurance 2.0 apparatus; that is not a flaw in itself, but readers outside that community should know the terminology comes with a specific agenda.\n\nWho is this for? People drafting AI safety frameworks, and researchers who want a check on whether current assurance ideas match decades of critical-systems practice. It deserves a serious referee: the position is defensible, the experience is real, and the authors flag their own limitations. My verdict would be accept with minor revision: add a caveat or citation for the Step 7 claim, and soften a couple of the unsourced generalisations.","headline":"Worth engaging: an experienced, honest position piece whose recommendations survive its one unsourced empirical anchor.","tokens_in":17869,"tokens_out":1596,"would_cite":true,"duration_ms":16551,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI assurance will go wrong in three predictable ways unless it adopts the discipline of critical systems engineering.","keywords":["AI assurance","safety frameworks","assurance cases","critical systems engineering","risk tolerability","hazard analysis","Assurance 2.0","foundation models"],"falsifier":"A documented modern aircraft accident caused by a failure in Step 7 (verification), rather than by requirements validation or earlier steps, would falsify the paper's historical premise; alternatively, a demonstration that a frontier AI system's risk cannot be bounded by any design-basis event because the environment is intrinsically open-ended would test its central transferability claim.","tokens_in":16963,"feed_emoji":"🛡️","tokens_out":4747,"duration_ms":43287,"temperature":0.7,"pith_summary":"This paper argues that current AI safety frameworks are heading for three avoidable failures: they target the wrong risks, use assurance techniques too weak for the risks they do address, and communicate their confidence poorly. Drawing on how aircraft and nuclear systems achieve safety, it contends that AI assurance must treat safety as a property of the whole socio-technical system, not of the model, and must set explicit risk-tolerability targets before choosing methods. It advocates structured assurance cases that make arguments deductively tight, expose doubts as defeaters, and weigh evidence by how much it actually raises confidence. The payoff, if the paper is right, is a concrete checklist for making frontier AI safety frameworks more demanding and more honest about what they know.","feed_headline":"AI assurance will fail without aircraft-grade safety discipline","feed_subtitle":"Critical-systems engineers say frameworks must widen boundaries, set risk tolerability, and argue sceptically.","key_machinery":"The load-bearing object is the structured assurance case in the 'Assurance 2.0' style: a tree of claims, argument, and evidence in which every argument step is expected to be deductive, every doubt is recorded as a defeater that must be refuted or accepted as residual risk, and evidence is scored by how much it increases confidence in the useful claim rather than the measured claim. Alongside it stands the eight-step safety-engineering process (environment, requirements, hazard analysis, safety requirements, validation, specification, verification, assurance case), which supplies the vocabulary for deciding what counts as a relevant risk and how much confidence is enough.","core_discovery":"The central claim is that AI assurance will go wrong in three specific ways unless it adopts the discipline of engineered critical systems: it will address the wrong risks (because system boundaries are drawn too narrowly and 'existential' risks crowd out everyday harms), its techniques will be inadequate (because red-teaming and fine-tuning deliver very low confidence and there are no theories linking measured behaviour to deployed behaviour), and it will fail to communicate its claims (because confidence is not stated relative to the criticality of the deployment decision). The paper's remedy is to transfer the eight-step critical-systems process—environment definition, hazard analysis, safety requirements, verification, and an overall assurance case—and to run that process with the sceptical, deductive machinery of Assurance 2.0, including explicit defeaters and confirmation-theoretic evidence assessment.","pith_inferences":["If the paper's diagnosis is right, current capability-based risk ratings that focus on frontier models will miss most of the harm, because the same model embedded in many narrow applications multiplies modest risks into intolerable ones.","The paper's emphasis on theories that connect measured evidence to useful claims suggests that AI evaluation should invest in coverage and extrapolation arguments (for example, from evaluated subsets to the full operational distribution) rather than accumulating more red-team results.","A testable extension would be to audit existing corporate safety frameworks against the eight-step checklist and the Assurance 2.0 requirements; a framework that lacks an explicit hazard analysis or tolerability target would be predictably weak.","The four-state dependability model implies that safety frameworks should be judged not only by how they prevent loss but by how quickly they detect, contain, and recover from a bad deployment decision."],"forward_implications":["Safety frameworks should define the system as the socio-technical deployment context, not the model, and should include representative narrow-AI applications of frontier models as part of their remit.","Frameworks should set explicit tolerability targets (for example, failure rates many orders of magnitude below everyday software) and use design-basis events and threats to bound the risk analysis.","Assurance should be built as rigorous cases with deductive argument steps, defeaters for doubt, and confirmation-theoretic evidence, rather than red-teaming and fine-tuning alone.","Deployment decisions should assess the criticality of the decision itself—how much harm can occur before a bad decision is detected and recovered—alongside the criticality of the system.","Guards, monitors, diverse secondaries, and defence in depth should be part of the architecture whenever the AI component itself cannot be strongly assured."],"supporting_citations":[{"why":"The interim Seoul report whose risk framing and commitments the paper argues are too narrow, and whose 'safety is a system property' point it builds on.","marker":"[2]"},{"why":"The paper's source for the Assurance 2.0 method of confidence, defeaters, and deductive argument steps.","marker":"[10]"},{"why":"The companion dependability-perspective report that applies the eight-step model to AI systems.","marker":"[9]"},{"why":"The 'Guaranteed Safe AI' proposal, used as the partial recognition of the process that still omits hazard analysis, requirements validation, and the overall assurance case.","marker":"[15]"},{"why":"The 'r2p2' risk-tolerability framework, cited as the basis for explicit decisions about what risks are acceptable.","marker":"[19]"},{"why":"The historical source for structured safety cases and their role in challenge and deployment decision-making.","marker":"[38]"}],"fun_headline_variants":["AI assurance fails without critical-systems rigour","Three ways AI assurance goes wrong without safety engineering","AI assurance must widen boundaries and argue sceptically","Critical systems playbook: the fix for AI assurance failures","AI assurance without critical-systems discipline is doomed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the discipline that kept modern aircraft safe in service—especially the claim that no modern aircraft accident has come from a verification failure—can be carried over to frontier AI systems even though their deployment environments, architectures, and internal behaviour are largely unknown.","fun_headline_variants_meta":{"raw":{"variants":["AI assurance fails without critical-systems rigour","Three ways AI assurance goes wrong without safety engineering","AI assurance must widen boundaries and argue sceptically","Critical systems playbook: the fix for AI assurance failures","AI assurance without critical-systems discipline is doomed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000519,"raw_usage":{"total_tokens":2522,"prompt_tokens":960,"completion_tokens":1562,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":1489}},"tokens_in":576,"tokens_out":1562,"duration_ms":12095,"temperature":1.0,"reasoning_tokens":1489,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:39:29.926708+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A documented modern aircraft accident caused by a failure in Step 7 (verification), rather than by requirements validation or earlier steps, would falsify the paper's historical premise; alternatively, a demonstration that a frontier AI system's risk cannot be bounded by any design-basis event because the environment is intrinsically open-ended would test its central transferability claim.","supporting_citations":[{"cited_title":"International Scientific Report on the Safety of Advanced AI , Interim Report","cited_arxiv_id":null,"evidence_quote":"The interim Seoul report whose risk framing and commitments the paper argues are too narrow, and whose 'safety is a system property' point it builds on."},{"cited_title":"Confidence in Assurance 2.0 Cases","cited_arxiv_id":"2409.10665","evidence_quote":"The paper's source for the Assurance 2.0 method of confidence, defeaters, and deductive argument steps."},{"cited_title":"Technical report, Health and Safety Executive, Stationaery Office, Norwich UK, 2001","cited_arxiv_id":null,"evidence_quote":"The 'r2p2' risk-tolerability framework, cited as the basis for explicit decisions about what risks are acceptable."},{"cited_title":"The interpretation and evaluation of assurance cases","cited_arxiv_id":null,"evidence_quote":"The historical source for structured safety cases and their role in challenge and deployment decision-making."}],"review_version":1}