Pith. sign in

REVIEW 3 major objections 7 minor 1 cited by

Where AI Assurance Might Go Wrong: Initial lessons from engineering of critical systems

T0 review · 3 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read AI assurance will go wrong in three predictable ways unless it adopts the discipline of critical systems engineering.

desk verdict Worth engaging: an experienced, honest position piece whose recommendations survive its one unsourced empirical anchor. read the letter →

arxiv 2502.03467 v1 pith:PVQQJFE3 submitted 2025-01-07 cs.CY cs.AIcs.SE

classification cs.CYcs.AIcs.SE
keywords AIassurancesafetyframeworkscasescriticalsystemsengineeringrisktolerabilityhazardanalysis2.0foundationmodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that current AI safety frameworks are heading for three avoidable failures: they target the wrong risks, use assurance techniques too weak for the risks they do address, and communicate their confidence poorly. Drawing on how aircraft and nuclear systems achieve safety, it contends that AI assurance must treat safety as a property of the whole socio-technical system, not of the model, and must set explicit risk-tolerability targets before choosing methods. It advocates structured assurance cases that make arguments deductively tight, expose doubts as defeaters, and weigh evidence by how much it actually raises confidence. The payoff, if the paper is right, is a concrete checklist for making frontier AI safety frameworks more demanding and more honest about what they know.

What carries the argument

The load-bearing object is the structured assurance case in the 'Assurance 2.0' style: a tree of claims, argument, and evidence in which every argument step is expected to be deductive, every doubt is recorded as a defeater that must be refuted or accepted as residual risk, and evidence is scored by how much it increases confidence in the useful claim rather than the measured claim. Alongside it stands the eight-step safety-engineering process (environment, requirements, hazard analysis, safety requirements, validation, specification, verification, assurance case), which supplies the vocabulary for deciding what counts as a relevant risk and how much confidence is enough.

What would settle it

A documented modern aircraft accident caused by a failure in Step 7 (verification), rather than by requirements validation or earlier steps, would falsify the paper's historical premise; alternatively, a demonstration that a frontier AI system's risk cannot be bounded by any design-basis event because the environment is intrinsically open-ended would test its central transferability claim.

Watch

Extended reading notes

Core claim

The central claim is that AI assurance will go wrong in three specific ways unless it adopts the discipline of engineered critical systems: it will address the wrong risks (because system boundaries are drawn too narrowly and 'existential' risks crowd out everyday harms), its techniques will be inadequate (because red-teaming and fine-tuning deliver very low confidence and there are no theories linking measured behaviour to deployed behaviour), and it will fail to communicate its claims (because confidence is not stated relative to the criticality of the deployment decision). The paper's remedy is to transfer the eight-step critical-systems process—environment definition, hazard analysis, safety requirements, verification, and an overall assurance case—and to run that process with the sceptical, deductive machinery of Assurance 2.0, including explicit defeaters and confirmation-theoretic evidence assessment.

Load-bearing premise

The paper assumes that the discipline that kept modern aircraft safe in service—especially the claim that no modern aircraft accident has come from a verification failure—can be carried over to frontier AI systems even though their deployment environments, architectures, and internal behaviour are largely unknown.

Editorial extensions

If this is right

  • Safety frameworks should define the system as the socio-technical deployment context, not the model, and should include representative narrow-AI applications of frontier models as part of their remit.
  • Frameworks should set explicit tolerability targets (for example, failure rates many orders of magnitude below everyday software) and use design-basis events and threats to bound the risk analysis.
  • Assurance should be built as rigorous cases with deductive argument steps, defeaters for doubt, and confirmation-theoretic evidence, rather than red-teaming and fine-tuning alone.
  • Deployment decisions should assess the criticality of the decision itself—how much harm can occur before a bad decision is detected and recovered—alongside the criticality of the system.
  • Guards, monitors, diverse secondaries, and defence in depth should be part of the architecture whenever the AI component itself cannot be strongly assured.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's diagnosis is right, current capability-based risk ratings that focus on frontier models will miss most of the harm, because the same model embedded in many narrow applications multiplies modest risks into intolerable ones.
  • The paper's emphasis on theories that connect measured evidence to useful claims suggests that AI evaluation should invest in coverage and extrapolation arguments (for example, from evaluated subsets to the full operational distribution) rather than accumulating more red-team results.
  • A testable extension would be to audit existing corporate safety frameworks against the eight-step checklist and the Assurance 2.0 requirements; a framework that lacks an explicit hazard analysis or tolerability target would be predictably weak.
  • The four-state dependability model implies that safety frameworks should be judged not only by how they prevent loss but by how quickly they detect, contain, and recover from a bad deployment decision.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper distills the authors' experience in critical-systems assurance into an eight-step safety-engineering process, claims that traditional aviation accidents are attributable to requirements validation (Step 5) rather than verification failures (Step 7), and uses this base rate to argue that AI Safety Frameworks are likely to go wrong by addressing the wrong risks, relying on techniques inadequate for the risks they address, and communicating claims poorly. The paper recommends broadening system boundaries, making risk tolerability explicit, using design basis events and threats, adopting architectural guards and defence in depth, and building assurance cases based on the authors' Assurance 2.0 framework. It maps these recommendations to two questions from the FAISC call and closes with suggestions for how AI safety frameworks should evolve.

Significance. If the paper's lessons are accepted, its main contribution is to redirect AI-safety attention from verification-centric activities such as red-teaming and testing toward hazard analysis, requirements validation, and explicit risk-tolerability reasoning, and to give concrete architectural and assurance-case vocabulary for doing so. The paper is transparent about being selective and experience-based, and it offers several testable empirical claims (the aviation accident base rate; the absence of these topics in current frameworks), which is a strength because it allows readers to check the evidence behind the recommendations. Its significance is limited by the qualitative, analogy-based nature of the argument and by the fact that the positive recommendation (Assurance 2.0) is the authors' own framework, presented without independent evaluation; nevertheless, as an initial position paper it is a useful and readable contribution to the FAISC dialogue.

major comments (3)
  1. [Section 2, paragraph beginning 'There is extensive historical experience.'] The assertion that 'there have been no accidents of modern aircraft due to failures of Step 7 (Verification)' and that 'all modern aircraft failures have been attributed to Step 5' is load-bearing: it anchors the later argument that testing and red-teaming are not the bottleneck for AI assurance, and that hazard analysis and requirements validation should be prioritized. The assertion is given without citation and is partly definitional. Because the Step 5/Step 7 boundary is drawn in terms of whether the safety requirements captured the hazard, an accident caused by a dangerous behavior that a more contextual verification activity (e.g., integrated system testing, flight testing, or in-service analysis of learned behavior) would have caught will still be coded as Step 5 if the requirement was incomplete. The paper should either cite systematic accident-taxonomy studies supporting the base rate, or weaken the claim into an explicitly coded statement and discuss how classification may differ for machine-learning systems. The 737 MAX example supports the need for better hazard analysis, but it does not by itself establish the universal 'no Step 7 accidents' prior.
  2. [Section 3, paragraph beginning 'None of the corporate or national frameworks.'] The universal negative that 'none of the corporate or national frameworks that we have examined make any mention of these topics' is central to the paper's critique of current AI Safety Frameworks, but the examined frameworks are not identified, the review period is not given, and the criterion for 'mention' is not defined. A reader cannot verify whether this is a measured absence or a selective sample. Please provide a list (or at least a table) of the frameworks reviewed with dates and the specific topics searched for, or replace the universal claim with a bounded statement such as 'in the frameworks we reviewed, these topics were absent or barely developed.'
  3. [Sections 3.2.1 and 4] The paper recommends Design Basis Events and Design Basis Threats as a way to bound the open-ended risk space for AI, while acknowledging two paragraphs later that 'for AI applications it may be difficult to enumerate a set with adequate coverage.' This is a real tension in the transfer argument: design-basis reasoning is only useful if a justifiable worst-case set can be constructed, and the paper does not say how such a set would be built or validated for frontier AI. Since this recommendation underlies the call for broader system boundaries and risk tolerability analysis, the authors should either provide a worked example of a design basis for one AI application, or state more explicitly that the concept is being offered as an open question rather than a ready-made solution.
minor comments (7)
  1. [Section 2, items 3 and 5] 'e.q.' should be 'e.g.' and 'Requirements V alidation' contains a stray space.
  2. [Section 2] The phrase 'All modern aircraft failures' is ambiguous; please specify whether it means accidents, fatal accidents, or all system failures, and clarify the date range or aircraft generation covered by 'modern.'
  3. [References [20] and [22]] These two entries appear to be duplicates of the same IAEA report; please consolidate them.
  4. [Reference [19]] There is a typo: 'Stationaery Office' should be 'Stationery Office.'
  5. [Section 3.1] The statement that 'the only things certified by the FAA are airplanes and engines (and propellers)' is an oversimplification of FAA certification (which also includes type designs and Technical Standard Order authorizations); consider rewording to avoid an unnecessary quibble.
  6. [Section 3.1.1] 'CrowdStrike crash' should be 'CrowdStrike outage' or 'CrowdStrike incident,' since the event was not a crash in the technical sense.
  7. [Section 2, Step 8 and Section 3.4.1] The term 'indefeasible assurance' is used before it is defined; please define it at first use or add a pointer to the Appendix.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an experience-based position essay whose critical-systems lessons are independently grounded; self-citations are pointers to the authors' own framework, not load-bearing proof steps.

full rationale

This paper does not present a formal derivation, fitted model, or prediction that could reduce to its inputs by construction. Its central claims are that AI assurance may address the wrong risks, use inadequate techniques, and fail to communicate; these are argued from the authors' experience and from established safety-engineering literature (Laprie, Perrow, HSE r2p2, IAEA, ISO 21448, etc.). The one empirical anchor that is load-bearing is the claim in Section 2 that 'there have been no accidents of modern aircraft due to failures of Step 7 (Verification)' and that 'all modern aircraft failures have been attributed to Step 5 (Requirements Validation).' This is unsourced and could be contested, but it is not circular: it is a historical generalization about accident attribution, not a claim made true by the definitions of the eight steps. A stronger verification process might have caught some Step 5 failures, but that is an argument about adequacy, not about circularity of the paper's reasoning. The paper's advocacy of Assurance 2.0 is self-referential in the sense that the authors invented the framework and cite their own papers [8, 10, 12, 43], but those citations are used as pointers to the framework's details, not as independent evidence that the framework works. The surrounding lessons—broader system boundaries, hazard analysis, risk tolerability, assurance cases—are drawn from conventional critical-systems practice and do not depend on Assurance 2.0 being accepted. No equations are reused as conclusions, no fitted parameter is relabeled as a prediction, and no uniqueness theorem from the authors' prior work is invoked to forbid alternatives. The paper even acknowledges that Assurance 2.0 builds on the existing CAE approach, which further weakens any charge of renaming. Therefore no circular step meets the required standard of quoting a specific reduction of a result to its own inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The paper has no free parameters or fitted values; it is a qualitative argument. It relies on domain assumptions about safety engineering practice and its transferability to AI, which are argued from experience but not formally demonstrated.

assumptions (4)
  • domain assumption Safety is a system property; hazards are located in the environment.
    Section 3.1 argues that safety only makes sense in conjunction with an environment, which drives the recommendation for broader system boundaries.
  • domain assumption Modern aircraft accidents, including 737 MCAS, were due to requirements validation failures rather than verification failures.
    Section 2 uses this historical claim to argue that AI assurance should focus on hazard analysis and requirements, not just model verification.
  • domain assumption Critical systems require confidence levels orders of magnitude higher than everyday systems, and everyday evaluation techniques do not scale in rigor.
    Section 3.3 asserts that moving from everyday to critical systems requires different engineering, and uses this to argue that current AI evaluation methods are insufficient.
  • domain assumption The 8-step safety engineering process is a valid starting point for analyzing AI systems.
    Section 2 presents the traditional process and Section 3 applies it to AI, assuming its relevance.
invented entities (1)
  • AFGI (Artificial Fairly General Intelligence)
    purpose: To name and draw attention to near-term risks from 'good enough' AI that do not require malicious actors, such as unemployment and degraded job performance.
    Introduced in Section 3 as a conceptual category; no falsifiable handle outside the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Where AI Assurance Might Go Wrong: Initial lessons from engineering of critical systems." pith.science (2026). https://pith.science/paper/PVQQJFE3

@misc{pith2026250203467,
  author       = {Pith},
  title        = {Pith review of: Where AI Assurance Might Go Wrong: Initial lessons from engineering of critical systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PVQQJFE3}},
  note         = {Machine review of arXiv:2502.03467}
}
read the original abstract

We draw on our experience working on system and software assurance and evaluation for systems important to society to summarise how safety engineering is performed in traditional critical systems, such as aircraft flight control. We analyse how this critical systems perspective might support the development and implementation of AI Safety Frameworks. We present the analysis in terms of: system engineering, safety and risk analysis, and decision analysis and support. We consider four key questions: What is the system? How good does it have to be? What is the impact of criticality on system development? and How much should we trust it? We identify topics worthy of further discussion. In particular, we are concerned that system boundaries are not broad enough, that the tolerability and nature of the risks are not sufficiently elaborated, and that the assurance methods lack theories that would allow behaviours to be adequately assured. We advocate the use of assurance cases based on Assurance 2.0 to support decision making in which the criticality of the decision as well as the criticality of the system are evaluated. We point out the orders of magnitude difference in confidence needed in critical rather than everyday systems and how everyday techniques do not scale in rigour. Finally we map our findings in detail to two of the questions posed by the FAISC organisers and we note that the engineering of critical systems has evolved through open and diverse discussion. We hope that topics identified here will support the post-FAISC dialogues.

Figures

Figures reproduced from arXiv: 2502.03467 by the authors.

Figure 1
Figure 1. Resilience Type 2: resilience beyond design basis threats, events and use. This might be split into known threats that are considered incredible or ignored for some reason, and other “black swan” threats that are true unknowns. Often we are able to engineer systems successfully to cope with Type 1 re￾silience using methods of redundancy and fault tolerance. Type 2 resilience is a more formidable challenge. We may ch… view at source ↗
Figure 2
Figure 2. Four State Model for Dependability assurance that this is so. These correspond to Steps 1–6 and the accompanying parts of Step 8 in the engineering and assurance outline presented earlier. Safety engineering attempts to eliminate hazards (e.g., if fire is a hazard, then remove flammable material and sources of ignition), or to mitigate them (e.g., add a fire extinguishing system, but then failure of that system beco… view at source ↗
Figure 3
Figure 3. Ideas from Engineering Critical Systems 22 [PITH_FULL_IMAGE:figures/full_fig_p022_3.png] view at source ↗
Figures from the paper (1 more)
Figure 1
Figure 1. Figure 1: Assurance 2.0 Building Blocks and “Helping Hand” Mnemonic (from [ [PITH_FULL_IMAGE:figures/full_fig_p031_1.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PRBench: A Standardized Probabilistic Robustness Benchmark

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Across 222 models, standard adversarial training matches or beats dedicated probabilistic-robustness training on probabilistic robustness, while probabilistic methods show lower generalization gaps but near-zero adver...

Reference graph

Works this paper leans on

43 extracted references · 39 canonical work pages · cited by 1 Pith paper

  1. [1]

    Socio-technical systems: From design methods to systems engineering

    Gordon Baxter and Ian Sommerville. Socio-technical systems: From design methods to systems engineering. Interacting With Computers , 23(1):4--17, 2011

  2. [2]

    International Scientific Report on the Safety of Advanced AI , Interim Report

    Yoshua Bengio, editor. International Scientific Report on the Safety of Advanced AI , Interim Report . AI Seoul Summit, May 2024. \ ://www.gov.uk/government/publications/international-scientific-report-on-the-safety-of-advanced-ai

  3. [3]

    Requirements-driven model checking and test generation for comprehensive verification

    Devesh Bhatt, Hao Ren, Anitha Murugesan, Jason Biatek, Srivatsan Varadarajan, and Natarajan Shankar. Requirements-driven model checking and test generation for comprehensive verification. In NASA Formal Methods Symposium , Volume 13260 of Springer-Verlag Lecture Notes in Computer Science , pages 576--596, Springer-Verlag, Pasadena, CA, May 2022

  4. [4]

    Toward a formalism for conservative claims about the dependability of software-based systems

    Peter Bishop, Robin Bloomfield, Bev Littlewood, Andrey Povyakalo, and David Wright. Toward a formalism for conservative claims about the dependability of software-based systems. IEEE Transactions on Software Engineering , 37(5):708--717, 2011

  5. [5]

    R. E. Bloomfield and W. D. Ehrenberger. Validation and licensing of intelligent software. In Man Machine Interface in the Nuclear Industry . International Atomic Energy Agency, Vienna, Austria, February 1988. \ ://inis.iaea.org/search/search.aspx?orig_q=RN:20045938

  6. [6]

    Safety and assurance cases: Past, present and possible future---an Adelard perspective

    Robin Bloomfield and Peter Bishop. Safety and assurance cases: Past, present and possible future---an Adelard perspective. In Chris Dale and Tom Anderson, editors, Advances in System Safety: Proceedings of the Nineteenth Safety-Critical Systems Symposium , pages 51--67, Springer, Bristol, UK, February 2010

  7. [7]

    Evaluating the resilience and security of boundaryless, evolving socio-technical systems of systems

    Robin Bloomfield and Ilir Gashi. Evaluating the resilience and security of boundaryless, evolving socio-technical systems of systems. Technical report, Centre for Software Reliability, City University, London, UK, May 2008

  8. [8]

    Assurance 2.0: A Manifesto

    Robin Bloomfield and John Rushby. Assurance 2.0: A Manifesto . In Mike Parsons and Mark Nicholson, editors, Systems and Covid-19: Proceedings of the 29th Safety-Critical Systems Symposium (SSS'21) , pages 85--108, Safety-Critical Systems Club, York, UK, February 2021. Preprint available as arXiv:2004.10474

Show all 43 references
  1. [9]

    Assurance of AI systems from a dependability perspective

    Robin Bloomfield and John Rushby. Assurance of AI systems from a dependability perspective. Technical Report SRI-CSL-2024-02, Computer Science Laboratory, SRI International, Menlo Park, CA, July 2024. Also arXiv:2407.13948

  2. [10]

    Confidence in Assurance 2.0 Cases

    Robin Bloomfield and John Rushby. Confidence in Assurance 2.0 Cases . In Ana Cavalcanti and James Baxter, editors, The Practice of Formal Methods: Essays in Honour of Cliff Jones , Part I , Volume 14780 of Springer-Verlag Lecture Notes in Computer Science , pages 1--23, Spring...

  3. [11]

    Models are central to AI assurance

    Robin Bloomfield and John Rushby. Models are central to AI assurance. In ASSURE 2024, Proceedings of IEEE 35th International Symposium on Software Reliability Engineering Workshops (ISSREW) , pages 199--202, Tsukuba, Japan, October 2024

  4. [12]

    Assurance 2.0 home page

    Robin Bloomfield, John Rushby, et al. Assurance 2.0 home page . \ ://www.csl.sri.com/users/rushby/assurance2.0

  5. [13]

    Computer trading and systemic risk: a nuclear perspective

    Robin Bloomfield and Anne Wetherilt. Computer trading and systemic risk: a nuclear perspective. Forsight Driver Review DR26, Government Office for Science, London, UK, 2012. https://openaccess.city.ac.uk/id/eprint/1950/1/12-1059-dr26-computer-trading-and-systemic-risk-nuclear-...

  6. [14]

    On the opportunities and risks of foundation models

    Rishi Bommasani et al. On the opportunities and risks of foundation models. arXiv:2108.07258 , August 2021

  7. [15]

    Towards guaranteed safe AI : A framework for ensuring robust and reliable AI systems

    David Dalrymple et al. Towards guaranteed safe AI : A framework for ensuring robust and reliable AI systems. arXiv:2405.06624 , June 2024

  8. [16]

    Viability and Resilience of Complex Systems: Concepts, Methods and Case Studies from Ecology and Society

    Guillaume Deffuant and Nigel Gilbert. Viability and Resilience of Complex Systems: Concepts, Methods and Case Studies from Ecology and Society . Springer, 2011

  9. [17]

    A Tutorial on Runtime Verification

    Yli \`e s Falcone, Klaus Havelund, and Giles Reger. A Tutorial on Runtime Verification . In Manfred Broy, Doron Peled, and Georg Kalus, editors, Engineering Dependable Software Systems (Marktoberdorf Summer School Lectures, 2012) , pages 141--175. IOS Press, 2013

  10. [18]

    New nuclear power plants: Generic design assessment

    Office for Nuclear Regulation. New nuclear power plants: Generic design assessment. Technical Guidance ONR-GDA-GD-007, Bootle, UK, May 2019. https://onr.org.uk/media/c2delysl/onr-gda-007.pdf

  11. [19]

    Technical report, Health and Safety Executive, Stationaery Office, Norwich UK, 2001

    Reducing risks, protecting people: HSE 's decision-making process. Technical report, Health and Safety Executive, Stationaery Office, Norwich UK, 2001. https://www.hse.gov.uk/enforce/assets/docs/r2p2.pdf

  12. [20]

    IAEA Nuclear Energy Series NP-T-3.27, International Atomic Energy Agency, Vienna, Austria, 2018

    Dependability assessment of software for safety instrumentation and control systems at nuclear power plants. IAEA Nuclear Energy Series NP-T-3.27, International Atomic Energy Agency, Vienna, Austria, 2018. https://www-pub.iaea.org/MTCD/Publications/PDF/P1808_web.pdf

  13. [21]

    Human factors integration (HFI)

    Nuclear Safety Inspector. Human factors integration (HFI) . Technical Assessment Guide NIS-TAST-GD-058, Office for Nuclear Regulation, Bootle, UK, March 2023. \ ://onr.org.uk/media/documents/guidance/ns-tast-gd-058.docx

  14. [22]

    International Atomic Energy Agency, 2018

    Dependability Assessment of Software for Safety Instrumentation and Control Systems at Nuclear Power Plants . International Atomic Energy Agency, 2018. Nuclear Energy Series, NP-T-3.27

  15. [23]

    Technical Standard PAS 21448, International Organization for Standardization (ISO), 2019

    Road Vehicles: Safety of the Intended Functionality . Technical Standard PAS 21448, International Organization for Standardization (ISO), 2019

  16. [24]

    Millett, editors

    Daniel Jackson, Martyn Thomas, and Lynette I. Millett, editors. Software for Dependable Systems: Sufficient Evidence? National Academies Press, Washington, DC, May 2007

  17. [25]

    Real-Time Systems: Design Principles for Distributed Embedded Applications

    Hermann Kopetz and Wilfried Steiner. Real-Time Systems: Design Principles for Distributed Embedded Applications . Springer, 2022

  18. [26]

    Proofs and Refutations

    Imre Lakatos. Proofs and Refutations . Cambridge University Press, Cambridge, England, 1976

  19. [27]

    J. C. Laprie, editor. Dependability: Basic Concepts and Terminology in English , French, German, Italian and Japanese , Volume 5 of Springer-Verlag, Vienna, Austria Dependable Computing and Fault-Tolerant Systems . Springer-Verlag, Vienna, Austria, February 1991

  20. [28]

    Open Minded: Working Out the Logic of the Soul

    Jonathan Lear. Open Minded: Working Out the Logic of the Soul . Harvard University Press, 1999

  21. [29]

    Plato and the Nerd : The Creative Partnership of Humans and Technology

    Edward Ashford Lee. Plato and the Nerd : The Creative Partnership of Humans and Technology . MIT Press, 2017

  22. [30]

    Our big problem is not misinformation; it's knowingness

    Jonathan Malesic. Our big problem is not misinformation; it's knowingness. Psyche , March 2023. https://psyche.co/ideas/our-big-problem-is-not-misinformation-its-knowingness

  23. [31]

    Jeffrey C. Mogul. Emergent (mis)behavior vs. complex software systems. ACM SIGOPS Operating Systems Review , 40(4):293--304, 2006

  24. [32]

    Normal Accidents: Living with High Risk Technologies

    Charles Perrow. Normal Accidents: Living with High Risk Technologies . Basic Books, New York, NY, 1984

  25. [33]

    Technical report, Royal Academy of Engineering, London UK, Undated

    Building Resilience : Lessons from the Academy's Review of the National Security Risk Assessment Methodology . Technical report, Royal Academy of Engineering, London UK, Undated. https://raeng.org.uk/media/g31bttwt/raeng-building-resilience.pdf

  26. [34]

    Gene I. Rochlin. Defining ``High Reliability'' Organizations in Practice: a Taxonomic Prologue . In Karlene H. Roberts, editor, New Challenges to Understanding Organizations . Macmillan New York, 1993

  27. [35]

    Quality measures and assurance for AI software

    John Rushby. Quality measures and assurance for AI software. Technical Report SRI-CSL-88-7R, Computer Science Laboratory, SRI International, Menlo Park, CA, September 1988. Also available as NASA Contractor Report 4187

  28. [36]

    Critical system properties: Survey and taxonomy

    John Rushby. Critical system properties: Survey and taxonomy. Reliability Engineering and System Safety , 43(2):189--219, 1994

  29. [37]

    Runtime certification

    John Rushby. Runtime certification. In Martin Leucker, editor, Eighth Workshop on Runtime Verification: RV08 , Volume 5289 of Springer-Verlag Lecture Notes in Computer Science , pages 21--35, Springer-Verlag, Budapest, Hungary, April 2008

  30. [38]

    The interpretation and evaluation of assurance cases

    John Rushby. The interpretation and evaluation of assurance cases. Technical Report SRI-CSL-15-01, Computer Science Laboratory, SRI International, Menlo Park, CA, July 2015. Available at http://www.csl.sri.com/users/rushby/papers/sri-csl-15-1-assurance-cases.pdf

  31. [39]

    Technical Standard J3016, SAE International, April 2021

    Surface vehicle recommended practice. Technical Standard J3016, SAE International, April 2021

  32. [40]

    Scott D. Sagan. The Limits of Safety: Organizations, Accidents, and Nuclear Accident . Princeton Studies in International History and Politics. Princeton University Press, Princeton, NJ, 1993

  33. [41]

    Development of the ISO 21448

    Adam Schnellbach and Gerhard Griessnig. Development of the ISO 21448. In Alastair Walker, Rory V. O'Connor, and Richard Messnarz, editors, Systems, Software and Services Process Improvement (EuroSPI) , pages 585--593, Springer, Edinburgh, Scotland, September 2019

  34. [42]

    Open Systems Dependability: Dependability Engineering for Ever-Changing Systems

    Mario Tokoro, editor. Open Systems Dependability: Dependability Engineering for Ever-Changing Systems . CRC Press, 2013

  35. [43]

    Clarissa : Foundations, tools and automation for assurance cases

    Srivatsan Varadarajan, Robin Bloomfield, John Rushby, Gopal Gupta, Anitha Murugesan, Robert Stroud, Kateryna Netkachova, and Isaac Hong Wong. Clarissa : Foundations, tools and automation for assurance cases. In 42nd AIAA/IEEE Digital Avionics Systems Conference , Barcelona, Sp...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.