REVIEW 3 major objections 5 minor 1 cited by
Systematic Hazard Analysis for Frontier AI using STPA
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper argues that frontier AI companies' unstructured hazard identification processes can miss causal factors, and that applying STPA—a systems-theoretic hazard analysis method—to an AI Control scenario reveals those factors, thereby…
desk verdict A useful STPA worked example on the AI Control safety case, with a headline robustness claim that outruns the evidence in the paper. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Unsafe Control Action (UCA), defined as a control action that, in a particular context and worst-case environment, leads to a hazard. The machinery is STPA's four-stage process: (1) define losses and system-level hazards; (2) model the control structure as controllers, controlled processes, and feedback loops; (3) derive UCAs by testing every control action against four failure modes—not provided, provided but causing a hazard, provided at the wrong time or order, or provided for the wrong duration; and (4) construct Loss Scenarios that spell out the causal factors, across categories such as technical and human factors, that could lead to each UCA. This decomposition is what gives the analysis its systematic coverage and traceability.
What would settle it
Conduct a full STPA enumeration of the AI Control scenario without author selection of UCAs, and compare the complete set of loss scenarios and causal factors to the original safety-case sketch; the robustness claim stands only if the full enumeration yields at least one causal factor absent from the sketch that the chosen four-UCAs cross-section does not already cover.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that a structured STPA walkthrough of the AI Control scenario yields causal factors that the original safety-case sketch did not address. The analysis surfaces specific Unsafe Control Actions such as the scaffold layer's memory reset being ineffective (UCA11.1), the untrusted agent encoding state in local file storage across resets (UCA24.6), a delayed shutdown after an exfiltration attempt (UCA7.3), and a human audit feedback loop that can be manipulated (UCA17.2). The corresponding Loss Scenarios expose combined technical and human vulnerabilities, organisational authority gaps, and conflicting goals, leading the author to claim that STPA broadens the scope, improves traceability, and strengthens robustness of safety assurance for frontier AI. The paper is careful to note that STPA is not claimed to be unique; the claim is that the systematic decomposition makes hidden causal factors more apparent and provides a basis for cross-checking existing safety cases.
Load-bearing premise
The analysis assumes that the four Unsafe Control Actions chosen for detailed loss-scenario exploration are a representative cross-section of all the hazards in the AI Control scenario, so the causal factors highlighted as evidence of STPA's breadth could reflect the author's selection rather than the method's full coverage.
Editorial extensions
If this is right
- Frontier AI developers could use STPA outputs to inform capability thresholds, model evaluations, and emergency procedures, while feeding those governance processes' results back as inputs to the STPA.
- Safety cases built from STPA elements gain traceability from losses down to loss scenarios, making it practical to update assurance arguments when system contexts or assumptions change.
- The structured hierarchical decomposition makes hazard analysis more amenable to LLM-assisted generation, potentially reducing the cost of scaling safety analysis to additional threat models.
- STPA can be applied incrementally, focusing on the highest-priority losses first, so developers can see early benefits without a full, resource-heavy analysis.
- In the event of a serious incident, an existing STPA control structure model can be used to support rapid incident response and a CAST investigation, because both use the same system model.
Reading between the lines
- If STPA were applied to a real deployment with a full control structure model rather than the paper's illustrative boundary, the number of UCAs and loss scenarios would likely be far larger, meaning the demonstrated breadth here may understate the method's output.
- A controlled comparison—independent teams analysing the same scenario with STPA versus an unstructured expert workshop—would be the direct test of the robustness claim; the paper does not provide one.
- The scalability claim rests on the reliability of reasoning models, which the paper itself flags as experimental; a pilot that measures the agreement between LLM-generated STPA artefacts and expert review would settle whether the automation holds.
- If frontier AI developers published their hazard analysis methodologies in their safety frameworks, as the paper recommends, STPA or similar systematic methods could become a standard for third-party scrutiny of risk claims.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper evaluates whether STPA, a systems-theoretic hazard analysis methodology from safety engineering, can strengthen frontier AI safety assurance. The author applies STPA to the AI Control scenario described by Korbak et al. (2025), walking through the four STPA stages: identifying losses and hazards (focusing on L1, unauthorized IP exfiltration, and H1, LLM agent transmitting sensitive data), modeling the control structure (Figures 3-4), deriving Unsafe Control Actions for each control action, and developing loss scenarios from four selected UCAs (UCA7.3, UCA11.1, UCA17.2, UCA24.6). The paper claims STPA identifies causal factors that unstructured hazard analysis would miss, improving robustness, and recommends STPA as a complement to existing governance mechanisms such as capability thresholds, evaluations, and safety cases. Sections 4-6 discuss benefits, limitations, and recommendations; Appendix B sketches a design pattern for building Claims-Argument-Evidence safety cases from STPA outputs.
Significance. If the central claim were established, this paper would make a useful contribution to frontier AI safety methodology: it provides a concrete, traceable worked example of STPA applied to a realistic AI Control deployment scenario, with loss scenarios spanning human, organizational, operational, technical, and feedback-related factors (Section 3.5). Its strengths include a faithful four-stage STPA execution with explicit traceability from hazards to UCAs to loss scenarios (Tables 3-5), a candid statement that the four explored UCAs are not claimed to be the most significant (Section 3.3), and a constructive design pattern in Appendix B linking STPA outputs to CAE safety cases. The limitation that no baseline comparison with unstructured methods was conducted is effectively conceded in Section 6.2, where such a comparison is deferred to future pilot work. As it stands, the paper is an existence demonstration and methodological proposal; its comparative claim about robustness over unstructured analysis is not yet evidenced.
major comments (3)
- [Section 3.3, Table 4] The full set of Unsafe Control Actions derived from the control structure is not reported. The text states that 'the analysis found multiple UCAs for each control action,' yet Table 4 displays only one UCA for each of four control actions (CA7, CA11, CA17, CA24), and the UCAs for the remaining control actions identified in the control structure diagrams (Figures 3 and 4) are not shown. Because the paper's central claim concerns robustness against missed hazards, the reader cannot assess whether the method was applied systematically across the full control-action set; the complete UCA list should be provided in an appendix together with a count of UCAs per control action, so that the subset selected for Stage 4 can be situated within the full output of Stage 3.
- [Sections 3.3-3.4, Table 5] The choice of the four UCAs taken forward to loss-scenario development is made by the author without documented selection criteria. The paper says only that the UCAs 'were selected to present a cross-section of loss scenarios and causal factors,' and it explicitly disclaims that they are the most significant ('we do not claim these are more significant than others'). Since the loss scenarios in Table 5 are the primary evidence for the claim that STPA surfaces causal factors that would otherwise be missed, the paper needs to justify that the selection is representative of the hazards in the scenario (for example, by coverage of hazard types, by severity ordering, or by traceability to the prioritized loss L1), or alternatively to develop loss scenarios for the full UCA set; otherwise the diversity of causal factors found in Section 3.5 could be an artifact of the selection rather than of the method's completeness.
- [Abstract and Section 7 vs. Section 6.2] The central claim that 'STPA is able to identify causal factors that may be missed by unstructured hazard analysis methodologies thereby improving robustness' has no supporting comparative evidence in the manuscript. No unstructured hazard analysis (expert workshop, brainstorming session, forecasting exercise) was run on the same scenario; the only comparison offered is with Korbak et al. (2025), which is a safety-case sketch, a structured argumentative artifact whose incompleteness its own authors acknowledge, not the output of an unstructured method. Section 6.2, step 5, defers 'compared with the risks and hazards identified by current techniques' to a future pilot study, effectively conceding that the evidence required for the abstract's claim does not yet exist. The abstract and conclusion should be revised to state what the demonstration actually supports: STPA produces a structured, traceable set of plausible causal factors that can complement existing safety-case work, while whether unstructured methods would miss these factors remains an open empirical question.
minor comments (5)
- [Abstract] The abstract contains rendering artifacts with missing whitespace ('WeevaluateSTPA'sability...' and 'theLoss Scenariosthat...'); the abstract should be typeset cleanly.
- [Table 4] Table 4 lists one UCA per control action, but the text states that multiple UCAs were found for each control action; the table should be annotated to indicate whether it shows the complete UCA set for CA7, CA11, CA17 and CA24, or only a selected subset.
- [Figures 3 and 4] The control structure diagrams are difficult to read at the resolution provided, particularly the arrow labels and entity names; higher-resolution figures or a supplementary vector version would make the Stage 2 traceability argument easier to verify.
- [Section 3.5] The same loss-scenario identifier is used for multiple distinct causal factors (LS7.3.1 appears under both 'Organisational' and 'Operational'; LS11.1.1 under both 'Technical' and 'Feedback'); assigning unique sub-identifiers (e.g., LS7.3.1a, LS7.3.1b) would make the causal-factor inventory unambiguous.
- [Table 3] Only 4 of the 18 hazards are shown, with no explanation of why these four were selected; adding the complete hazard list in an appendix, or a note on the selection rationale, would support the paper's traceability emphasis and allow the reader to see how H1 relates to the omitted hazards.
Circularity Check
No circular derivation: the STPA outputs are structured analyses of an external scenario, not fitted inputs or self-citations.
full rationale
The paper's derivation chain is not circular in the sense defined by the review criteria. The STPA methodology is external to the author (Leveson & Thomas, 2018), and the scenario and threat model come from Korbak et al. (2025), an external source. The Unsafe Control Actions and Loss Scenarios are generated by applying STPA's structured process to that scenario; they are not fitted parameters, renamed inputs, or consequences of an assumed conclusion. The paper explicitly does not claim uniqueness for STPA ('We do not claim that STPA is unique in being able to identify these causal factors'), and no load-bearing argument depends on a self-citation by the same author. The core weakness is evidentiary rather than circular: the abstract's claim that STPA 'identifies causal factors that may be missed by unstructured hazard analysis methodologies' is not supported by a comparison against any actual unstructured analysis of the same scenario. The paper's own pilot recommendation (Section 6.2, step 5) defers exactly this comparison to future work ('compared with the risks and hazards identified by current techniques'), and the four explored UCAs are an author-selected cross-section (Section 3.3: 'we do not claim these are more significant than others--they were selected to present a cross-section'). These are limitations of empirical support and representativeness, not cases where the result reduces by construction to its inputs or where a fitted value is relabeled as a prediction. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption STPA is a valid and effective methodology for identifying hazards in complex sociotechnical systems.
- domain assumption The threat model and scenario from Korbak et al. (2025) are a faithful representation of a relevant frontier AI deployment.
- domain assumption Frontier AI safety frameworks described in the cited company documents lack systematic hazard identification methodology.
- ad hoc to paper The selected subset of UCAs is representative of the hazards in the scenario.
Cite this review
Pith. "Pith review of Systematic Hazard Analysis for Frontier AI using STPA." pith.science (2026). https://pith.science/paper/IKDFZQ7M
@misc{pith2026250601782,
author = {Pith},
title = {Pith review of: Systematic Hazard Analysis for Frontier AI using STPA},
year = {2026},
howpublished = {\url{https://pith.science/paper/IKDFZQ7M}},
note = {Machine review of arXiv:2506.01782}
}
read the original abstract
All of the frontier AI companies have published safety frameworks where they define capability thresholds and risk mitigations that determine how they will safely develop and deploy their models. Adoption of systematic approaches to risk modelling, based on established practices used in safety-critical industries, has been recommended, however frontier AI companies currently do not describe in detail any structured approach to identifying and analysing hazards. STPA (Systems-Theoretic Process Analysis) is a systematic methodology for identifying how complex systems can become unsafe, leading to hazards. It achieves this by mapping out controllers and controlled processes then analysing their interactions and feedback loops to understand how harmful outcomes could occur (Leveson & Thomas, 2018). We evaluate STPA's ability to broaden the scope, improve traceability and strengthen the robustness of safety assurance for frontier AI systems. Applying STPA to the threat model and scenario described in 'A Sketch of an AI Control Safety Case' (Korbak et al., 2025), we derive a list of Unsafe Control Actions. From these we select a subset and explore the Loss Scenarios that lead to them if left unmitigated. We find that STPA is able to identify causal factors that may be missed by unstructured hazard analysis methodologies thereby improving robustness. We suggest STPA could increase the safety assurance of frontier AI when used to complement or check coverage of existing AI governance techniques including capability thresholds, model evaluations and emergency procedures. The application of a systematic methodology supports scalability by increasing the proportion of the analysis that could be conducted by LLMs, reducing the burden on human domain experts.
Figures
Forward citations
Cited by 1 Pith paper
-
HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study
A hybrid FMEA/FTA safety-analysis framework for foundation-model-based autonomous driving, illustrated on a GenAD and GAIA-2 style reference architecture.
Reference graph
Works this paper leans on
-
[1]
Ahlbrecht, A., Sprockhoff, J., & Durak, U. (2024). A system-theoretic assurance frame- work for safety-driven systems engineering.Software and Systems Modeling,24(1), 253– 270.https://doi.org/10.1007/s10270-024-01209-6
-
[2]
(2025, March 31).Anthropic’s Responsible Scaling Policy v2.1
Anthropic. (2025, March 31).Anthropic’s Responsible Scaling Policy v2.1. Anthropic
work page 2025
-
[3]
Ayvali, E. (2025).Systems-Theoretic Process Analysis: Un- covering Systemic AI Risks.https://medium.com/@eayvali/ systems-theoretic-process-analysis-uncovering-systemic-ai-risks-6d99ed3de9f7
work page 2025
-
[4]
Bloomfield, R., & Netkachova, K. (2014). Building Blocks for Assurance Cases.2014 IEEE International Symposium on Software Reliability Engineering Workshops, 186– 191.https://doi.org/10.1109/ISSREW.2014.72 23
-
[5]
(2025).Emerging Practices in Frontier AI Safety
Buhl, M., Bucknall, B., & Masterson, T. (2025).Emerging Practices in Frontier AI Safety
work page 2025
-
[6]
D., Sett, G., Koessler, L., Schuett, J., & Anderljung, M
Buhl, M. D., Sett, G., Koessler, L., Schuett, J., & Anderljung, M. (2024).Safety cases for frontier AI(arXiv:2410.21572). arXiv.http://arxiv.org/abs/2410.21572
arXiv 2024
-
[7]
Campos, S., Papadatos, H., Roger, F., Touzet, C., Quarks, O., & Murray, M. (2025). A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management(arXiv:2502.06656). arXiv.https://doi. org/10.48550/arXiv.2502.06656
-
[8]
(2025).AI for AI Safety.https://joecarlsmith.substack.com/p/ ai-for-ai-safety
Carlsmith, J. (2025).AI for AI Safety.https://joecarlsmith.substack.com/p/ ai-for-ai-safety
work page 2025
Show all 53 references
-
[9]
(2019).Providing for Safety in Rail Megaprojects
Chatzimichailidou, M., & Dunsford, R. (2019).Providing for Safety in Rail Megaprojects. WSP.https://www.wsp.com/-/media/Insights/Global/Documents/ Providing-for-Safety-in-Rail-Megaprojects.pdf
2019
-
[10]
(2024).Safety Cases: How to Justify the Safety of Advanced AI Systems(arXiv:2403.10462)
Clymer, J., Gabrieli, N., Krueger, D., & Larsen, T. (2024).Safety Cases: How to Justify the Safety of Advanced AI Systems(arXiv:2403.10462). arXiv.http://arxiv.org/abs/ 2403.10462
2024 arXiv
-
[11]
N., Rahim, E., & Brewster, S
Dawson, M., Burrell, D. N., Rahim, E., & Brewster, S. (2010).INTEGRATING SOFTWARE ASSURANCE INTO THE SOFTWARE DEVELOPMENT LIFE CY- CLE (SDLC).3(6)
2010
- [12]
-
[13]
Endsley, M. R. (1995). Toward a Theory of Situation Awareness in Dynamic Systems. Human Factors: The Journal of the Human Factors and Ergonomics Society,37(1), 32–64.https://doi.org/10.1518/001872095779049543
1995 doi
-
[14]
Falzone, T., & Thomas, J. P. (2021).STPA at Google. Google
2021
-
[15]
D., Schuett, J., Korbak, T., Wang, J., Hilton, B., & Irv- ing, G
Goemans, A., Buhl, M. D., Schuett, J., Korbak, T., Wang, J., Hilton, B., & Irv- ing, G. (2024).Safety case template for frontier AI: A cyber inability argument (arXiv:2411.08088). arXiv.https://doi.org/10.48550/arXiv.2411.08088
-
[16]
(2025).GDM Frontier Safety Framework 2.0
Google Deepmind. (2025).GDM Frontier Safety Framework 2.0
2025
-
[17]
S., & Lehman, S
Graydon, M. S., & Lehman, S. M. (2025).Examining Proposed Uses of LLMs to Produce or Assess Assurance Arguments
2025
-
[18]
Greenblatt, R. (2025, January 23).AI companies are unlikely to make high-assurance safety cases if timelines are short—Ryan Greenblatt, LessWrong.pdf.https://www.lesswrong.com/posts/neTbrpBziAsTH5Bn7/ ai-companies-are-unlikely-to-make-high-assurance-safety 24
2025
- [19]
-
[20]
(2009).The Nimrod Review
Haddon-Cave, C. (2009).The Nimrod Review
2009
-
[21]
D., Korbak, T., & Irving, G
Hilton, B., Buhl, M. D., Korbak, T., & Irving, G. (2024).Safety Cases: A Scalable Approach to Frontier AI Safety
2024
-
[22]
INCOSE. (2015). A Complexity Primer for Systems Engineers.INCOSE
2015
-
[23]
(2023).INCOSE Systems Engineering Handbook 5th Edition
INCOSE. (2023).INCOSE Systems Engineering Handbook 5th Edition
2023
-
[24]
(2018).ISO 31000_2018 Risk Management Guidelines.pdf
ISO. (2018).ISO 31000_2018 Risk Management Guidelines.pdf
2018
-
[25]
W., & Holloway, C
Johnson, C. W., & Holloway, C. M. (2003). The ESA/NASA SOHO mission interrup- tion: Using the STAMP accident analysis technique for a software related ’mishap’. Software: Practice and Experience,33(12), 1177–1198.https://doi.org/10.1002/ spe.544
2003
-
[26]
(2024).A Sketch of Potential Tripwire Capabilities for AI
Karnofsky, H. (2024).A Sketch of Potential Tripwire Capabilities for AI
2024
-
[27]
(2023).Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems
Khlaaf, H. (2023).Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems
2023
-
[28]
(2023).Risk assessment at AGI companies: A review of pop- ular risk assessment techniques from other safety-critical industries(arXiv:2307.08823)
Koessler, L., & Schuett, J. (2023).Risk assessment at AGI companies: A review of pop- ular risk assessment techniques from other safety-critical industries(arXiv:2307.08823). arXiv.https://doi.org/10.48550/arXiv.2307.08823
- [29]
-
[30]
(2019).The CAST Handbook
Leveson, N. (2019).The CAST Handbook
2019
-
[31]
Leveson, N., & Thomas, J. P. (2018).STPA Handbook
2018
-
[32]
P., Franch, X., & Nakagawa, E
Martínez-Fernández, S., Ayala, C. P., Franch, X., & Nakagawa, E. Y. (2015). A Survey on the Benefits and Drawbacks of AUTOSAR.Proceedings of the First International Workshop on Automotive Software Architecture, 19–26.https://doi.org/10.1145/ 2752489.2752493
2015
-
[33]
(2025).Meta Frontier AI Framework
Meta. (2025).Meta Frontier AI Framework
2025
-
[34]
(2025).Common Elements of Frontier AI Safety Policies, March 2025
METR. (2025).Common Elements of Frontier AI Safety Policies, March 2025
2025
-
[35]
Monkhouse, Helen Elizabeth, & Ward, David. (2024). A Hazard Analysis Ap- proach for Automated Driving Shared Control.SAE, 12.https://doi.org/10.4271/ 2024-01-2056 25
2024
-
[36]
(2022).Risk Analysis and Assessment Modeling Language (RAAML) Libraries and Profiles, v1.0
Object Management Group. (2022).Risk Analysis and Assessment Modeling Language (RAAML) Libraries and Profiles, v1.0
2022
-
[37]
(2024).OpenAI Preparedness Framework
OpenAI. (2024).OpenAI Preparedness Framework. OpenAI
2024
- [38]
-
[39]
Rose, R. L. (2024).Limitations of Commercial Aviation Safety Assessment Standards Uncovered in the Wake of the Boeing 737 MAX Accidents
2024
-
[40]
(2025).AI policy needs an increased focus on inci- dent preparedness
Shaffer Shane, T., & Robinson, B. (2025).AI policy needs an increased focus on inci- dent preparedness. The Centre for Long Term Resilience.https://doi.org/10.71172/ dwaa-x7wy
2025
-
[41]
Spencer, M. B. (2012).Engineering Financial Safety: A System-Theoretic Case Study from the Financial Crisis
2012
-
[42]
(2007).The Black Swan: The Impact of the Highly Improbable
Taleb, N. (2007).The Black Swan: The Impact of the Highly Improbable
2007
-
[43]
(2024).STPA Step 4 Building Scenarios A Formal Scenario Ap- proach.https://psas.scripts.mit.edu/home/wp-content/uploads/2024/ STPA-Scenarios-New-Approach.pdf
Thomas, J. (2024).STPA Step 4 Building Scenarios A Formal Scenario Ap- proach.https://psas.scripts.mit.edu/home/wp-content/uploads/2024/ STPA-Scenarios-New-Approach.pdf
2024
-
[44]
Thomas, J. P. (2024).Evaluation of System-Theoretic Process Analysis (STPA) for Im- proving Aviation Safety.https://rosap.ntl.bts.gov/view/dot/78914/dot_78914_ DS1.pdf
2024
-
[45]
(2016).STAMP Applied to Fukushima Daiichi Nuclear Disaster and the Safety of Nuclear Power Plants in Japan
Uesako, D. (2016).STAMP Applied to Fukushima Daiichi Nuclear Disaster and the Safety of Nuclear Power Plants in Japan. A Glossary •Accident: (We avoid using this term as we want to cover for example deliberate misuse). •Control Action: A command or directive issued by a contro...
2016
-
[46]
The top level claim states the system is safe to deploy in a certain context
-
[47]
If they can be prevented then the top level claim will be valid
This is supported by enumeration or decomposition arguments: breaking down the overall system into a number of losses which would be considered unacceptable to the stakeholders. If they can be prevented then the top level claim will be valid. 27
-
[48]
This produces a second level of claims–each one stating a specific loss is prevented
-
[49]
For each of these losses, a decomposition or enumeration argument is made that it can be prevented by preventing the hazards that contribute to it
-
[50]
A claim that each hazard is prevented is supported by an argument that specific control actions prevent the hazard from occurring
-
[51]
An argument is then made that each control action could be unsafe in certain ways, identifying unsafe control actions
-
[52]
For each of these unsafe control actions, another argument is then made for the loss scenarios that contribute to them
-
[53]
Finally, an evidence incorporation argument is made that specific mitigations are in place or safety requirements have been met. By following such a design pattern, it is possible to build confidence that the safety case has not missed any causal factors (loss scenarios) that ...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.