REVIEW 5 minor 249 references
Model benchmarks cannot certify AI safety; safety lives in the whole sociotechnical system.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 02:17 UTC pith:Y72N2NNU
load-bearing objection A solid sociotechnical synthesis that deserves a serious referee—the critique of component-level safety is well supported, the positive organizational-governance agenda is more prescription than proof, and the paper mostly admits that itself.
Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that safety is an emergent property of a composed sociotechnical system, not a property of individual components—models, datasets, or their performance scores. At any level of performance, or for any tradeoff between failure modes ('at any AUC'), a system that causes harm in its context of use is unsafe. Safety therefore exists only for the whole assemblage of technology, people, organizations, incentives, and environment; without a context of use, claims about safe component performance are 'technical artifice.' The paper distills six 'unlearned lessons' from industrial disasters and maps each to common AI practices, arguing that the same organizational dynamics
What carries the argument
The analytic mechanism is the systems-safety lens, which distinguishes component reliability from system-level emergent behavior. The load-bearing principle is 'safe components do not imply safe systems': failures arise from interactions among correctly functioning parts, so verification, benchmarking, and reliability assurance are necessary but never sufficient for safety. This principle drives the paper's taxonomy of six unlearned lessons, which it uses to connect disaster case studies to current AI practices and to motivate sociotechnical interventions such as safety cultures, curmudgeons, traceability, and structural incentives.
Load-bearing premise
The load-bearing premise is that the organizational failure mechanisms documented in other safety-critical industries—and the accuracy of the retold disaster accounts—transfer to modern AI development and deployment; if that analogy fails, the taxonomy is illustrative rather than evidence-bearing.
What would settle it
An empirical study showing that real-world AI harms are fully explained by model performance metrics, with no residual variation attributable to organizational safety culture, would falsify the central claim. More directly, a demonstrated counterexample of a high-AUC, component-focused system with zero recorded harms and no organizational safety practices would undermine the thesis that safety requires system-level governance.
If this is right
- Safety claims about AI must be contextual: a system is safe only within a defined context of use, not in the abstract.
- Benchmarks, audits, red-teaming, and documentation are useful only when embedded in organizational processes that act on them; as standalone rituals they can create the illusion of safety.
- Organizations deploying AI should institutionalize protected dissent, traceable decision records, pre-mortems, blameless postmortems, and leaders with authority to block launches.
- AI safety research should broaden from component-level fixes to system-level interventions based on accident-modeling methods.
- Regulation should aim to strengthen organizational risk perception and control, not merely mandate technical tests or disclosures.
Where Pith is reading between the lines
- If the paper is right, model cards and benchmark scores are the wrong primary audit objects; governance practices—like whether a safety review can veto a release—would carry most of the predictive signal for real-world harm.
- A testable extension: compare deployments matched on model performance but differing in safety-culture practices; the paper predicts the organizational dimension, not AUC, explains downstream harm rates.
- The same logic implies that procurement standards and regulatory frameworks should require assurance cases with explicit claims, evidence, and skeptical review rather than metric dashboards.
- A further consequence is that incident reporting and whistleblower protections, long used in aviation and medicine, may be more consequential for AI safety than any technical alignment method.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that AI safety is a property of full sociotechnical systems rather than of individual models or technical components. Drawing on safety science and a set of well-known disasters—Challenger, Chernobyl, USS Scorpion, Three Mile Island, Boeing 737 MAX, Silicon Valley Bank, the Uber autonomous-vehicle crash, and others—the authors identify six “unlearned lessons” (poor risk perception; hazardous incentives; permanent rush cultures; suppression of bad news; safe components not implying safe systems; and the “kick the can” displacement of responsibility onto operators). They then translate these lessons into a set of organizational and governance-oriented practices: authentic safety culture, curmudgeons and critics, traceable processes, psychological safety and diversity, meaningful external participation, treating failure as normal, and structural incentives for self-regulation. The paper uses Site Reliability Engineering as a precedent, critiques component-level responsible-AI tools such as documentation and benchmarks, and proposes a research agenda centered on applying systems-safety methods such as STAMP to AI systems. The central normative claim is that reliability or benchmark performance is not safety and that safety can only be assessed in context of the composed sociotechnical system.
Significance. If the paper's position is accepted, it would usefully reframe responsible-AI evaluation away from model metrics and toward organizational and governance structures. The paper's strengths are its careful synthesis of an existing safety-science literature, its concrete taxonomy of recurring failure etiologies, and its repeated honesty about open empirical questions: §4.4 notes that documentation practices require validation, §5.3 calls governance-practice validation a critical open problem, and §5.4 explicitly asks whether STAMP can translate to AI systems. These disclaimers substantially disarm the otherwise obvious objection that the positive program is unproven: the paper presents itself as a grounded research agenda rather than a completed empirical demonstration. It also makes a valuable intervention by distinguishing safety from reliability and by arguing that component-level tools such as AUC, benchmarks, and alignment techniques cannot substitute for system-level analysis. The paper is synthetic rather than experimental, but it is a significant contribution to the responsible-AI literature.
minor comments (5)
- [§3.5] The Chernobyl account is internally inconsistent and should be clarified. The text says operators made “seemingly insignificant changes to components of the reactor’s control rods with inadequate records” but then states that investigation “did not uncover specific component failures or operator deviation from procedure.” Chernobyl is one of the paper's recurring anchor cases, and the standard accident literature (including INSAG-7) does identify operator deviations in addition to design flaws. Please rewrite this passage to align with the cited sources or replace the Chernobyl example with a less contested instance of component-level correctness failing to imply system safety.
- [§6] The concluding sentence “System behavior safety exists only for this composed assemblage” is a strong (and defensible) conceptual claim, but it could be misread as an empirically established result. Since §4.4 and §5.3 explicitly acknowledge that the proposed organizational practices are not yet validated, consider adding one sentence in §6 that distinguishes the conceptual redefinition of safety from the empirical research agenda the paper proposes.
- [§5.3] Minor technical errors: “ISO 420001” should be “ISO/IEC 42001,” and “International Organizations for Standardization” should be “International Organization for Standardization.”
- [§5.1] Typo: “Preparadness” should be “Preparedness” in the list of themes from the SRE literature.
- [References] The Khlaaf (2023) reference lists the arXiv identifier “2606.29390,” which is inconsistent with a 2023 publication year. Please verify the identifier and date.
Circularity Check
No significant circularity: the sociotechnical-safety thesis is supported by external safety-science literature and case evidence; self-citations are background, not load-bearing.
full rationale
The paper is an argumentative synthesis, not a derivation with fitted parameters, equations, or predictions. Its central thesis—component reliability (e.g., AUC) does not entail system safety, which is an emergent property of the whole sociotechnical assemblage—is supported by independent external evidence: Vaughan's Challenger study, Perrow's normal accidents, Leveson's STAMP literature, Elish's moral crumple zone, and the Obermeyer, Boeing, SVB, and Uber case accounts. The paper explicitly adopts a broad stipulative definition of 'AI system' that includes organizational processes (Section 1), so the conclusion that organizational factors matter is partly built into the framing; but the load-bearing, non-circular step is the demonstration that failures occur even when components work to specification (Section 3.5, Chernobyl, cybersecurity incidents), which comes from outside the paper and does not presuppose the conclusion. There are several self-citations (Kroll 2020/2021; Geiger et al. 2018/2024; Smart & Kasirzadeh 2024; Jatho & Kroll 2022; Rismani et al. 2023; Abdu & Jacobs 2026), but none is the sole support for the central claim: they are used for background concepts such as traceability, accountability, documentation practice, and prior STAMP applications, and the main argument rests on the independently citable safety-science canon. The paper also explicitly flags its own open questions ('can the utility of STAMP translate to AI systems?', Section 5.4; documentation 'is, in fact, an empirical question requiring validation', Section 4.4), so it does not present its proposed interventions as already validated. The skeptic's complaint—that the positive thesis about organizational/governance locus outruns the evidence—is a scope-of-claim or evidence-weight concern, which the review rules classify as correctness risk, not circularity. Accordingly, no circular step meets the quoted-reduction standard; score 2 reflects minor non-load-bearing self-citation only.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Accident case studies (Challenger, Chernobyl, SVB, Boeing) provide transferable causal knowledge about organizational failure that applies to AI systems.
- domain assumption Safety is an emergent system-level property not decomposable into component reliability.
- domain assumption The cited retrospective accounts of disaster causation (Vaughan on Challenger, Plokhy on Chernobyl, Washington Post on SVB) are accurate and uncontested.
- ad hoc to paper The six-part 'unlearned lessons' taxonomy is a valid and useful partition of failure etiologies.
read the original abstract
As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the technology's capabilities, limitations, and attendant risks in context requires analysis of full sociotechnical systems. Sociotechnical analysis of risks in highly complex systems provides clear lessons for the design and evaluation of AI systems, transcending a technical focus on reliable or "responsibly designed" components to understand risks at a systems level. Human-made catastrophes have been studied for decades because of the severity of these events: consider Chernobyl, Three Mile Island, Fukushima-Daiichi, Bhopal, the Challenger disaster. A common misconception is that these kinds of events are freak accidents, resulting from the inherently unforeseeable interactions in complex systems. Closer examination reveals that the risks and hazards were well-known beforehand but not acted upon due to social structural, political and economic factors. We outline several areas where the development and use of AI can benefit from learning these unlearned lessons: improved risk perception, communication, and analysis at the organizational level; traceability of requirements and responsibilities; and holistic approaches to responsibility and safety that include social and organizational dynamics as first-order engineering concerns. For each area, we offer concrete unlearned lessons and exemplify how they led to failure in prior accidents as well as examples of how these lessons remain unlearned for modern computing systems, particularly AI.
Reference graph
Works this paper leans on
-
[1]
UCLA Law Review , volume=
The Public Harms of Private Surveillance , author=. UCLA Law Review , volume=. 2025 , URL=
2025
-
[2]
First Monday , doi =
Ahmed, Shazeda and Ja. First Monday , doi =
-
[3]
Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =
Ali, Sanna J and Christin, Ang. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =
2023
-
[4]
Communications of the ACM , volume =
Computational complexity and information asymmetry in financial products , author =. Communications of the ACM , volume =. 2011 , publisher =
2011
-
[5]
Responsive
Ayres, Ian and Braithwaite, John , year =. Responsive
-
[6]
Analysis, Design and Evaluation of Man--Machine Systems , pages =
Ironies of automation , author =. Analysis, Design and Evaluation of Man--Machine Systems , pages =. 1983 , publisher =
1983
-
[7]
Gebru, Timnit and Torres, Émile P. , urldate =. The. doi:10.5210/fm.v29i4.13636 , shorttitle =
-
[8]
2020 , publisher =
Can artificial intelligence transform higher education? , author =. 2020 , publisher =
2020
-
[9]
arXiv preprint arXiv:2501.17805 , year =
Bengio, Yoshua and Mindermann, S. arXiv preprint arXiv:2501.17805 , year =
-
[10]
doi:10.48550/arXiv.2511.19863 , year =
Bengio, Yoshua and Clare, Stephen and Prunkl, Carina and Andriushchenko, Maksym and Bucknall, Ben and Fox, Philip and Maslej, Nestor and McGlynn, Conor and Murray, Malcolm and Rismani, Shalaleh and others , journal =. doi:10.48550/arXiv.2511.19863 , year =
-
[11]
2016 , URL =
Site reliability engineering: How Google runs production systems , author =. 2016 , URL =
2016
-
[12]
Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , pages =
Robot rights? Let's talk about human welfare instead , author =. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , pages =
-
[13]
2021 IEEE Winter Conference on Applications of Computer Vision (WACV) , pages =
Large image datasets: A pyrrhic win for computer vision? , author =. 2021 IEEE Winter Conference on Applications of Computer Vision (WACV) , pages =. 2021 , organization =
2021
-
[14]
1975 , publisher =
The mythical man-month: essays on software engineering , author =. 1975 , publisher =
1975
-
[15]
2018 , publisher =
Artificial unintelligence: How computers misunderstand the world , author =. 2018 , publisher =
2018
-
[16]
Big data & society , volume =
How the machine ‘thinks’: Understanding opacity in machine learning algorithms , author =. Big data & society , volume =. 2016 , publisher =
2016
-
[17]
Annual Review of Sociology , volume =
The Society of Algorithms , author =. Annual Review of Sociology , volume =. 2020 , publisher =
2020
-
[18]
1981 , publisher =
Checkland, Peter , title =. 1981 , publisher =
1981
-
[19]
2016 , publisher =
Man-made catastrophes and risk information concealment , author =. 2016 , publisher =
2016
-
[20]
Journal of International Business Studies , volume =
Firm self-regulation through international certifiable standards: Determinants of symbolic versus substantive implementation , author =. Journal of International Business Studies , volume =. 2006 , publisher =
2006
-
[21]
Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages =
Reviewable automated decision-making: A framework for accountable algorithmic systems , author =. Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages =. 2021 , doi =
2021
-
[22]
Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =
Understanding accountability in algorithmic supply chains , author =. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2023 , doi =
2023
-
[23]
2019 , publisher =
Between truth and power: The legal constructions of informational capitalism , author =. 2019 , publisher =
2019
-
[24]
Proceedings of the 2022 ACM conference on fairness, accountability, and transparency , pages=
Accountability in an algorithmic society: relationality, responsibility, and robustness in machine learning , author=. Proceedings of the 2022 ACM conference on fairness, accountability, and transparency , pages=. 2022 , doi =
2022
-
[25]
2020 , location =
Design justice: Community-led practices to build the worlds we need , author =. 2020 , location =
2020
-
[26]
2002 , publisher =
Dekker, Sidney and Woods, David D , journal =. 2002 , publisher =
2002
-
[27]
2016 , publisher =
Drift into failure: From hunting broken components to understanding complex systems , author =. 2016 , publisher =
2016
-
[28]
Government Information Quarterly , volume =
Integral system safety for machine learning in the public sector: An empirical account , author =. Government Information Quarterly , volume =. 2024 , publisher =
2024
-
[29]
Oxford Handbook of
System Safety and Artificial Intelligence , author =. Oxford Handbook of. 2022 , doi =
2022
-
[30]
Toward Sociotechnical
Dobbe, Roel and Wolters, Anouk , journal =. Toward Sociotechnical. 2024 , publisher =
2024
-
[31]
2025 , doi =
Dobbe, Roel , journal =. 2025 , doi =
2025
-
[32]
Industry and Innovation , volume =
Learning from a drastic failure: the case of the Airbus A380 program , author =. Industry and Innovation , volume =. 2014 , publisher =
2014
-
[33]
ACM SIGSOFT Software Engineering Notes , volume =
The Ariane 5 software failure , author =. ACM SIGSOFT Software Engineering Notes , volume =. 1997 , publisher =
1997
-
[34]
International Conference on Computer Safety, Reliability, and Security , pages =
Byzantine fault tolerance, from theory to reality , author =. International Conference on Computer Safety, Reliability, and Security , pages =. 2003 , organization =
2003
-
[35]
American Journal of Sociology , volume=
Legal ambiguity and symbolic structures: Organizational mediation of civil rights law , author=. American Journal of Sociology , volume=. 1992 , publisher=
1992
-
[36]
American journal of Sociology , volume =
The endogeneity of legal regulation: Grievance procedures as rational myth , author =. American journal of Sociology , volume =. 1999 , publisher =
1999
-
[37]
Explaining Compliance: Business Responses to Regulation , pages=
To comply or not to comply—That isn’t the question: How organizations construct the meaning of compliance , author=. Explaining Compliance: Business Responses to Regulation , pages=. 2011 , publisher=
2011
-
[38]
Praise the machine! Punish the human! The contradictory history of accountability in automated aviation , author =
-
[39]
Engaging Science, Technology, and Society , volume = 5, year =
Moral crumple zones: Cautionary tales in human-robot interaction , author =. Engaging Science, Technology, and Society , volume = 5, year =
-
[40]
Repairing innovation: A study of integrating
Elish, Madeleine Clare and Watkins, Elizabeth Anne , institution =. Repairing innovation: A study of integrating
-
[41]
1996 , publisher =
Tyranny of the bottom line: Why corporations make good people do bad things , author =. 1996 , publisher =
1996
-
[42]
1963 , publisher =
Engines of culture: Philanthropy and art museums , author =. 1963 , publisher =
1963
-
[43]
Handbook of ethics, values, and technological design: sources, theory, values and application domains , pages =
Design for values and operator roles in sociotechnical systems , author =. Handbook of ethics, values, and technological design: sources, theory, values and application domains , pages =. 2015 , publisher =
2015
-
[44]
2019 , publisher =
Power without knowledge: a critique of technocracy , author =. 2019 , publisher =
2019
-
[45]
Atul Gawande , publisher =
-
[46]
Communications of the ACM , volume =
Datasheets for datasets , author =. Communications of the ACM , volume =. 2021 , publisher =
2021
-
[47]
International Journal of Communication , author =
Making Algorithms Public: Reimagining Auditing from Matters of Fact to Matters of Concern , volume =. International Journal of Communication , author =. 2024 , pages =
2024
-
[48]
Berkeley Tech
Through the handoff lens: competing visions of autonomous futures , author =. Berkeley Tech. LJ , volume =. 2020 , publisher =
2020
-
[49]
To Be High-Risk, or Not To Be—Semantic Specifications and Implications of the
Golpayegani, Delaram and Pandit, Harshvardhan J and Lewis, Dave , booktitle =. To Be High-Risk, or Not To Be—Semantic Specifications and Implications of the. 2023 , doi =
2023
-
[50]
2021 , publisher =
The dawn of everything: A new history of humanity , author =. 2021 , publisher =
2021
-
[51]
Safety science , volume =
The nature of safety culture: a review of theory and research , author =. Safety science , volume =. 2000 , publisher =
2000
-
[52]
Law & policy , volume =
Industry self-regulation: an institutional perspective , author =. Law & policy , volume =. 1997 , publisher =
1997
-
[53]
Proceedings Fourth Safety Congress , pages =
Standardization of safeguards , author =. Proceedings Fourth Safety Congress , pages =. 1915 , doi =
1915
-
[54]
STAT News , author =
-
[55]
Benjamin, Ruha , date =
-
[56]
2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages =
Enhanced well-being assessment as basis for the practical implementation of ethical and rights-based normative principles for AI , author =. 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages =. 2020 , organization =
2020
-
[57]
Proceedings of the ACM on Human-Computer Interaction , volume =
Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and Desiderata , author =. Proceedings of the ACM on Human-Computer Interaction , volume =. 2022 , publisher =
2022
-
[58]
2021 , doi=
Hendrycks, Dan and Carlini, Nicholas and Schulman, John and Steinhardt, Jacob , journal =. 2021 , doi=
2021
-
[59]
2020 , publisher =
Herkert, Joseph and Borenstein, Jason and Miller, Keith , journal =. 2020 , publisher =
2020
-
[60]
Ethics and Information Technology , volume =
ChatGPT is bullshit , author =. Ethics and Information Technology , volume =. 2024 , publisher =
2024
-
[61]
Jama , volume =
Deep learning—a technology with the potential to transform health care , author =. Jama , volume =. 2018 , publisher =
2018
-
[62]
Accountability: Power, ethos and the technologies of managing , volume =
The ‘awful idea of accountability’: inscribing people into the measurement of objects , author =. Accountability: Power, ethos and the technologies of managing , volume =. 1996 , publisher =
1996
-
[63]
Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure , author =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , doi =
2021
-
[64]
2022 , journal =
Artificial Intelligence: Too Fragile to Fight? , author =. 2022 , journal =
2022
-
[65]
Concrete Safety for
Jatho, Edgar W and Mailloux, Logan O and Williams, Eugene D and McClure, Patrick and Kroll, Joshua A , journal =. Concrete Safety for. 2023 , doi =
2023
-
[66]
Journal of business ethics , volume=
The social license to operate , author=. Journal of business ethics , volume=. 2016 , publisher=
2016
-
[67]
Toward Comprehensive Risk Assessments and Assurance of
Khlaaf, Heidy , institution =. Toward Comprehensive Risk Assessments and Assurance of. 2023 , url =
2023
-
[68]
arXiv preprint arXiv:2506.11135 , year =
Large Language Models and Emergence: A Complex Systems Perspective , author =. arXiv preprint arXiv:2506.11135 , year =
-
[69]
Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences , volume =
The fallacy of inscrutability , author =. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences , volume =. 2018 , publisher =
2018
-
[70]
2020 , publisher =
Accountability in computer systems , author =. 2020 , publisher =
2020
-
[71]
2021 Conference on Fairness, Accountability, and Transparency , pages =
Outlining Traceability: A Principle for Operationalizing Accountability in Computing Systems , author =. 2021 Conference on Fairness, Accountability, and Transparency , pages =. 2021 , doi =
2021
-
[72]
Responsible
Kroll, Joshua A , journal =. Responsible
-
[73]
2022 , institution =
Understanding, Assessing, and Mitigating Safety Risks in Artificial Intelligence Systems , author =. 2022 , institution =
2022
-
[74]
Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =
The Misuse of AUC: What High Impact Risk Assessment Gets Wrong , author =. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages =
2023
-
[75]
Where are the missing masses? The sociology of a few mundane artifacts , shorttitle =
Latour, Bruno , editor =. Where are the missing masses? The sociology of a few mundane artifacts , shorttitle =. Shaping Technology/Building Society: Studies in Sociotechnical Change , publisher =
-
[76]
ACM Computing Surveys (CSUR) , volume =
A survey of DevOps concepts and challenges , author =. ACM Computing Surveys (CSUR) , volume =. 2019 , publisher =
2019
-
[77]
An investigation of the
Leveson, Nancy and Turner, Clark S , journal =. An investigation of the. 1993 , publisher =
1993
-
[78]
Safety science , volume =
A new accident model for engineering safer systems , author =. Safety science , volume =. 2004 , publisher =
2004
-
[79]
Journal of spacecraft and Rockets , volume =
Role of software in spacecraft accidents , author =. Journal of spacecraft and Rockets , volume =. 2004 , doi=
2004
-
[80]
Organization studies , volume =
Moving beyond normal accidents and high reliability organizations: A systems approach to safety in complex systems , author =. Organization studies , volume =. 2009 , publisher =
2009
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.