REVIEW 3 major objections 6 minor 179 references
AI Safety for Everyone
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Across 383 peer-reviewed papers, AI safety is mostly concrete robustness work, not existential-risk theory.
desk verdict A useful empirical map of peer-reviewed AI safety work, but the relative-frequency claim about existential risk is partly an artifact of the search query. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the systematic literature review pipeline: a keyword hierarchy (AI terms AND safe*/robust*/align* AND lifecycle terms) run on two major indexing databases, narrowed to 383 papers by title/abstract/full-text screening, enriched by inductive coding and snowball sampling. From that corpus the authors derive their eight risk types and ten methodology categories. The empirical distribution of those categories—not any single theorem—does the argumentative work: it shows near-term technical safety concerns dominate and that they map onto traditional systems-safety concerns such as reliability, control, and adversarial resilience.
What would settle it
Re-run the same three-pronged keyword query (AI AND safe*/robust*/align* AND lifecycle terms) on a broad preprint server plus the main online alignment forums for the same period and compare the risk-type shares. If papers on existential risk, superintelligence, and AGI confinement outnumber any single near-term category—or if the combined non-peer-reviewed share overturns the 23.2% figure for noise and outliers—the review's central empirical claim fails.
Extended reading notes
Core claim
The paper's central claim is that the empirical record of peer-reviewed AI safety research undercuts the narrative that AI safety is primarily about existential risk from advanced AI. Reviewing 2,666 database hits plus snowballed papers down to 383, the authors identify eight families of safety risk and find the largest share concerns noise and outliers (23.2%), followed by lack of monitoring, system misspecification, and lack of control enforcement, with explicitly existential or AGI-confinement work accounting for a small fraction of the literature. They further show the field's methods—applied algorithms, safe-RL agent simulations, analysis frameworks, mechanistic interpretability, and design frameworks—mirror the practices of reliability engineering, control theory, and systems safety. On this evidence they argue AI safety should be understood pluralistically, as continuous with the safety engineering tradition rather than as a separate enterprise fixated on extinction scenarios.
Load-bearing premise
The whole picture rests on the assumption that the peer-reviewed, index-database corpus queried with the words safe, robust, or align fairly represents what 'AI safety' research is; because most existential-risk and alignment work appears on preprints and non-peer-reviewed forums, the measured small share of existential risk could be an artifact of where the authors looked.
Editorial extensions
If this is right
- If the review's distribution is representative, regulators and funders should treat AI safety as a broad engineering discipline, not only as existential-risk mitigation.
- The shared technical vocabulary across time horizons, e.g. corrigibility and adversarial robustness in reinforcement learning, supports integrated safety research rather than a near-term/long-term dichotomy.
- Reconnecting AI safety to systems-safety practice in aviation, medicine, and nuclear power gives policymakers a ready stock of certification and assurance frameworks to adapt.
- A pluralistic definition of AI safety can broaden participation from researchers who do not subscribe to existential-risk narratives, without eliminating existential risk as a research topic.
- The authors stress that their findings do not negate the importance of existential risk, only its status as the defining lens.
Reading between the lines
- A direct test of the paper's scope claim would be to run the same query hierarchy on a broad preprint server plus the main online alignment forums and recalculate the risk-type shares; if existential-risk and AGI-confinement papers then outnumber every near-term category, the quantitative 'sliver' claim would be an artifact of the peer-reviewed-only sample, a limitation the paper itself flags.
- The governance consequence the authors leave implicit: if AI safety is continuous with systems safety, then standards bodies with existing safety-case patterns for machine-learning components have a head start on AI regulation, and 'AI safety certification' becomes more tractable than under an existential-risk framing.
- The four-cluster map suggests a research hypothesis: adversarial robustness and safe exploration may be two sides of the same formal problem (worst-case perturbations versus constraint violations), and a unification would strengthen the paper's continuity claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a systematic literature review of peer-reviewed AI safety research indexed in Web of Science and Scopus, using a multi-stage query that combines AI/ML terms with the safety-related keywords safe*, robust*, and align*, and adds lifecycle terms. From 383 selected papers, the authors inductively code eight risk types and ten methodology categories, reporting that near-term concerns such as robustness to noise and outliers dominate while existential risk appears in only a small minority of papers. On this basis, the paper argues that AI safety is a pluralistic field naturally connected to traditional technological and systems safety, and that the public narrative focusing on existential risk is too narrow. The authors propose a more inclusive framing and outline future research directions, including expanding the review to non-peer-reviewed sources.
Significance. If the empirical findings are accepted, the paper would make a useful contribution by providing a structured, reproducible map of peer-reviewed AI safety research and by grounding debates about the field's scope in data rather than anecdote. The authors are commendably explicit about their methodology and state their intention to release the annotated corpus and analysis code. However, the force of the central claim depends on whether the corpus is representative of AI safety as a field; the current manuscript does not yet demonstrate this, because the selection criteria and the borrowed risk taxonomy partly determine the observed distribution. The pluralism thesis is plausible and important, but the evidence as presented is provisional.
major comments (3)
- [§2.1, §4, §6] The central relative-frequency claim in §6—that AI safety research 'challenges the narrative that associates AI safety primarily with mitigating existential risks'—is computed from the 383-paper corpus selected by q3 = q1 ∧ (safe* ∨ robust* ∨ align*) and q4. Because q2 includes the broad wildcard 'robust*', the denominator is inflated by mainstream machine-learning, control, and reliability-engineering papers on noise, outliers, and adversarial perturbations that may not self-identify as AI safety research. The authors apply exclusion criterion (B) (motivation), but they do not report how many papers were excluded at each stage, nor do they provide a baseline comparison against general ML papers to show that the risk distribution in Figure 3 is specific to AI safety rather than an artifact of the robust* term. Without a sensitivity analysis (for example, re-running the review with q2 restricted to 'safe*' or to the phrase 'AI safety', or with additional terms), the claim that 'noise and outliers' is the dominant risk type (23.2%) and the resulting argument in §6 are not fully supported.
- [§2.1; Introduction, Limitations paragraph] The search protocol also systematically underrepresents existential-risk and long-term alignment research. Much of this work is disseminated on arXiv, LessWrong, and the AI Alignment Forum, and it often uses vocabulary such as 'existential risk', 'x-risk', 'superintelligence', or 'misalignment' without necessarily co-occurring with safe*, robust*, or align* in WoS/Scopus-indexed metadata. The Limitations paragraph concedes the exclusion of preprints and forums, but this concession is not quantified, and the paper's conclusion in §6 is not correspondingly hedged. A supplemental search of arXiv (and a citation-based snowball from known x-risk seeds) with the resulting papers coded into the same taxonomy would allow the authors to bound the effect of this omission and would materially strengthen the central claim.
- [§4, Fig. 3] The risk taxonomy is borrowed from DeepMind's blog post [67], and several categories overlap with the search terms themselves. 'Noise and outliers' and 'Lack of monitoring' are natural homes for papers captured by robust* and safe*, so the relative frequencies in Figure 3 partially reflect the query design. The authors should discuss this potential circularity and, ideally, re-code a random subset of papers with a taxonomy not derived from the safety-robustness-alignment vocabulary to verify that the qualitative ordering of risk types is stable. At minimum, the text should acknowledge that the taxonomy choice and the keyword filter are not independent.
minor comments (6)
- [Fig. 1 caption] The caption reads 'World cloud of morphologically standardised terms'; 'World' should be 'Word'.
- [§6] The text contains two typos: 'episemically-inclusive' should be 'epistemically inclusive', and 'demistify' should be 'demystify'.
- [§2.1, q4d] The wildcard 'decomission*' is misspelled; the intended term is 'decommission*'.
- [§5] 'Generalisaion' should be 'generalization' in the sentence describing domain-generalisation methods.
- [§2.1] The query is described in WoS notation only; the authors should state how the syntax was adapted for Scopus, including any differences in wildcard handling and field tags.
- [Limitations] The snowball sampling description does not report how many of the 117 snowballed papers were included in the final 383; reporting this overlap would help readers assess the contribution of the snowballing step.
Circularity Check
No significant circularity: the corpus is operationalized rather than derived, and the field-breadth conclusion is an empirical generalization supported by independent paper-level screening; minor self-citations are not load-bearing.
full rationale
This paper is a systematic literature review, not a formal derivation, so equation-level circularity does not directly apply. The corpus is defined in Section 2.1 by q3 = q1 ∧ q2, where q2 requires safe*, robust*, or align* in WoS/Scopus metadata. One might worry that the prominence of 'noise and outliers' (23.2%) and 'adversarial attacks' categories in Figure 3 is an artifact of including robust* as an inclusion keyword. However, the final 383-paper set was not produced by q3 alone: the authors applied exclusion criteria (A)-(C) requiring a primary focus on AI, a stated motivation for safe AI algorithms, and generalizability, and then manually annotated each full text. The risk-type frequencies are therefore a measured property of the screened corpus, not a logically forced consequence of the keyword filter. The risk taxonomy is borrowed from DeepMind [67], an external source, and applying it is standard deductive coding; the authors concede annotator bias in the Limitations section. The paper also explicitly concedes that excluding arXiv, LessWrong, and the AI Alignment Forum may underrepresent existential-risk work; that is a generalizability threat, not a circularity, because the paper's central claim is explicitly about 'primarily peer-reviewed research' as indexed by WoS/Scopus, and the existence of the 383 concrete safety papers is an independent empirical fact. Self-citations ([1], [16], [78], [148]) are used as illustrative examples of recent discourse, explainability research, and mechanistic interpretability; none is load-bearing, and no uniqueness theorem or prior-work-only premise is invoked to force the conclusions. The finding that AI safety research extends traditional technological safety is an interpretive synthesis of the annotated corpus, not a renaming of the query inputs. Overall, the paper shows minor self-citation but no circular step that reduces its central empirical claim to its own construction.
Assumptions & free parameters
assumptions (5)
- domain assumption The WoS and Scopus peer-reviewed corpus plus snowball sampling is representative of AI safety research as a whole.
- ad hoc to paper The keyword filter safe*, robust*, align* captures the relevant AI safety literature.
- domain assumption The eight risk categories adapted from DeepMind's taxonomy are a valid coding scheme.
- domain assumption Subjective title, abstract, and full-text screening can be applied consistently.
- domain assumption A more inclusive and pluralistic conception of AI safety is normatively desirable.
Cite this review
Pith. "Pith review of AI Safety for Everyone." pith.science (2026). https://pith.science/paper/YOOEDBHE
@misc{pith2026250209288,
author = {Pith},
title = {Pith review of: AI Safety for Everyone},
year = {2026},
howpublished = {\url{https://pith.science/paper/YOOEDBHE}},
note = {Machine review of arXiv:2502.09288}
}
read the original abstract
Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of potential existential threats. However, this framing has three potential drawbacks: it may exclude researchers and practitioners who are committed to AI safety but approach the field from different angles; it could lead the public to mistakenly view AI safety as focused solely on existential scenarios rather than addressing a wide spectrum of safety challenges; and it risks creating resistance to safety measures among those who disagree with predictions of existential AI risks. Through a systematic literature review of primarily peer-reviewed research, we find a vast array of concrete safety work that addresses immediate and practical concerns with current AI systems. This includes crucial areas like adversarial robustness and interpretability, highlighting how AI safety research naturally extends existing technological and systems safety concerns and practices. Our findings suggest the need for an epistemically inclusive and pluralistic conception of AI safety that can accommodate the full range of safety considerations, motivations, and perspectives that currently shape the field.
Reference graph
Works this paper leans on
-
[67]
A., Maini, V
Ortega, P. A., Maini, V. & the DeepMind safety team. Building safe artificial intel- ligence: specification, robustness, and assurance. https://deepmindsafetyresearch. medium.com/building-safe-artificial-intelligence-52f5f75058f1 (2018). Accessed: 2024-05-13
2018
-
[1]
Two types of AI existential risk: Decisive and accumulative
Kasirzadeh, A. Two types of AI existential risk: Decisive and accumulative. arXiv preprint arXiv:2401.07836 (2024)
arXiv 2024
-
[2]
& Nelson, A
Lazar, S. & Nelson, A. AI safety on whose terms? Science 381, 138–138 (2023)
2023
-
[3]
& Wang, M
Ahmed, S., Ja´ zwi´ nska, K., Ahlawat, A., Winecoff, A. & Wang, M. Building the epistemic community of AI safety. First Monday (forthcoming). URL https://papers.ssrn.com/sol3/papers.cfm?abstract id=4641526
-
[4]
Existential risks: Analyzing human extinction scenarios and related hazards
Bostrom, N. Existential risks: Analyzing human extinction scenarios and related hazards. Journal of Evolution and Technology 9, 1–30 (2002)
2002
-
[5]
Superintelligence: Paths, dangers, strategies (Oxford University Press, 2014)
Bostrom, N. Superintelligence: Paths, dangers, strategies (Oxford University Press, 2014)
2014
-
[6]
The precipice: Existential risk and the future of humanity (Hachette Books, 2020)
Ord, T. The precipice: Existential risk and the future of humanity (Hachette Books, 2020)
2020
-
[7]
The logic of effective altruism
Boston Review. The logic of effective altruism. Boston Review (2015). URL https://www.bostonreview.net/forum/peter-singer-logic-effective-altruism/
2015
Show all 179 references
-
[8]
Effective altruism and its critics
Gabriel, I. Effective altruism and its critics. Journal of Applied Philosophy 34, 457–473 (2017)
2017
-
[9]
AI is an existential threat—just not the way you think
Eisikovits, N. AI is an existential threat—just not the way you think. https://www.scientificamerican.com/article/ ai-is-an-existential-threat-just-not-the-way-you-think/ (2023). Accessed: July 12, 2023
2023
-
[10]
Statement on ai risk
Center for AI Safety. Statement on ai risk. https://www.safe.ai/work/ statement-on-ai-risk (2023). Accessed on 2024-05-08
2023
-
[11]
Training, A. S. Ai safety training. https://aisafety.training/ (2024). Accessed on 2024-05-08
2024
-
[12]
AI safety — wikipedia, the free encyclopedia (2024)
Wikipedia contributors. AI safety — wikipedia, the free encyclopedia (2024). URL https://en.wikipedia.org/wiki/AI safety. [Online; accessed 21-January-2024]
2024
-
[13]
& Norvig, P
Ag¨ uera y Arcas, B. & Norvig, P. Artificial general intelligence is already here. Noema Magazine (2023). Available at: https://www.noemamag.com/ artificial-general-intelligence-is-already-here/. 17
2023
-
[14]
Roose, K. A.I. Poses ‘Risk of Extinction’, Industry Leaders Warn. The New York Times (2023). URL https://www.nytimes.com/2023/05/30/technology/ ai-threat-warning.html
2023
-
[15]
PM should make ethics a priority at AI safety summit, say tech professionals (2023)
BCS Comment. PM should make ethics a priority at AI safety summit, say tech professionals (2023). URL https://www.bcs.org/articles-opinion-and-research/ pm-should-make-ethics-a-priority-at-ai-safety-summit-say-tech-professionals/. Accessed: 27 January, 2024
2023
-
[16]
& Gohdes, A
Gilardi, F., Kasirzadeh, A., Bernstein, A., Staab, S. & Gohdes, A. We need to understand the effect of narratives about generative ai. Nature Human Behaviour 1–2 (2024)
2024
-
[17]
Bender, E. M. Talking about a ‘schism’ is ahistorical (2023). URL https://medium. com/@emilymenonbender/talking-about-a-schism-is-ahistorical-3c454a77220f
2023
-
[18]
Krause, S. S. Aircraft safety (McGraw-Hill Professional Publishing, 2003)
2003
-
[19]
Boyd, D. D. A review of general aviation safety (1984–2017). Aerospace medicine and human performance 88, 657–664 (2017)
2017
-
[20]
& Restani, P
Pifferi, G. & Restani, P. The safety of pharmaceutical excipients. Il Farmaco 58, 541–550 (2003)
2003
-
[21]
Leveson, N. et al. Applying system engineering to pharmaceutical safety. Journal of Healthcare Engineering 3, 391–414 (2012)
2012
-
[22]
& Van Ouytsel, J
De Kimpe, L., Walrave, M., Ponnet, K. & Van Ouytsel, J. Internet safety. The international encyclopedia of media literacy 1–11 (2019)
2019
-
[23]
Salim, H. M. Cyber safety: A systems thinking and systems theory approach to managing cyber security risks. Ph.D. thesis, Massachusetts Institute of Technology (2014)
2014
-
[24]
Leveson, N. G. Engineering a safer world: Systems thinking applied to safety (The MIT Press, 2016)
2016
-
[25]
Varshney, K. R. Engineering safety in machine learning , 1–5 (IEEE, 2016)
2016
-
[26]
Rismani, S. et al. From plane crashes to algorithmic harm: applicability of safety engineering frameworks for responsible ml , 1–18 (2023)
2023
-
[27]
System safety and artificial intelligence , 1584–1584 (2022)
Dobbe, R. System safety and artificial intelligence , 1584–1584 (2022)
2022
-
[28]
Rismani, S. et al. Beyond the ml model: Applying safety engineering frameworks to text-to-image development , 70–83 (2023). 18
2023
-
[29]
Amodei, D. et al. Concrete problems in AI safety.arXiv preprint arXiv:1606.06565 (2016)
2016 arXiv
-
[30]
Raji, I. D. & Dobbe, R. Concrete problems in AI safety, revisited. arXiv preprint arXiv:2401.10899 (2023)
2023 arXiv
-
[31]
& Charters, S
Kitchenham, B. & Charters, S. Guidelines for performing Systematic Literature Reviews in Software Engineering. EBSE Technical Report EBSE-2007-01, School of Computer Science and Mathematics, Keele University, Keele, UK (2007)
2007
-
[32]
Wohlin, C. Guidelines for snowballing in systematic literature studies and a replication in software engineering , EASE ’14, 1–10 (Association for Computing Machinery, New York, NY, USA, 2014)
2014
-
[33]
& Amodei, D
Irving, G., Christiano, P. & Amodei, D. AI safety via debate. arXiv (2018)
2018
-
[34]
Ng, A. Y. & Russell, S. J. Algorithms for Inverse Reinforcement Learning, ICML ’00, 663–670 (Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2000)
2000
-
[35]
Hendrycks, D. et al. Aligning AI With Shared Human Values (ICLR, 2021). 2008.02275
2021 arXiv
-
[36]
Yampolskiy, R. V. Artificial Intelligence Safety and Cybersecurity: A Timeline of AI Failures. arXiv (2016)
2016
-
[37]
J., Abbeel, P
Hadfield-Menell, D., Russell, S. J., Abbeel, P. & Dragan, A. Cooperative Inverse Reinforcement Learning, Vol. 29 (Curran Associates, Inc., 2016)
2016
-
[38]
Xu, H., Zhu, T., Zhang, L., Zhou, W. & Yu, P. S. Machine Unlearning: A Survey. ACM Computing Surveys 56, 9:1–9:36 (2023)
2023
-
[39]
& Tegmark, M
Russell, S., Dewey, D. & Tegmark, M. Research Priorities for Robust and Beneficial Artificial Intelligence. AI magazine 36, 105–114 (2015)
2015
-
[40]
& Abrecht, S
Willers, O., Sudholt, S., Raafatnia, S. & Abrecht, S. Casimiro, A., Ortmeier, F., Schoitsch, E., Bitsch, F. & Ferreira, P. (eds) Safety Concerns and Mitigation Approaches Regarding the Use of Deep Learning in Safety-Critical Perception Tasks. (eds Casimiro, A., Ortmeier, F., S...
2020
-
[41]
Mohseni, S. et al. Taxonomy of Machine Learning Safety: A Survey and Primer. ACM Computing Surveys 55, 1–38 (2022)
2022
-
[42]
& Steinhardt, J
Hendrycks, D., Carlini, N., Schulman, J. & Steinhardt, J. Unsolved Problems in ML Safety. arXiv (2022). 19
2022
-
[43]
Boyatzis, R. E. Transforming qualitative information: Thematic analysis and code development (sage, 1998)
1998
-
[44]
van Eck, N. J. & Waltman, L. Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics 84, 523–538 (2010)
2010
-
[45]
V., Strong, J
Oster Jr, C. V., Strong, J. S. & Zorn, C. K. Analyzing aviation safety: Problems, challenges, opportunities. Research in transportation economics 43, 148–164 (2013)
2013
-
[46]
S., Corrigan, J
Donaldson, M. S., Corrigan, J. M. & Kohn, L. T. To err is human: building a safer health system (2000)
2000
-
[47]
Bates, D. W. et al. The safety of inpatient health care. New England Journal of Medicine 388, 142–153 (2023)
2023
-
[48]
Marais, K., Dulac, N., Leveson, N. et al. Beyond normal accidents and high reliability organizations: The need for an alternative approach to safety in complex systems, 1–16 (Citeseer, 2004)
2004
-
[49]
Griffor, E. Handbook of system safety and security: cyber risk and risk manage- ment, cyber security, threat analysis, functional safety, software systems, and cyber physical systems (Syngress, 2016)
2016
-
[50]
& Rohokale, V
Prasad, R. & Rohokale, V. Cyber security: the lifeline of information and communication technology (Springer, 2020)
2020
-
[51]
& Armstrong, S
Soares, N., Fallenstein, B., Yudkowsky, E. & Armstrong, S. Corrigibility (Austin, Texas, USA, 2015)
2015
-
[52]
& Orseau, L
Ring, M. & Orseau, L. Schmidhuber, J., Th´ orisson, K. R. & Looks, M. (eds) Delusion, Survival, and Intelligent Agents . (eds Schmidhuber, J., Th´ orisson, K. R. & Looks, M.) Artificial General Intelligence , Lecture Notes in Computer Science, 11–20 (Springer, Berlin, Heidelbe...
2011
-
[53]
Ganin, Y. et al. in Domain-Adversarial Training of Neural Networks (ed.Csurka, G.) Domain Adaptation in Computer Vision Applications 189–209 (Springer International Publishing, Cham, 2017)
2017
-
[54]
& Chellappa, R
Balaji, Y., Sankaranarayanan, S. & Chellappa, R. MetaReg: Towards Domain Generalization using Meta-Regularization, Vol. 31 (Curran Associates, Inc., 2018)
2018
-
[55]
J., Shlens, J
Goodfellow, I. J., Shlens, J. & Szegedy, C. Explaining and Harnessing Adversarial Examples (2015). 1412.6572
2015 arXiv
-
[56]
& Vladu, A
Madry, A., Makelov, A., Schmidt, L., Tsipras, D. & Vladu, A. Towards Deep Learning Models Resistant to Adversarial Attacks. arXiv (2019). 20
2019
-
[57]
& Amodei, D
Ray, A., Achiam, J. & Amodei, D. Benchmarking Safe Exploration in Deep Reinforcement Learning. OpenAI (2023). Unpublished article
2023
-
[58]
& Krause, A
Turchetta, M., Berkenkamp, F. & Krause, A. Wallach, H. et al. (eds) Safe Exploration for Interactive Machine Learning . (eds Wallach, H. et al. ) Advances in Neural Information Processing Systems 32 (NIPS 2019) , Vol. 32 (2019)
2019
-
[59]
Ouyang, L. et al. Training language models to follow instructions with human feedback. arXiv (2022). OpenAI
2022
-
[60]
Abbeel, P. & Ng, A. Y. Apprenticeship learning via inverse reinforcement learning, ICML ’04, 1 (Association for Computing Machinery, New York, NY, USA, 2004)
2004
-
[61]
A., Chaudhuri, K
Izzo, Z., Smart, M. A., Chaudhuri, K. & Zou, J. Approximate Data Deletion from Machine Learning Models , 2008–2016 (PMLR, 2021)
2008
-
[62]
& Davidson, S
Wu, Y., Dobriban, E. & Davidson, S. DeltaGrad: Rapid retraining of machine learning models, 10355–10366 (PMLR, 2020)
2020
-
[63]
Adebayo, J. et al. Sanity Checks for Saliency Maps , Vol. 31 (Curran Associates, Inc., 2018)
2018
-
[64]
Kim, B. et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCA V), 2668–2677 (PMLR, 2018)
2018
-
[65]
& Gimpel, K
Hendrycks, D. & Gimpel, K. A Baseline for Detecting Misclassified and Out-of- Distribution Examples in Neural Networks (arXiv, 2018). 1610.02136
2018 arXiv
-
[66]
& Dietterich, T
Hendrycks, D. & Dietterich, T. Benchmarking Neural Network Robustness to Common Corruptions and Perturbations (arXiv, 2019). 1903.12261
2019 arXiv
-
[68]
Meng, Y. et al. Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-Training , 10367–10378 (2021)
2021
-
[69]
& Guo, P
Wang, K. & Guo, P. A Robust Automated Machine Learning System with Pseudoinverse Learning. Cognitive Computation 13, 724–735 (2021)
2021
-
[70]
& Murphy, T
Cappozzo, A., Greselin, F. & Murphy, T. B. A robust approach to model-based classification based on trimming and constraints: Semi-supervised learning in presence of outliers and label noise. Advances in Data Analysis and Classification 14, 327–354 (2020). 21
2020
-
[71]
& Wang, Y
Li, W. & Wang, Y. A robust supervised subspace learning approach for output- relevant prediction and detection against outliers. Journal of Process Control 106, 184–194 (2021)
2021
-
[72]
& Krause, A
Curi, S., Bogunovic, I. & Krause, A. Combining Pessimism with Optimism for Robust and Efficient Model-Based Deep Reinforcement Learning , Vol. 139, 2254–2264 (2021)
2021
-
[73]
& Mintz, Y
Dobbe, R., Gilbert, TK. & Mintz, Y. Hard choices in artificial intelligence. Artificial Intelligence 300 (2021)
2021
-
[74]
& Feldman, V
Dwork, C. & Feldman, V. Privacy-preserving Prediction, 1693–1702 (PMLR, 2018)
2018
-
[75]
Elhage, N. et al. A mathematical framework for transformer circuits. Transformer Circuits Thread (2021). Unpublished article
2021
-
[76]
& Mnih, A
Kim, H. & Mnih, A. Disentangling by Factorising. arXiv (2019)
2019
-
[77]
& Habli, I
Ward, F. & Habli, I. An Assurance Case Pattern for the Interpretability of Machine Learning in Safety-Critical Systems , Vol. 12235 LNCS, 395–407 (2020)
2020
-
[78]
& Schafer, B
Gyevnar, B., Ferguson, N. & Schafer, B. Bridging the transparency gap: What can explainable ai learn from the ai act? , 964–971 (IOS Press, 2023)
2023
-
[79]
& Kniesel-W¨ unsche, G.Safe-DS: A Domain Specific Language to Make Data Science Safe , 72–77 (2023)
Reimann, L. & Kniesel-W¨ unsche, G.Safe-DS: A Domain Specific Language to Make Data Science Safe , 72–77 (2023)
2023
-
[80]
& Lee, S.-W
Dey, S. & Lee, S.-W. A Multi-layered Collaborative Framework for Evidence- driven Data Requirements Engineering for Machine Learning-based Safety-critical Systems, 1404–1413 (2023)
2023
-
[81]
& Zimmert, J
Wei, C.-Y., Dann, C. & Zimmert, J. A Model Selection Approach for Corruption Robust Reinforcement Learning, Vol. 167, 1043–1096 (2022)
2022
-
[82]
& Singla, A
Ghosh, A., Tschiatschek, S., Mahdavi, H. & Singla, A. Towards Deployment of Robust Cooperative AI Agents: An Algorithmic Framework for Learning Adaptive Policies. New Zealand (2020)
2020
-
[83]
& Hutter, M
Everitt, T. & Hutter, M. Avoiding Wireheading with Value Reinforcement Learning (2016). 1605.03143
2016 arXiv
-
[84]
& Garrabrant, S
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J. & Garrabrant, S. Risks from learned optimization in advanced machine learning systems. CoRR abs/1906.01820 (2019). URL http://arxiv.org/abs/1906.01820
2019 arXiv
-
[85]
& Yampolskiy, R
Pistono, F. & Yampolskiy, R. V. The Age of Artificial Intelligence , Ch. Unethical Research: How to Create a Malevolent Artificial Intelligence (Vernon Press 22 Wilmington, DE, USA, 2016)
2016
-
[86]
& Habli, I
Picardi, C., Paterson, C., Hawkins, R., Calinescu, R. & Habli, I. Assurance argument patterns and processes for machine learning in safety-related systems , Vol. 2560, 23–30 (2020)
2020
-
[87]
& Zeilinger, MN
Wabersich, KJ., Hewing, L., Carron, A. & Zeilinger, MN. Probabilistic Model Predictive Safety Certification for Learning-Based Control. IEEE Transactions on Automatic Control 67, 176–188 (2022)
2022
-
[88]
& Topcu, U
Wen, M. & Topcu, U. Constrained Cross-Entropy Method for Safe Reinforcement Learning. IEEE Transactions on Automatic Control 66, 3123–3137 (2021)
2021
-
[89]
Analyzing Information Leakage of Updates to Natural Language Models, 363–375 (2020)
Zanella-B´ eguelin, S.et al. Analyzing Information Leakage of Updates to Natural Language Models, 363–375 (2020). 1912.07942
2020 arXiv
-
[90]
& Dong, D
Wang, Z., Chen, C. & Dong, D. A Dirichlet Process Mixture of Robust Task Models for Scalable Lifelong Reinforcement Learning. IEEE Transactions on Cybernetics 1–12 (2022)
2022
-
[91]
Zou, A. et al. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv (2023)
2023
-
[92]
Ilyas, A. et al. Adversarial Examples Are Not Bugs, They Are Features , Vol. 32 (Curran Associates, Inc., 2019)
2019
-
[93]
& Yin, YL
He, RD., Han, ZY., Yang, Y. & Yin, YL. Not All Parameters Should Be Treated Equally: Deep Safe Semi-supervised Learning under Class Distribution Mismatch , 6874–6883 (2022)
2022
-
[94]
& Vigna, G
Aghakhani, H., Meng, D., Wang, Y.-X., Kruegel, C. & Vigna, G. Bullseye polytope: A scalable clean-label poisoning attack with improved transferability , 159–178 (2021)
2021
-
[95]
Liu, Y. et al. Backdoor Defense with Machine Unlearning , 280–289 (IEEE Press, London, United Kingdom, 2022)
2022
-
[96]
& Hein, M
Meinke, A. & Hein, M. Towards neural networks that provably know when they don’t know (2020). 1909.12180
2020 arXiv
-
[97]
Abdelfattah, S., Kasmarik, K. & Hu, J. A robust policy bootstrapping algo- rithm for multi-objective reinforcement learning in non-stationary environments. Adaptive Behavior 28, 273–292 (2020)
2020
-
[98]
& Topcu, U
Djeumou, F., Cubuktepe, M., Lennon, C. & Topcu, U. Task-Guided Inverse Reinforcement Learning Under Partial Information (2021). 2105.14073. 23
2021 arXiv
-
[99]
Yampolskiy, R. V. Leakproofing the Singularity: Artificial intelligence confinement problem. Journal of Consciousness Studies 19, 194–214 (2012)
2012
-
[100]
Yampolskiy, R. V. in Artificial Intelligence Safety Engineering: Why Machine Ethics Is a Wrong Approach (ed.M¨ uller, V. C.)Philosophy and Theory of Artificial Intelligence Studies in Applied Philosophy, Epistemology and Rational Ethics, 389–396 (Springer, Berlin, Heidelberg, 2013)
2013
-
[101]
Yampolskiy, R. V. Taxonomy of Pathways to Dangerous AI (2015). 1511.03246
2015 arXiv
-
[102]
Wheatley, S., Sovacool, B. K. & Sornette, D. Reassessing the safety of nuclear power. Energy Research & Social Science 15, 96–100 (2016)
2016
-
[103]
& Ho, WK
Tay, EB., Gan, OP. & Ho, WK. Rauch, HE. (ed.) A study on real-time artificial intelligence. (ed.Rauch, HE.) Artificial Intelligence in Real-Time Control 1997 , 109–114 (1998)
1998
-
[104]
Bach, J., Goertzel, B
Hibbard, B. Bach, J., Goertzel, B. & Ikl´ e, M. (eds) Avoiding Unintended AI Behaviors. (eds Bach, J., Goertzel, B. & Ikl´ e, M.) Artificial General Intelligence , Lecture Notes in Computer Science, 107–116 (Springer, Berlin, Heidelberg, 2012)
2012
-
[105]
Bieger, J., Goertzel, B
Sezener, CE. Bieger, J., Goertzel, B. & Potapov, A. (eds) Inferring Human Values for Safe AGI Design. (eds Bieger, J., Goertzel, B. & Potapov, A.) Artificial General Intelligence (AGI 2015) , Vol. 9205, 152–155 (2015)
2015
-
[106]
& Guerraoui, R
El Mhamdi, E. & Guerraoui, R. When Neurons Fail, 1028–1037 (2017)
2017
-
[107]
& Grote, T
Freiesleben, T. & Grote, T. Beyond generalization: A theory of robustness in machine learning. Synthese 202 (2023)
2023
-
[108]
& Shah, J
Sanneman, L. & Shah, J. Transparent Value Alignment , 557–560 (ACM, Stockholm Sweden, 2023)
2023
-
[109]
& Mitchell, T
Murphy, B., Talukdar, P. & Mitchell, T. Kay, M. & Boitet, C. (eds) Learning Effective and Interpretable Semantic Models using Non-Negative Sparse Embed- ding. (eds Kay, M. & Boitet, C.) Proceedings of COLING 2012, 1933–1950 (The COLING 2012 Organizing Committee, Mumbai, India, 2012)
2012
-
[110]
& Hovy, E
Subramanian, A., Pruthi, D., Jhamtani, H., Berg-Kirkpatrick, T. & Hovy, E. SPINE: SParse Interpretable Neural Embeddings (2017). 1711.08792
2017 arXiv
-
[111]
& Negahban, S
Shaham, U., Yamada, Y. & Negahban, S. Understanding adversarial training: Increasing local stability of supervised models through robust optimization. Neurocomputing 307, 195–204 (2018)
2018
-
[112]
& Srinivasa, C
Wu, G., Hashemi, M. & Srinivasa, C. PUMA: Performance Unchanged Model Augmentation for Training Data Removal (2022). 2203.00846. 24
2022 arXiv
-
[113]
& Yang, L
Jing, S. & Yang, L. A robust extreme learning machine framework for uncertain data classification. Journal of Supercomputing 76, 2390–2416 (2020)
2020
-
[114]
& Luo, ZZ
Gan, HT., Li, ZH., Fan, YL. & Luo, ZZ. Dual Learning-Based Safe Semi- Supervised Learning. IEEE Access 6, 2615–2621 (2018)
2018
-
[115]
Engstrom, L. et al. Adversarial Robustness as a Prior for Learned Representations. arXiv (2019)
2019
-
[116]
& Lowd, D
Brophy, J. & Lowd, D. Machine Unlearning for Random Forests , 1092–1104 (PMLR, 2021)
2021
-
[117]
S., Tarun, A
Chundawat, V. S., Tarun, A. K., Mandal, M. & Kankanhalli, M. Zero-Shot Machine Unlearning. IEEE Transactions on Information Forensics and Security 18, 2345–2354 (2023)
2023
-
[118]
& Jha, S
Chen, J., Li, Y., Wu, X., Liang, Y. & Jha, S. Oliver, N., P´ erez-Cruz, F., Kramer, S., Read, J. & Lozano, J. A. (eds) ATOM: Robustifying Out-of-Distribution Detection Using Outlier Mining . (eds Oliver, N., P´ erez-Cruz, F., Kramer, S., Read, J. & Lozano, J. A.) Machine Learn...
2021
-
[119]
& Horvitz, E
Lakkaraju, H., Kamar, E., Caruana, R. & Horvitz, E. Identifying unknown unknowns in the open world: Representations and policies for guided exploration , Vol. 31 (2017)
2017
-
[120]
& Huang, QM
Zhuo, JB., Wang, SH., Zhang, WG. & Huang, QM. Deep Unsupervised Convolutional Domain Adaptation , 261–269 (2017)
2017
-
[121]
& Bishop, N
Bossens, DM. & Bishop, N. Explicit Explore, Exploit, or Escape (E4): Near- optimal safety-constrained reinforcement learning in polynomial time. Machine Learning 112, 817–858 (2023)
2023
-
[122]
& Trimpe, S
Massiani, PF., Heim, S., Solowjow, F. & Trimpe, S. Safe Value Functions. IEEE Transactions on Automatic Control 68, 2743–2757 (2023)
2023
-
[123]
& Shroff, N
Shi, M., Liang, Y. & Shroff, N. A Near-Optimal Algorithm for Safe Reinforcement Learning Under Instantaneous Hard Constraints , Vol. 202, 31243–31268 (2023)
2023
-
[124]
Hunt, N. et al. Verifiably Safe Exploration for End-to-End Reinforcement Learning (2021)
2021
-
[125]
J., Shen, A., Bastani, O
Ma, Y. J., Shen, A., Bastani, O. & Dinesh, J. Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning, Vol. 36, 5404–5412 (2022)
2022
-
[126]
Zwane, S. et al. Safe Trajectory Sampling in Model-Based Reinforcement Learning, Vol. 2023-August (2023). 25
2023
-
[127]
& Lauer, M
Fischer, J., Eyberg, C., Werling, M. & Lauer, M. Sampling-based Inverse Reinforcement Learning Algorithms with Safety Constraints , 791–798 (2021)
2021
-
[128]
& Zhou, M
Zhou, Z., Liu, G. & Zhou, M. A Robust Mean-Field Actor-Critic Reinforcement Learning Against Adversarial Perturbations on Agent States. IEEE Transactions on Neural Networks and Learning Systems 1–12 (2023)
2023
-
[129]
Aligning individual and collective welfare in complex socio- technical systems by combining metaheuristics and reinforcement learning
Bazzan, ALC. Aligning individual and collective welfare in complex socio- technical systems by combining metaheuristics and reinforcement learning. Engineering Applications of Artificial Intelligence 79, 23–33 (2019)
2019
-
[130]
J., Haupt, A
Christoffersen, P. J., Haupt, A. A. & Hadfield-Menell, D. Get It in Writing: Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL, AAMAS ’23, 448– 456 (International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 2023)
2023
-
[131]
Christiano, P. F. et al. Deep Reinforcement Learning from Human Preferences , Vol. 30 (Curran Associates, Inc., 2017)
2017
-
[132]
& Lipton, Z
Kaushik, D., Hovy, E. & Lipton, Z. C. Learning the Difference that Makes a Difference with Counterfactually-Augmented Data (arXiv, 2020). 1909.12434
2020 arXiv
-
[133]
& Sreenath, K
Li, ZY., Zeng, J., Thirugnanam, A. & Sreenath, K. Bridging Model-based Safety and Model-free Reinforcement Learning through System Identification of Low Dimensional Linear Models. Robotics: Science and System 18 (2022)
2022
-
[134]
& IEEE A Contact-Safe Reinforcement Learning Framework for Contact-Rich Robot Manipulation , 2476–2482 (2022)
Zhu, X., Kang, SC., Chen, JY. & IEEE A Contact-Safe Reinforcement Learning Framework for Contact-Rich Robot Manipulation , 2476–2482 (2022)
2022
-
[135]
& Inam, R
Terra, A., Riaz, H., Raizer, K., Hata, A. & Inam, R. Safety vs. Efficiency: AI-Based Risk Mitigation in Collaborative Robotics , 151–160 (2020)
2020
-
[136]
& Wagner, D
Carlini, N. & Wagner, D. Towards Evaluating the Robustness of Neural Networks , 39–57 (IEEE Computer Society, 2017)
2017
-
[137]
& Clune, J
Nguyen, A., Yosinski, J. & Clune, J. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images, 427–436 (IEEE, Boston, MA, USA, 2015)
2015
-
[138]
Kaufmann, M. et al. Testing Robustness Against Unforeseen Adversaries. arXiv (2019)
2019
-
[139]
Hendrycks, D. et al. What Would Jiminy Cricket Do? Towards Agents That Behave Morally (2021)
2021
-
[140]
Gardner, M. et al. Cohn, T., He, Y. & Liu, Y. (eds) Evaluating Models’ Local Decision Boundaries via Contrast Sets . (eds Cohn, T., He, Y. & Liu, Y.) Findings of the Association for Computational Linguistics: EMNLP 2020 , 1307–1323 26 (Association for Computational Linguistics...
2020
-
[141]
Spears, D. F. in Assuring the Behavior of Adaptive Agents (eds Rouff, C. A., Hinchey, M., Rash, J., Truszkowski, W. & Gordon-Spears, D.) Agent Technol- ogy from a Formal Perspective NASA Monographs in Systems and Software Engineering, 227–257 (Springer, London, 2006)
2006
-
[142]
& Putzer, H.A Safety Case Pattern for Systems with Machine Learning Components , Vol
Wozniak, E., Cˆ arlan, C., Acar-Celik, E. & Putzer, H.A Safety Case Pattern for Systems with Machine Learning Components , Vol. 12235 LNCS, 370–382 (2020)
2020
-
[143]
& Verma, R
Gholampour, P. & Verma, R. Adversarial Robustness of Phishing Email Detection Models, 67–76 (2023)
2023
-
[144]
& Steinhardt, J
Nanda, N., Chan, L., Lieberum, T., Smith, J. & Steinhardt, J. Progress measures for grokking via mechanistic interpretability. arXiv (2023)
2023
-
[145]
Olsson, C. et al. In-context learning and induction heads. Transformer Circuits Thread (2022). OpenAI
2022
-
[146]
& Fei-Fei, L
Karpathy, A., Johnson, J. & Fei-Fei, L. Visualizing and Understanding Recurrent Networks (arXiv, 2016). 1506.02078
2016 arXiv
-
[147]
S., Barrett, D
Morcos, A. S., Barrett, D. G. T., Rabinowitz, N. C. & Botvinick, M. On the importance of single directions for generalization. arXiv (2018)
2018
-
[148]
G., Cohen, S
Gyevnar, B., Wang, C., Lucas, C. G., Cohen, S. B. & Albrecht, S. V. Causal Explanations for Sequential Decision-Making in Multi-Agent Systems , AAMAS ’24, 771–779 (International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, 2024)
2024
-
[149]
& Iwane, H
Okawa, Y., Sasaki, T. & Iwane, H. Automatic Exploration Process Adjustment for Safe Reinforcement Learning with Joint Chance Constraint Satisfaction , Vol. 53, 1588–1595 (2020)
2020
-
[150]
& She, QS
Gan, HT., Luo, ZZ., Meng, M., Ma, YL. & She, QS. A risk degree-based safe semi-supervised learning algorithm. International Journal of Machine Learning and Cybernetics 7, 85–94 (2016)
2016
-
[151]
Zhang, YX. et al. Barrier Lyapunov Function-Based Safe Reinforcement Learning for Autonomous Vehicles With Optimized Backstepping. IEEE Transactions on Neural Networks and Learning Systems (2022)
2022
-
[152]
& Amini, M.-R
Nicolae, M.-I., Sebban, M., Habrard, A., Gaussier, E. & Amini, M.-R. Algorithmic robustness for semi-supervised ( ϵ, γ, τ )-good metric learning, Vol. 9489, 253–263 (2015)
2015
-
[153]
& Seto, E
Khan, WU. & Seto, E. A ”Do No Harm” Novel Safety Checklist and Research Approach to Determine Whether to Launch an Artificial Intelligence-Based 27 Medical Technology: Introducing the Biological-Psychological, Economic, and Social (BPES) Framework. Journal of Medical Internet ...
2023
-
[154]
& Werner, B
Schumeg, B., Marotta, F. & Werner, B. Proposed V-Model for Verification, Validation, and Safety Activities for Artificial Intelligence , 61–66 (2023)
2023
-
[155]
& Weyrich, M
Kamm, S., Sahlab, N., Jazdi, N. & Weyrich, M. A Concept for Dynamic and Robust Machine Learning with Context Modeling for Heterogeneous Manufactur- ing Data , Vol. 118, 354–359 (2023)
2023
-
[156]
& Nogueira, I
Costa, E., Rebello, C., Fontana, M., Schnitman, L. & Nogueira, I. A Robust Learning Methodology for Uncertainty-Aware Scientific Machine Learning Models. Mathematics 11 (2023)
2023
-
[157]
& Kyrki, V
Aksjonov, A. & Kyrki, V. A Safety-Critical Decision-Making and Control Framework Combining Machine-Learning-Based and Rule-Based Algorithms. SAE International Journal of Vehicle Dynamics Stability And NVH 7, 287–299 (2023)
2023
-
[158]
Antikainen, J. et al. Yue, T. & Mirakhorli, M. (eds) A Deployment Model to Extend Ethically Aligned AI Implementation Method ECCOLA . (eds Yue, T. & Mirakhorli, M.) 29th IEEE International Requirements Engineering Conference Workshops (REW 2021) , 230–235 (2021)
2021
-
[159]
& Abrahamsson, P
Vakkuri, V., Kemell, KK., Jantunen, M., Halme, E. & Abrahamsson, P. ECCOLA - A method for implementing ethically aligned AI systems. Journal of Systems and Software 182 (2021)
2021
-
[160]
& Asudeh, A
Zhang, H., Shahbazi, N., Chu, X. & Asudeh, A. FairRover: Explorative model building for fair and responsible machine learning (2021)
2021
-
[161]
Coston, A. et al. A Validity Perspective on Evaluating the Justified Use of Data-driven Decision-making Algorithms , 690–704 (2023)
2023
-
[162]
& Yung, M
Gittens, A., Yener, B. & Yung, M. An Adversarial Perspective on Accuracy, Robustness, Fairness, and Privacy: Multilateral-Tradeoffs in Trustworthy ML. IEEE Access 10, 120850–120865 (2022)
2022
-
[163]
& Critch, A
Taylor, J., Yudkowsky, E., LaVictoire, P. & Critch, A. Alignment for advanced machine learning systems. Ethics of Artificial Intelligence 342–382 (2016)
2016
-
[164]
& Yampolskiy, R
Sotala, K. & Yampolskiy, R. V. Responses to catastrophic AGI risk: A survey. Physica Scripta 90, 018001 (2014)
2014
-
[165]
Metacognition for artificial intelligence system safety-An approach to safe and desired behavior
Johnson, B. Metacognition for artificial intelligence system safety-An approach to safe and desired behavior. Safety Science 151 (2022)
2022
-
[166]
Hatherall, L. et al. Responsible Agency Through Answerability (2022). 28
2022
-
[167]
Embedding responsibility in intelligent systems: From AI ethics to responsible AI ecosystems
Stahl, BC. Embedding responsibility in intelligent systems: From AI ethics to responsible AI ecosystems. Scientific Reports 13 (2023)
2023
-
[168]
Counterfactual learning in enhancing resilience in autonomous agent systems
Samarasinghe, D. Counterfactual learning in enhancing resilience in autonomous agent systems. Frontiers in Artificial Intelligence 6 (2023)
2023
-
[169]
& Joyce, J
Diemert, S., Millet, L., Groves, J. & Joyce, J. Safety Integrity Levels for Artificial Intelligence, Vol. 14182 LNCS, 397–409 (2023)
2023
-
[170]
& Jia, R
Wang, J. & Jia, R. Data Banzhaf: A Robust Data Valuation Framework for Machine Learning, Vol. 206, 6388–6421 (2023)
2023
-
[171]
& Hutter, M
Everitt, T., Filan, D., Daswani, M. & Hutter, M. Self-Modification of Policy and Utility Function in Rational Agents. arXiv (2016)
2016
-
[172]
& Artus, G
Badea, C. & Artus, G. Bramer, M. & Stahl, F. (eds) Morality, Machines, and the Interpretation Problem: A Value-based, Wittgensteinian Approach to Building Moral Agents. (eds Bramer, M. & Stahl, F.) Artificial Intelligence, AI 2022 , Vol. 39, 124–137 (2022)
2022
-
[173]
Beneficial Artificial Intelligence Coordination by Means of a Value Sensitive Design Approach
Umbrello, S. Beneficial Artificial Intelligence Coordination by Means of a Value Sensitive Design Approach. Big Data and Cognitive Computing 3 (2019)
2019
-
[174]
& Fox, J
Yampolskiy, R. & Fox, J. Safety Engineering for Artificial General Intelligence. Topoi (2012)
2012
-
[175]
& Etzioni, O
Weld, D. & Etzioni, O. Barley, M. et al. (eds) The First Law of Robotics . (eds Barley, M. et al. ) Safety and Security in Multiagent Systems , Lecture Notes in Computer Science, 90–100 (Springer, Berlin, Heidelberg, 2009)
2009
-
[176]
The responsibility gap: Ascribing responsibility for the actions of learning automata
Matthias, A. The responsibility gap: Ascribing responsibility for the actions of learning automata. Ethics and Information Technology 6, 175–183 (2004)
2004
-
[177]
Muller, VC
Farina, L. Muller, VC. (ed.) Artificial Intelligence Systems, Responsibility and Agential Self-Awareness. (ed.Muller, VC.) Philosophy and Theory of Artificial Intelligence 2021 , Vol. 63, 15–25 (2022)
2022
-
[178]
Lee, A. T. Flight simulation: virtual environments in aviation (Routledge, 2017)
2017
-
[179]
Weidinger, L. et al. Sociotechnical safety evaluation of generative ai systems. arXiv preprint arXiv:2310.11986 (2023). 29
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.