REVIEW 4 major objections 4 minor 15 references
What Is AI Safety? What Do We Want It to Be?
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that a research project counts as AI safety exactly when it aims to prevent or reduce harms from AI systems, making bias, misinformation, and privacy core topics rather than peripheral ones.
desk verdict A genuinely useful conceptual-engineering defense of the broad harm-based definition of AI safety, but the Section 5 normative argument compares allocation rules rather than concepts and overstates its case. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is The Safety Conception, a constitutive criterion: a research project is AI safety research just in case it aims to prevent or reduce harms from AI systems' development and deployment. The argument's machinery is the two-purpose evaluative framework borrowed from conceptual engineering and ameliorative inquiry: any concept of AI safety is assessed by whether it (A) unifies paradigmatic research programs under one field and (B) conduces to reducing AI harms. The comparison is carried out through three ideal-typical scenarios—The Safe Scenario, The Catastrophic Scenario, and The Engineering Scenario—which differ in how research is prioritized and which experts are consulted; The Safety Conception corresponds to the scenario that prioritizes purely by expected harm reduction.
What would settle it
A concrete check: compare funding and policy-consultation records before and after a funder or government body adopts the broad definition; if the set of supported projects and invited experts does not shift toward social-harm research, the practical argument for The Safety Conception loses force.
Extended reading notes
Core claim
The paper's central claim is that The Safety Conception is the best conception of AI safety and therefore ought to be the operative concept, not merely the stated one. Using conceptual engineering, it evaluates candidate concepts against two purposes: (A) providing a unifying explanation of why paradigmatic AI safety research programs belong to the same field, and (B) being conducive to reducing the harms caused by AI systems. It argues that The Safety Conception does both better than The Catastrophic Conception and The Engineering Conception, and that any more restrictive or more permissive departure fares worse. If the argument is right, AI safety includes work on social harms such as bias, misinformation, privacy, and economic harms alongside work on catastrophic harms, and researchers in both areas should be integrated, including in political conversations about AI safety.
Load-bearing premise
The argument stands on the assumption that the only tests that matter for a definition of AI safety are how well it unifies the field and how much it helps reduce AI harms, and that labels shape what gets funded and who gets heard.
Editorial extensions
If this is right
- Research on algorithmic bias, misinformation, privacy, and economic harms qualifies as AI safety, so it should be presented and published in the same venues as research on catastrophic risks.
- AI labs should not maintain separate ethics and safety teams, and funders should support harm-reduction research that is not framed in terms of catastrophic or existential risk.
- Political deliberations about AI safety should include researchers working on social harms as well as those working on catastrophic harms.
- Requiring work on present systems to justify itself by future catastrophic relevance is a form of gatekeeping that the paper predicts makes the field less effective at reducing harm.
- Because the same model behavior (for example, generating hate speech and generating bomb instructions) can produce both social and catastrophic harms, drawing a hard line between them forfeits explanatory continuity.
Reading between the lines
- Not pursued in the paper but testable: if funders and venues adopted the broad conception, the portfolio of funded AI safety research should measurably shift toward sociotechnical and governance work; comparing funders with broad versus narrow definitions could test the paper's causal premise.
- The same conceptual-engineering method could be pointed at adjacent categories such as AI ethics or responsible AI, which would dissolve the safety/ethics split from the other direction and could produce different policy alliances.
- The paper assumes that harm-reducing merit is comparable between social and catastrophic risks; one implicit task for future work is a shared metric or decision procedure for comparing them, since the argument does not supply one.
- If the paper is right, marginalization of social-harm research is partly a labeling effect; an observable extension is to audit which researchers are invited into AI safety policy settings and whether invitations track the broad harm taxonomy.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks what AI safety as a discipline should be. It states and defends The Safety Conception: a research project belongs to the field of AI safety just in case it aims to prevent or reduce harms from AI systems' development and deployment. Sections 2 and 3 argue that this conception is widely held as the manifest concept but is not the operative concept in practice, documenting two rival tendencies: a focus on catastrophic or existential risks from future systems (The Catastrophic Conception) and an understanding of AI safety as a branch of safety engineering (The Engineering Conception). Section 4 adopts the methodology of conceptual engineering and stipulates two purposes for a concept of AI safety: (A) explanatory unification of paradigmatic AI safety research, and (B) conduciveness to reducing harms caused by AI systems. Section 5 argues that The Safety Conception serves both purposes better than the two rivals and then generalizes to the claim that any departure from The Safety Conception is worse. The paper concludes that AI safety should include social harms and catastrophic harms, and that greater disciplinary integration and broader inclusion in political conversations are required.
Significance. The paper is a serious, clearly written conceptual-engineering contribution to an important public and policy-relevant debate. If its normative argument succeeded, it would provide a principled basis for changing how research is funded, organized, and represented in policy discussions. Its independent descriptive contributions are substantial: the taxonomy of AI safety research in Section 2 is useful, and Section 3 gives concrete evidence that operative conceptions of AI safety diverge from the manifest conception. The paper is also commendably explicit about its normative methodology, and it does not rest on hidden technical machinery. However, the central normative argument currently compares resource-allocation scenarios rather than concepts, and the step from two rival conceptions to 'any departure' is underargued. The significance of the paper is therefore conditional on repairing those load-bearing steps.
major comments (4)
- [Section 5, Safe/Catastrophic/Engineering Scenarios] The comparison in Section 5 does not establish the definitional conclusion. The Safe Scenario is characterized by a resource-allocation rule: prioritize research by expected contribution to harm reduction and consult experts by harm-reduction expertise. The Catastrophic and Engineering Scenarios are characterized by restrictions built into their allocation rules. But one can accept The Catastrophic Conception as a classification of the field and still fund non-catastrophic harm-reduction research under a different heading; conversely, one can accept The Safety Conception and still allocate resources badly. The scenario comparison therefore supports at most the near-tautological claim that an unrestricted merit-based allocation rule is better than a stipulated restricted one. The bridge from that claim to the conclusion that The Safety Conception is the best concept must be an explicit, defended premise that operative concepts causally shape research prioritization and expert inclusion. That premise is asserted in Sections 1 and 4 but not argued with evidence specific to AI safety.
- [Section 5, generalizing beyond two rivals] The paper generalizes from two rivals to 'any departure from The Safety Conception' in two places: the purpose-(A) argument ('no conception of AI safety more restrictive than The Safety Conception could fare better' and 'any less restrictive conception would fare poorly') and the purpose-(B) argument ('any way of choosing research priorities or selecting experts other than the one embodied in The Safe Scenario is likely to be less effective'). Only The Catastrophic Conception and The Engineering Conception are examined. Not all departures have the features that make those two rivals fail. For example, a more restrictive conception that excludes only research with zero expected contribution to harm reduction would not obviously be worse under purpose (B), and a less restrictive conception organized around 'beneficial AI' could still be explanatorily unified. The authors should either restrict the conclusion to the two rivals and explicitly frame the thesis as comparative, or provide an argument covering the full space of alternatives.
- [Section 4, stipulation of purposes (A) and (B)] The conclusion that The Safety Conception is 'the best conception of AI safety' depends on the stipulation that a concept of AI safety should be evaluated by exactly purposes (A) and (B). The paper gives no argument that these two purposes are the only ones that matter or that they are the right weights relative to other plausible purposes, such as epistemic integrity, public trust, accountability, or clarity of responsibility assignment. If additional purposes are admitted, the ranking of candidate conceptions could change. The argument would be more defensible if the thesis were explicitly stated as conditional on the two stipulated purposes, or if a substantive argument for the completeness and weighting of the purpose set were provided.
- [Sections 1 and 4, sociological bridge premise] The normative claim that revising the operative concept of AI safety will reduce harm depends on a causal-sociological premise: disciplinary boundaries shape research priorities and policy inclusion in ways that matter for harm outcomes. The paper cites general work in sociology of disciplines (Trowler 2012), but it does not provide evidence that this link holds specifically for AI safety, where funding streams, venue structures, and policy processes may be shaped by many factors besides the operative concept. Without this premise, even a successful comparison of scenarios would not show that adopting The Safety Conception would have the practical consequences claimed in the conclusion. The authors should either provide empirical support for the premise or weaken the inference from conceptual choice to harm reduction.
minor comments (4)
- [Section 1, paragraph on FATE research] The acronym 'F AccTor F ATE' appears garbled and should be corrected to 'FAccT' or 'FAT*'.
- [Section 3.1, note 35] The MIRI quotation in note 35 lacks a URL, unlike the surrounding sources; this should be completed or the note reformatted.
- [Section 5, final paragraph] The phrase 'as we suspect it will' relies on an unstated empirical conjecture about the relative merit of catastrophic versus non-catastrophic interventions; if this conjecture is not defended, it should be marked as an open empirical question rather than used in the argument.
- [References] Several references have OCR artifacts or formatting inconsistencies (for example, the Rafailov reference and the Wijk et al. entry); these should be cleaned in the final version.
Circularity Check
The purpose-(B) argument builds its conclusion into the Safe Scenario's construction, but the explanatory purpose-(A) argument and the Sections 2–3 evidence give the central claim independent content.
-
self definitional
[Section 5, purpose (B), comparison of The Safe, Catastrophic, and Engineering Scenarios, and the concluding inference]
"In the first, which we might call The Safe Scenario, the AI safety research community is structured in accordance with The Safety Conception. Research projects are prioritized based on their expected contribution to the goal of preventing or mitigating harms from AI systems ... In so far as The Safe Scenario is the scenario that embodies The Safety Conception, we have reason to believe that The Safety Conception is the best conception of AI safety for purpose (B)."
By stipulation, The Safe Scenario is the allocation rule that directly optimizes purpose (B) — prioritize by expected contribution to harm reduction, consult experts by harm-reduction expertise — while the rival scenarios are the same rule with built-in restrictions. So the result that the first is more conducive to harm reduction is true by construction. The argument then transfers this scenario-level result to the concept-level claim via 'In so far as The Safe Scenario ... embodies The Safety Conception'. But a concept's extension does not by itself fix an allocation rule: one could accept The Catastrophic Conception as a classification and still fund non-catastrophic harm-reduction research under another heading, or accept The Safety Conception and allocate badly.
full rationale
The paper's overall thesis is not a fitted prediction and is not carried by a self-citation chain: the self-citations (Harding 2023; Goldstein and Kirk-Giannini 2023; Kirk-Giannini 2023/forthcoming) are illustrative, and no uniqueness theorem or machine-checked result is invoked. The purpose-(A) argument has independent content, appealing to continuities between central and marginal research programs, and Sections 2–3 provide non-circular evidence of tension between The Safety Conception and operative practice. The partial circularity is confined to the purpose-(B) argument in Section 5, where the Safe Scenario is constructed so that it directly embodies both The Safety Conception and the harm-merit allocation rule, making the result that it best serves harm reduction true by stipulation. The sociological premise that operative concepts shape research prioritization and policy inclusion is asserted rather than independently supported, and it is needed to bridge from allocation rules back to the definition of the field. Because the central claim also rests on the independent explanatory argument, the paper deserves a moderate partial-circularity score rather than a high one.
Assumptions & free parameters
assumptions (5)
- ad hoc to paper A concept of AI safety should be evaluated by purpose (A), explanatory unification, and purpose (B), conduciveness to reducing AI harms.
- domain assumption Disciplinary boundaries shape research expectations, collaboration, evaluation norms, and inclusion in policy discussions.
- domain assumption The Safety Conception is a live manifest concept among AI safety researchers.
- ad hoc to paper No conception of AI safety that is more or less restrictive than The Safety Conception could better serve purposes (A) and (B).
- ad hoc to paper Conceptual engineering is an appropriate methodology for deciding what AI safety should be.
Cite this review
Pith. "Pith review of What Is AI Safety? What Do We Want It to Be?." pith.science (2026). https://pith.science/paper/XILYL5DM
@misc{pith2026250502313,
author = {Pith},
title = {Pith review of: What Is AI Safety? What Do We Want It to Be?},
year = {2026},
howpublished = {\url{https://pith.science/paper/XILYL5DM}},
note = {Machine review of arXiv:2505.02313}
}
read the original abstract
The field of AI safety seeks to prevent or reduce the harms caused by AI systems. A simple and appealing account of what is distinctive of AI safety as a field holds that this feature is constitutive: a research project falls within the purview of AI safety just in case it aims to prevent or reduce the harms caused by AI systems. Call this appealingly simple account The Safety Conception of AI safety. Despite its simplicity and appeal, we argue that The Safety Conception is in tension with at least two trends in the ways AI safety researchers and organizations think and talk about AI safety: first, a tendency to characterize the goal of AI safety research in terms of catastrophic risks from future systems; second, the increasingly popular idea that AI safety can be thought of as a branch of safety engineering. Adopting the methodology of conceptual engineering, we argue that these trends are unfortunate: when we consider what concept of AI safety it would be best to have, there are compelling reasons to think that The Safety Conception is the answer. Descriptively, The Safety Conception allows us to see how work on topics that have historically been treated as central to the field of AI safety is continuous with work on topics that have historically been treated as more marginal, like bias, misinformation, and privacy. Normatively, taking The Safety Conception seriously means approaching all efforts to prevent or mitigate harms from AI systems based on their merits rather than drawing arbitrary distinctions between them.
Reference graph
Works this paper leans on
-
[5]
F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D
<https://news.artnet.com/art-world/ class- action-lawsuit-lensa-ai-prisma-labs-biometric-information-2257096> Christiano, P . F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in neural information processing systems ,
work page 2017
-
[7]
https://doi.org/10.48550/arXiv.2405.06624. Dembroff, R. (2016). What is sexual orientation? Philosophers’ Imprint 16: 1–27. Dobbe, R. I. J. (2022). System Safety and Artificial Intellig ence. In Justin B. Bullock, Y u-Che Chen, Johannes Himmelreich, V alerie M. Hudson, Anto n Korinek, Matthew M. Y oung, and Baobao Zhang (eds). The Oxford Handbook of AI Gov...
-
[8]
Harding, J. and Sharadin, N. (Forthcoming). What is it for a M achine Learning Model to Have a Capability? British Journal for the Philosophy of Science. Haslanger, S. (1995). Ontology and Social Construction. Philosophical Topics 23: 95–125. Reprinted in Haslanger (2012), pp. 83–112. Haslanger, S. (2000). Gender and race: (What) are they? (Wha t) do we w...
arXiv 1995
-
[9]
Hendrycks, D., Mazeika, M., and Woodside, T. (2023). An Over view of Catastrophic AI Risks. ArXiv preprint. <https://arxiv.org/abs/2306.12001> Hendrycks, D., Mazeika, M., Zou, A., Patel, S., Zhu, C., Nava rro, J., Song, D., Li, B. and Steinhardt, J. (2021). What would Jiminy Cricket do? Toward agents that behave mora lly. 35th Conference on Neural Informa...
arXiv 2023
-
[12]
Rafailov, R., Sharma, A., Mitchell, E., Manning, C
<https://time.com/6288245/openai-eu-lobbying-ai-act/ >. Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermo n, S., & Finn, C. (2024). Direct preference optimization: Y our language model is secretly a reward mode l. Advances in Neural Information Processing Systems
-
[13]
Raji, I. D., Smart, A., White, R. M., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., and Barnes, P . (2020). Closing the AI Accountability Gap: De fining an End-to-End Framework for Internal Algorithmic Auditing. Proceedings of the 2020 Conference on Fairness, Accountabi lity, and Transparency : 33–44. Rismani, S., Shelby, R., Smart, ...
work page 2020
-
[14]
Russell, S. (2019). Human Compatible: AI and the Problem of Control . Allen Lane. Schneier, B. and Sanders, N. (2023). The A.I. Wars Have Three Factions, and They All Crave Power. New York Times September 28,
work page 2019
-
[15]
Selbst, A. D., Boyd, D., Friedler, S. A., V enkatasubramania n, S., and V ertesi, J. (2019). Fairness and Abstrac- tion in Sociotechnical Systems. Proceedings of the Conference on Fairness, Accountability , and Transparency (F AT* ’19): 59–68. Shavit, Y . (2023). What Does It Take to Catch a Chinchilla? V e rifying Rules on Large-Scale Neural Network Trai...
arXiv 2019
Show all 15 references
-
[30]
and Chalmers, D
Clark, A. and Chalmers, D. (1998). The extended mind. Analysis 58: 7–19. Clark, J. and Amodei, D. (2016). Faulty reward function in th e wild. OpenAI blog post. <https://openai.com/research/faulty-reward-functions> Corso, A., Karamadian, D., V alentin, R., Cooper, M., and Koc ...
1998 arXiv
-
[36]
Manheim, K., & Kaplan, L. (2019). Artificial intelligence: R isks to privacy and democracy. Yale Journal of Law and Technology 21: 106–188. Nemitz P . (2018). Constitutional democracy and technology in the age of artificial intelligence. Philosophical Transactions of the Royal S...
2019 arXiv
-
[91]
Cappelen, H. (2018). Fixing Language: An Essay on Conceptual Engineering . Oxford University Press. Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber , J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A. (2019). On evaluating adversarial robustness. ArXiv pre...
2018 arXiv
-
[1995]
doi: 10.1207/s15516709cog1903_1
ISSN 03640213. doi: 10.1207/s15516709cog1903_1. Irving, G., Christiano, P ., & Amodei, D. (2018). AI safety vi a debate. ArXiv preprint. <https://arxiv.org/abs/1805.00899>. Isaac, M. G. (2020). How to conceptually engineer conceptua l engineering? Inquiry. Online First. Kasirz...
2018 arXiv
-
[2022]
Bowman, S., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner , S.,
<https://www.lesswrong.com/posts/EFp QcBmfm2bFfM4zM/ai-safety-and-neighboring- communities-a-quick-start-guide-as >. Bowman, S., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner , S., ... & Kaplan, J. (2022). Measuring progress on scalable oversight for large language models....
2022 arXiv
-
[2023]
Collag- ing
(2023). Artists and Illustrators Are Suing Th ree A.I. Art Generators for Scraping and “Collag- ing” Their Work Without Consent’. Artnet News February 16,
2023
-
[2024]
Aïmeur, E., Amri, S., and Brassard, G
<https://doi.org/10.5210/fm.v29i4.1362 6 >. Aïmeur, E., Amri, S., and Brassard, G. (2023). Fake news, dis information and misinformation in social media: A review. Social Network Analysis and Mining 13(30): 1–36. Allen, A. L. (2016). Protecting one's own privacy in a big dat a...
2023 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.