{"id":"7a05c243-2bb1-4531-8b83-fbe7c6389088","arxiv_id":"2505.02313","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI safety is best understood as all research aimed at preventing or reducing harms from AI systems, covering social harms and catastrophic risks together.","lead":"The paper argues that AI safety should be defined broadly as any research aimed at preventing or reducing harms from AI systems, rather than only work on catastrophic future risks or safety engineering. It offers a conceptual-engineering case for this definition and says the field should integrate work on social harms like bias, misinformation, and privacy with work on existential risks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The harm-reduction argument for The Safety Conception conflates the concept with a merit-based priority-setting scenario; the Section 5 comparison supports merit-based allocation, not the definitional claim.","rationale":"The reader correctly identified the stipulated purposes and the generalization from two rivals as the weakest assumptions; my concern sharpens this by showing the normative purpose (B) argument is internally misaligned: it compares scenarios defined by allocation rules, not concepts. This is independent of any dispute with the authors' preferred conclusions (1) and (2), which I think are reasonably motivated. The paper is a serious conceptual-engineering contribution with a clear structure and fair treatment of rival views. The concern is addressable in revision: the authors could either provide empirical evidence for the concept-to-allocation link or reformulate the B argument so that the concept itself, rather than a stipulated decision procedure, does the normative work. Because the gap is substantial but repairable, the appropriate verdict remains CONDITIONAL, matching the reader's judgment. I mark agreement partial because the reader's weakest assumption includes the sociological premise but not the scenario/concept conflation, which is the precise place where the argument's load is carried.","tokens_in":21126,"tokens_out":5569,"duration_ms":74213,"concrete_test":"Perform a controlled conceptual analysis of Section 5 by separating two variables: (i) concept extension (Safety, Catastrophic, Engineering) and (ii) resource-allocation rule (merit-based unrestricted vs restricted). Enumerate the six possible combinations and evaluate each against purpose (B). If every combination with the merit-based unrestricted allocation rule is equally effective regardless of concept extension, then the 'Safe Scenario' advantage is an artifact of the allocation rule, and Section 5 does not support The Safety Conception over its rivals. This can be done as a short philosophical model with cases; no new empirical data needed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that The Safety Conception is the best concept of AI safety is not established because the normative argument for purpose (B) in Section 5 compares the wrong objects. The 'Safe Scenario' is characterized by a resource-allocation rule — prioritize projects by expected contribution to harm reduction and consult experts by harm-reduction expertise — not by the extension of the concept 'AI safety.' The Catastrophic and Engineering Scenarios are characterized by restricted allocation rules built into the rival concepts by stipulation. This asymmetry drives the conclusion. One can accept The Catastrophic Conception as a classification of the field and still fund or consult researchers working on present non-catastrophic harms under a different heading; conversely, one can accept The Safety Conception and still allocate resources badly. So the scenario comparison establishes at most that an unrestricted, merit-based allocation rule is better than a restricted one, which is close to tautological and does not entail that The Safety Conception is the best concept. The only bridge from that conclusion to the definitional claim is the sociological premise that operative concepts shape research prioritization and policy inclusion; this premise is asserted in Sections 1 and 4 but not empirically supported for AI safety. Moreover, the generalization from two rivals to every departure from The Safety Conception (Section 5) is a leap, not an argument. These gaps are internal to the argument, not disagreements with external consensus. They leave claims (1) and (2) plausible but conditional.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks what AI safety as a discipline should be. It states and defends The Safety Conception: a research project belongs to the field of AI safety just in case it aims to prevent or reduce harms from AI systems' development and deployment. Sections 2 and 3 argue that this conception is widely held as the manifest concept but is not the operative concept in practice, documenting two rival tendencies: a focus on catastrophic or existential risks from future systems (The Catastrophic Conception) and an understanding of AI safety as a branch of safety engineering (The Engineering Conception). Section 4 adopts the methodology of conceptual engineering and stipulates two purposes for a concept of AI safety: (A) explanatory unification of paradigmatic AI safety research, and (B) conduciveness to reducing harms caused by AI systems. Section 5 argues that The Safety Conception serves both purposes better than the two rivals and then generalizes to the claim that any departure from The Safety Conception is worse. The paper concludes that AI safety should include social harms and catastrophic harms, and that greater disciplinary integration and broader inclusion in political conversations are required.","tokens_in":21404,"tokens_out":5521,"duration_ms":76179,"significance":"The paper is a serious, clearly written conceptual-engineering contribution to an important public and policy-relevant debate. If its normative argument succeeded, it would provide a principled basis for changing how research is funded, organized, and represented in policy discussions. Its independent descriptive contributions are substantial: the taxonomy of AI safety research in Section 2 is useful, and Section 3 gives concrete evidence that operative conceptions of AI safety diverge from the manifest conception. The paper is also commendably explicit about its normative methodology, and it does not rest on hidden technical machinery. However, the central normative argument currently compares resource-allocation scenarios rather than concepts, and the step from two rival conceptions to 'any departure' is underargued. The significance of the paper is therefore conditional on repairing those load-bearing steps.","major_comments":[{"comment":"The comparison in Section 5 does not establish the definitional conclusion. The Safe Scenario is characterized by a resource-allocation rule: prioritize research by expected contribution to harm reduction and consult experts by harm-reduction expertise. The Catastrophic and Engineering Scenarios are characterized by restrictions built into their allocation rules. But one can accept The Catastrophic Conception as a classification of the field and still fund non-catastrophic harm-reduction research under a different heading; conversely, one can accept The Safety Conception and still allocate resources badly. The scenario comparison therefore supports at most the near-tautological claim that an unrestricted merit-based allocation rule is better than a stipulated restricted one. The bridge from that claim to the conclusion that The Safety Conception is the best concept must be an explicit, defended premise that operative concepts causally shape research prioritization and expert inclusion. That premise is asserted in Sections 1 and 4 but not argued with evidence specific to AI safety.","section":"Section 5, Safe/Catastrophic/Engineering Scenarios"},{"comment":"The paper generalizes from two rivals to 'any departure from The Safety Conception' in two places: the purpose-(A) argument ('no conception of AI safety more restrictive than The Safety Conception could fare better' and 'any less restrictive conception would fare poorly') and the purpose-(B) argument ('any way of choosing research priorities or selecting experts other than the one embodied in The Safe Scenario is likely to be less effective'). Only The Catastrophic Conception and The Engineering Conception are examined. Not all departures have the features that make those two rivals fail. For example, a more restrictive conception that excludes only research with zero expected contribution to harm reduction would not obviously be worse under purpose (B), and a less restrictive conception organized around 'beneficial AI' could still be explanatorily unified. The authors should either restrict the conclusion to the two rivals and explicitly frame the thesis as comparative, or provide an argument covering the full space of alternatives.","section":"Section 5, generalizing beyond two rivals"},{"comment":"The conclusion that The Safety Conception is 'the best conception of AI safety' depends on the stipulation that a concept of AI safety should be evaluated by exactly purposes (A) and (B). The paper gives no argument that these two purposes are the only ones that matter or that they are the right weights relative to other plausible purposes, such as epistemic integrity, public trust, accountability, or clarity of responsibility assignment. If additional purposes are admitted, the ranking of candidate conceptions could change. The argument would be more defensible if the thesis were explicitly stated as conditional on the two stipulated purposes, or if a substantive argument for the completeness and weighting of the purpose set were provided.","section":"Section 4, stipulation of purposes (A) and (B)"},{"comment":"The normative claim that revising the operative concept of AI safety will reduce harm depends on a causal-sociological premise: disciplinary boundaries shape research priorities and policy inclusion in ways that matter for harm outcomes. The paper cites general work in sociology of disciplines (Trowler 2012), but it does not provide evidence that this link holds specifically for AI safety, where funding streams, venue structures, and policy processes may be shaped by many factors besides the operative concept. Without this premise, even a successful comparison of scenarios would not show that adopting The Safety Conception would have the practical consequences claimed in the conclusion. The authors should either provide empirical support for the premise or weaken the inference from conceptual choice to harm reduction.","section":"Sections 1 and 4, sociological bridge premise"}],"minor_comments":[{"comment":"The acronym 'F AccTor F ATE' appears garbled and should be corrected to 'FAccT' or 'FAT*'.","section":"Section 1, paragraph on FATE research"},{"comment":"The MIRI quotation in note 35 lacks a URL, unlike the surrounding sources; this should be completed or the note reformatted.","section":"Section 3.1, note 35"},{"comment":"The phrase 'as we suspect it will' relies on an unstated empirical conjecture about the relative merit of catastrophic versus non-catastrophic interventions; if this conjecture is not defended, it should be marked as an open empirical question rather than used in the argument.","section":"Section 5, final paragraph"},{"comment":"Several references have OCR artifacts or formatting inconsistencies (for example, the Rafailov reference and the Wijk et al. entry); these should be cleaned in the final version.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"I see no publication-integrity concern. The paper is within scope for a venue interested in the societal dimensions of AI, and its descriptive sections are valuable. My recommendation is driven by the internal gaps in the Section 5 argument, which I believe are fixable within the paper's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper does something useful. It takes the standard 'AI safety = preventing harms from AI' definition, gives it a name (The Safety Conception), distinguishes it from the operative concept in the field, and builds a systematic conceptual-engineering argument for why it should be the operative concept. The descriptive/explanatory part is solid. The normative part has a structural flaw.\n\nWhat's genuinely new: the named, defended conclusion that The Safety Conception is the best concept, and the explicit two-purpose framework (explanatory unification, harm-reduction conduciveness). The paper is also honest: it cites Amodei et al., CSET, AISI, and Wikipedia as already endorsing the broad definition, and concedes its recommendations are continuous with Lazar and Nelson (2023).\n\nWhat it does well: the survey of research areas under the harm-reduction lens is clear; the manifest/operative distinction is applied carefully; and the concluding prescriptions (no separate ethics vs safety teams, funders open to non-catastrophic work) follow naturally.\n\nNow the soft spots, in proportion. The Section 5 argument for purpose (B) compares three scenarios that differ in their resource-allocation and expert-consultation rules, not merely in their concepts. The Safe Scenario is effectively 'prioritize everything by expected harm reduction'; the rivals are defined with restricted allocation rules built in. So the Safe Scenario wins largely by construction. That supports merit-based unrestricted allocation, not the claim that the broad concept itself is best. You could accept the Catastrophic Conception as a classification and still fund bias research under a different heading. The bridge from 'merit-based allocation is better' to 'this concept is better' is the sociological premise that operative concepts shape prioritization and policy inclusion—asserted, not demonstrated for AI safety. Also, the generalization from two rival conceptions to any departure is a leap; a power/legitimacy-based conception, which they themselves mention, isn't addressed.\n\nThese are real gaps, but the central descriptive claim holds up and the normative claim is plausible. The paper deserves a serious referee; I'd ask the authors to either support the sociological premise or soften the conclusion to a conditional one.\n\nRecommendation: send to peer review.","headline":"A genuinely useful conceptual-engineering defense of the broad harm-based definition of AI safety, but the Section 5 normative argument compares allocation rules rather than concepts and overstates its case.","tokens_in":21882,"tokens_out":2766,"would_cite":true,"duration_ms":30409,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a research project counts as AI safety exactly when it aims to prevent or reduce harms from AI systems, making bias, misinformation, and privacy core topics rather than peripheral ones.","keywords":["AI safety","conceptual engineering","harm reduction","catastrophic risk","safety engineering","AI governance","social harms of AI","disciplinary boundaries"],"falsifier":"A concrete check: compare funding and policy-consultation records before and after a funder or government body adopts the broad definition; if the set of supported projects and invited experts does not shift toward social-harm research, the practical argument for The Safety Conception loses force.","tokens_in":20906,"feed_emoji":"🛡️","tokens_out":10201,"duration_ms":98294,"temperature":0.7,"pith_summary":"This paper tries to settle a live boundary question: what makes a research project part of AI safety? It defends The Safety Conception, which says a project belongs to the field just in case it aims to prevent or reduce harms from the development and deployment of AI systems. The paper argues that rival conceptions—one focused only on catastrophic risks from future systems, one that treats AI safety as a branch of safety engineering—are worse on both explanatory and practical grounds. A sympathetic reader should care because the paper's conclusion makes bias, misinformation, privacy, and economic harms core AI safety topics, with consequences for funding, publication venues, and who gets listened to in policy.","feed_headline":"AI safety should mean preventing every AI harm, not just catastrophe","feed_subtitle":"That puts bias, misinformation, and privacy in the same field as existential risk, reshaping funding and policy.","key_machinery":"The load-bearing object is The Safety Conception, a constitutive criterion: a research project is AI safety research just in case it aims to prevent or reduce harms from AI systems' development and deployment. The argument's machinery is the two-purpose evaluative framework borrowed from conceptual engineering and ameliorative inquiry: any concept of AI safety is assessed by whether it (A) unifies paradigmatic research programs under one field and (B) conduces to reducing AI harms. The comparison is carried out through three ideal-typical scenarios—The Safe Scenario, The Catastrophic Scenario, and The Engineering Scenario—which differ in how research is prioritized and which experts are consulted; The Safety Conception corresponds to the scenario that prioritizes purely by expected harm reduction.","core_discovery":"The paper's central claim is that The Safety Conception is the best conception of AI safety and therefore ought to be the operative concept, not merely the stated one. Using conceptual engineering, it evaluates candidate concepts against two purposes: (A) providing a unifying explanation of why paradigmatic AI safety research programs belong to the same field, and (B) being conducive to reducing the harms caused by AI systems. It argues that The Safety Conception does both better than The Catastrophic Conception and The Engineering Conception, and that any more restrictive or more permissive departure fares worse. If the argument is right, AI safety includes work on social harms such as bias, misinformation, privacy, and economic harms alongside work on catastrophic harms, and researchers in both areas should be integrated, including in political conversations about AI safety.","pith_inferences":["Not pursued in the paper but testable: if funders and venues adopted the broad conception, the portfolio of funded AI safety research should measurably shift toward sociotechnical and governance work; comparing funders with broad versus narrow definitions could test the paper's causal premise.","The same conceptual-engineering method could be pointed at adjacent categories such as AI ethics or responsible AI, which would dissolve the safety/ethics split from the other direction and could produce different policy alliances.","The paper assumes that harm-reducing merit is comparable between social and catastrophic risks; one implicit task for future work is a shared metric or decision procedure for comparing them, since the argument does not supply one.","If the paper is right, marginalization of social-harm research is partly a labeling effect; an observable extension is to audit which researchers are invited into AI safety policy settings and whether invitations track the broad harm taxonomy."],"forward_implications":["Research on algorithmic bias, misinformation, privacy, and economic harms qualifies as AI safety, so it should be presented and published in the same venues as research on catastrophic risks.","AI labs should not maintain separate ethics and safety teams, and funders should support harm-reduction research that is not framed in terms of catastrophic or existential risk.","Political deliberations about AI safety should include researchers working on social harms as well as those working on catastrophic harms.","Requiring work on present systems to justify itself by future catastrophic relevance is a form of gatekeeping that the paper predicts makes the field less effective at reducing harm.","Because the same model behavior (for example, generating hate speech and generating bomb instructions) can produce both social and catastrophic harms, drawing a hard line between them forfeits explanatory continuity."],"supporting_citations":[{"why":"Establishes an early technical agenda for AI safety framed as reducing unintended and harmful behavior, used as evidence that The Safety Conception is at least a manifest concept.","marker":"Amodei et al. (2016)"},{"why":"Supplies the conceptual engineering methodology of assessing and improving representational devices.","marker":"Cappelen (2018)"},{"why":"Introduces ameliorative inquiry, the method of assessing concepts by the purposes they ideally serve, grounding purpose (B).","marker":"Haslanger (2005)"},{"why":"Supplies the sociological premise that disciplinary boundaries shape reading, collaboration, evaluation, and policy influence.","marker":"Trowler (2012)"},{"why":"Exemplifies The Catastrophic Conception by defining ML safety through long-term and long-tail risks, a rival the paper argues against.","marker":"Hendrycks et al. (2021)"},{"why":"Provides a taxonomy of generative AI harms and a sociotechnical safety-engineering framing, used both as a rival conception and as evidence of harm categories.","marker":"Weidinger et al. (2023)"},{"why":"Gives the standard definition of safety engineering with hazard identification, risk analysis, and risk management, which The Engineering Conception builds on.","marker":"Roland and Moriarty (1990)"},{"why":"Characterizes the AI Safety Problem in terms of threats to human survival or permanent disempowerment, representing the catastrophic narrowing the paper rejects.","marker":"Dalrymple et al. (2024)"}],"fun_headline_variants":["AI safety is more than catastrophe prevention","Rethink AI safety: all harms count","Widen AI safety to bias and privacy","AI safety isn't just existential risk"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands on the assumption that the only tests that matter for a definition of AI safety are how well it unifies the field and how much it helps reduce AI harms, and that labels shape what gets funded and who gets heard.","fun_headline_variants_meta":{"raw":{"variants":["AI safety is more than catastrophe prevention","Rethink AI safety: all harms count","Widen AI safety to bias and privacy","AI safety isn't just existential risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1339,"prompt_tokens":962,"completion_tokens":377,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":323}},"tokens_in":578,"tokens_out":377,"duration_ms":4158,"temperature":1.0,"reasoning_tokens":323,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:55:22.275070+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: compare funding and policy-consultation records before and after a funder or government body adopts the broad definition; if the set of supported projects and invited experts does not shift toward social-harm research, the practical argument for The Safety Conception loses force.","supporting_citations":[],"review_version":1}