{"id":"e6a80e83-4ea4-4604-a4d9-d30e50e564ec","arxiv_id":"2506.18932","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper argues that AI safety targets unintentional failures while AI security targets intentional attacks, and that both dimensions must be integrated to achieve trustworthy AI.","lead":"This paper proposes intent-based definitions that separate AI safety (preventing accidental harm) from AI security (stopping intentional attacks). The distinction is meant to help researchers, funders, and policymakers deploy clearer and safer AI risk management.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Definitional ambiguity about whose intent counts undermines the claimed precise boundary between AI safety and AI security.","rationale":"The paper offers a clear and useful synthesis of a widely discussed distinction, and its case studies and analogy are pedagogically helpful. However, the central claim's strength depends on Definitions 1 and 2 being precise enough to draw boundaries. The text never defines the bearer of intent: the AI system, the developer, or the adversary. This is not merely a practical measurement problem; it makes the classification of a single event underdetermined. For example, a prompt injection involves an output that is both unintended by the system operator (so Definition 1 applies) and intentionally crafted by an adversary (so Definition 2 applies). The paper acknowledges this overlap in Section 4.2, which weakens its Section 4.1 claim that the distinction is primary and its abstract claim of 'precise research boundaries.' Since the paper is conceptual and makes no empirical or formal claims, the fix is straightforward: explicitly state whose intent is the classifier and add an operational rule for ambiguous cases. The reader's conditional verdict already captures the need for this clarification, so we do not change the verdict.","tokens_in":10886,"tokens_out":7716,"duration_ms":80801,"concrete_test":"Analytical re-derivation: take the three canonical cases from Section 4.2 (prompt injection of an LLM, hijacking of an autonomous vehicle, and exploitation of a known model bias) and classify each using Definitions 1 and 2 with 'unintended' replaced first by 'unintended by the AI system' and then by 'unintended by the system operator'. If any case changes classification across the two readings, the definitions are under-specified and the claimed precise boundary (Sections 1 and 4.1) is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section 4.1) is that the primary distinction between AI Safety and AI Security lies in the origin and nature of the risk: safety handles unintentional behaviors, security handles intentional adversarial actions. This requires that 'unintended' in Definition 1 and 'intentional' in Definition 2 refer to a well-defined, singular attribution of intent. The paper never specifies whose intent matters. Definition 1 says AI Safety avoids 'unintended harmful outcomes'; if read literally as the AI's own intent, the paper's own safety concern about power-seeking AI (Section 5.1, [14]) violates it, since a power-seeking agent intentionally subverts shutdown. If read as the developer's intent, then a user deliberately exploiting a known bias to get harmful output is also 'unintended by the developer', yet Definition 2 classifies it as security because of the user's intent. Consequently, a single event can satisfy both definitions: the system's output is unintended by the operator (safety failure) and intentionally elicited by an adversary (security failure). Section 4.2 explicitly concedes that 'misuse often lies at the intersection' of unintentional system flaws and intentional exploitation. Thus the claimed precise boundary is not derivable from the definitions; it rests on an unstated and untested rule for attributing intent. The paper's own examples therefore undercut the central claim's precision.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a conceptual/taxonomic contribution arguing that AI Safety and AI Security should be distinguished by the origin and nature of the risk. It proposes intent-based definitions (Definition 1: AI Safety as avoiding unintended harmful outcomes despite uncertainties; Definition 2: AI Security as resilience against intentional attacks on data, algorithms, or operations), reviews the historical broadening of the term \"AI Safety,\" illustrates the distinction with message-transmission and building analogies, maps distinct research agendas and protective mechanisms, and concludes that the two domains are interdependent and should be managed within a unified risk framework. The authors explicitly acknowledge in Section 4.2 that misuse often lies at the intersection of unintentional system flaws and intentional exploitation.","tokens_in":11057,"tokens_out":5669,"duration_ms":58716,"significance":"If the proposed taxonomy could be made precise, it would help align research agendas, funding, and regulation with distinct classes of AI risk, and the paper is a readable synthesis of the ongoing safety-security debate. Its strengths are a clear organization, a substantial set of references to relevant literature, and an honest acknowledgment of edge cases. As it stands, however, the claimed \"precise boundaries\" are not supported because the central notion of intent is ambiguous and the two definitions are not mutually exclusive for the misuse cases the paper itself identifies. The contribution is therefore currently a useful heuristic rather than a rigorous delineation.","major_comments":[{"comment":"The central distinction rests on intent, but the manuscript never states whose intent is the relevant one. Definition 1's \"unintended\" is ambiguous among the AI system's own goals, the developer's or operator's intent, and the user's intent, while Definition 2's \"intentional attacks\" refers to an adversary. These readings diverge: a power-seeking AI (Section 5.1, reference [14]) deliberately resists shutdown, so under a system-intent reading it is not a safety failure; conversely, a user deliberately exploiting a known bias (the kind of case mentioned in Section 4.2) is unintended by the developer but intentionally caused by the user, so the same event satisfies both definitions. The paper needs an explicit attribution rule for intent, or a classification procedure, before the claimed \"precise boundary\" can be supported.","section":"§3, Definitions 1 and 2; §4.1"},{"comment":"The paper concedes that \"misuse often lies at the intersection\" of unintentional system flaws and intentional exploitation, yet provides no criterion for resolving such cases. Definitions 1 and 2 are not mutually exclusive as stated: a prompt-injection attack that succeeds because of an alignment flaw is simultaneously an unintended harmful outcome of the system and an intentional attack by an adversary. Section 4.1's examples do not cover this case, so the proposed boundary cannot classify the central phenomenon—AI misuse—that the paper says it clarifies.","section":"§4.2"},{"comment":"Section 5.1 lists \"power-seeking behaviors\" and \"misaligned objectives\" under AI Safety, but these are goal-directed behaviors rather than accidental system failures. This conflicts with Definition 1's characterization of safety as preventing \"unintended harmful outcomes\" and indicates a conflation between \"not intended by the designers\" and \"accidental system malfunction.\" The authors should clarify that safety includes failures of alignment even when the system competently pursues a misaligned goal, or they should restrict the definition accordingly.","section":"§3.1 and §5.1"}],"minor_comments":[{"comment":"The word \"illustratred\" should be \"illustrated.\"","section":"§7"},{"comment":"In the paragraph on autonomous vehicles, \"A Vs\" should be \"AVs.\"","section":"§6"},{"comment":"The definition attributed to \"Bengio et al.\" is a quotation from the International AI Safety Report [6], which is an institutional multi-author report; the attribution should be clarified.","section":"§3.2"},{"comment":"The message-transmission analogy treats safety failures as stochastic noise, but many AI safety failures, such as misgeneralization or goal misalignment, are not well modeled by random corruption; the analogy should be explicitly labeled as heuristic only.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more of a position paper than a technical research contribution. For a cs.CY venue that accepts conceptual synthesis, the revised version could be a useful reference, but the incremental contribution over Qi et al. [39] should be stated more explicitly, and the paper's main claim of \"precise boundaries\" needs either to be operationalized or softened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful conceptual paper that sharpens a distinction the field already half-has. The definitions are clean enough to spark productive debate, but the abstract's promise of 'precise research boundaries' is stronger than what the definitions actually support.\n\nWhat's new: most of the substance is a crisper restatement of Qi et al. (2024)'s call to keep safety and security separate, with formal-looking definitions, a mapping of research topics to each side, and genuinely pedagogical analogies (checksum vs. MAC, building structure vs. locks). The paper is honest that the two domains interlock: security breaches cause safety failures, and safety flaws become attack vectors. That is a fair and useful synthesis, not a breakthrough.\n\nSoft spots: the core definitions hinge on 'unintended' versus 'intentional,' but the paper never specifies whose intent counts. Definition 1 says AI Safety avoids 'unintended harmful outcomes'—unintended by the developer? The user? The AI itself? This matters. A malicious user deliberately exploiting a known bias produces an outcome unintended by the developer (so safety) and intentional by an adversary (so security). The paper explicitly concedes misuse sits at the intersection (Section 4.2), yet the abstract promises precise boundaries. That is an overclaim. The definitions are informal, with no operational rule for ambiguous cases. For a conceptual paper this is not fatal, but it should be fixed by pinning intent to the system's designers/operators and explicitly acknowledging residual ambiguity.\n\nMinor: the related-work discussion is fair, and the citation pattern looks clean. The paper is a synthesizing essay, not an empirical or formal contribution, so judge it on clarity and framing rather than novelty.\n\nWho it's for: researchers and policy people needing a shared vocabulary. I'd consider bringing it to a reading group to argue about edge cases, and I might cite it as a framing reference, though not as a technical result.\n\nRecommendation: send it to peer review. With modest revisions—tone down 'precise' and specify the intent attribution—it could be a solid reference piece.","headline":"Useful, clearly-written conceptual synthesis, but the claimed precision overreaches because the paper never pins down whose intent distinguishes safety from security.","tokens_in":11602,"tokens_out":2127,"would_cite":false,"duration_ms":22444,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI safety vs. AI security: intent draws the line","keywords":["AI safety","AI security","intentional vs unintentional harm","adversarial attacks","AI risk taxonomy","trustworthy AI","AI misuse"],"falsifier":"A documented AI incident in which the same harmful output arises simultaneously from an unintended design flaw and a deliberate exploit, and where a panel of safety and security experts cannot agree on which category it belongs to, would show the intent axis is not precise enough to carry the claimed boundary.","tokens_in":10632,"feed_emoji":"🛡️","tokens_out":3167,"duration_ms":29712,"temperature":0.7,"pith_summary":"This paper argues that AI Safety and AI Security are not interchangeable terms but distinct research domains separated by the origin of the risk: AI Safety addresses accidental, unintended harm from system flaws, misalignment, or distributional shifts, while AI Security addresses intentional adversarial actions such as data poisoning, evasion attacks, prompt injection, and model theft. The authors propose formal definitions of both concepts, show how they are interdependent, and argue that keeping them distinct is essential for focused research, effective regulation, and trustworthy AI. If the distinction holds, it gives regulators and funders a clean way to route risks to the right mitigation tools.","feed_headline":"AI safety vs. AI security: intent draws the line","feed_subtitle":"A new framework says accidents belong to safety, deliberate attacks to security, and both are needed for trustworthy AI.","key_machinery":"The argument is carried by an intent-based dichotomy formalized in two named definitions: Definition 1 for AI Safety and Definition 2 for AI Security. These definitions do the classificatory work throughout the paper, supported by two analogies: the communication model where a checksum (CRC) detects accidental corruption while a message authentication code (MAC) resists deliberate tampering, and the building analogy where structural integrity maps to safety while locks, alarms, and perimeter defenses map to security. The intent axis is used to classify research topics, case studies, and the relationship between the two domains.","core_discovery":"The paper's central claim is that the primary distinction between AI Safety and AI Security lies in the origin and nature of the risk: safety concerns accidental or unintended behaviors, while security concerns intentional adversarial actions. Definition 1 defines AI Safety as the property of avoiding unintended harmful outcomes despite uncertainties in inputs, goals, training data, or deployment conditions. Definition 2 defines AI Security as the property of remaining resilient against intentional attacks on data, algorithms, or operations, preserving confidentiality, integrity, and availability. The paper argues that misuse cases often sit at the intersection of these two categories because safety flaws can be exploited by adversaries, and security breaches can produce safety failures, yet the intent-based boundary is still the correct organizing axis for the field.","pith_inferences":["A testable prediction: expert disagreement over classifying ambiguous incidents, such as a model that generates harmful content both because of training-data bias and because of a jailbreak, will persist, suggesting the boundary is a spectrum rather than a clean partition.","A practical extension: AI incident reporting could tag every event on two independent axes—cause (accidental vs. intentional) and harm type—so that the field inherits the safety/security separation that aviation and nuclear power have long used.","The intent axis could be operationalized empirically: if an incident is fully resolved by retraining or re-specification it leans safety; if it is resolved only by authentication, filtering, or access controls it leans security.","The paper's framework implies that a single AI system may require separate assurance cases for safety and security, each with its own evidence and ownership, which current practice rarely distinguishes."],"forward_implications":["Research funding and agendas can be split by threat origin: test-and-align work goes to safety; adversarial robustness, access control, and incident response go to security.","Policy tools can be matched to risk type: pre-deployment testing and ethical standards for safety; cybersecurity standards, red teaming, and breach notification for security.","Systems need both properties: a safe but insecure system can be hacked, while a secure but unsafe system can still harm through bias or misalignment.","Unified AI risk management must trace chains that cross the boundary, such as a security attack like prompt injection that defeats a safety filter and produces harmful output."],"supporting_citations":[{"why":"Supplies the canonical framing of AI safety as preventing accidents in machine learning systems and explicitly treats protection against malicious actors as a distinct, related field.","marker":"[2]"},{"why":"Provides the security-engineering definition centered on intentional adversaries, which Definition 2 builds on.","marker":"[3]"},{"why":"The broad International AI Safety Report usage that the paper argues against, and also the source of the formal definition of AI security quoted in Section 3.2.","marker":"[6]"},{"why":"Load-bearing example of an indirect prompt injection attack that subverts an LLM's safety, used to illustrate security failures causing safety failures.","marker":"[22]"},{"why":"The experimental car-hacking study used as the prime example of a security breach (remote vehicle control) precipitating a safety failure (crash).","marker":"[29]"},{"why":"The prior position paper that motivates the call for delineation and frames the unified risk-management approach the authors adopt in Section 7.","marker":"[39]"}],"fun_headline_variants":["AI safety vs security: accidents vs attacks","Intent draws the line between AI safety and security","Safety for mishaps, security for malice in AI","AI safety: errors, AI security: enemies","Threat origin separates AI safety and security"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that harmful AI events can be reliably classified by whether the cause was accidental or intentional, even though it concedes that many misuse cases sit at the intersection of the two.","fun_headline_variants_meta":{"raw":{"variants":["AI safety vs security: accidents vs attacks","Intent draws the line between AI safety and security","Safety for mishaps, security for malice in AI","AI safety: errors, AI security: enemies","Threat origin separates AI safety and security"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001358,"raw_usage":{"total_tokens":5454,"prompt_tokens":832,"completion_tokens":4622,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":4552}},"tokens_in":448,"tokens_out":4622,"duration_ms":31601,"temperature":1.0,"reasoning_tokens":4552,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:00:18.804296+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A documented AI incident in which the same harmful output arises simultaneously from an unintended design flaw and a deliberate exploit, and where a panel of safety and security experts cannot agree on which category it belongs to, would show the intent axis is not precise enough to carry the claimed boundary.","supporting_citations":[{"cited_title":"Anderson.Security Engineering: A Guide to Building Dependable Distributed Systems","cited_arxiv_id":null,"evidence_quote":"Provides the security-engineering definition centered on intentional adversaries, which Definition 2 builds on."},{"cited_title":"International AI safety report: The international scientific report on the safety of advanced AI","cited_arxiv_id":null,"evidence_quote":"The broad International AI Safety Report usage that the paper argues against, and also the source of the formal definition of AI security quoted in Section 3.2."},{"cited_title":"Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection","cited_arxiv_id":null,"evidence_quote":"Load-bearing example of an indirect prompt injection attack that subverts an LLM's safety, used to illustrate security failures causing safety failures."},{"cited_title":"Experimental security analysis of a modern automobile","cited_arxiv_id":null,"evidence_quote":"The experimental car-hacking study used as the prime example of a security breach (remote vehicle control) precipitating a safety failure (crash)."},{"cited_title":"Su, Mengdi Wang, Chaowei Xiao, Bo Li, Dawn Song, Peter Henderson, and Prateek Mittal","cited_arxiv_id":null,"evidence_quote":"The prior position paper that motivates the call for delineation and frames the unified risk-management approach the authors adopt in Section 7."}],"review_version":1}