{"id":"3873aa7a-9f49-49ea-bf90-d02cf0d10706","arxiv_id":"2607.25648","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Safety governance designed for narrow AI presupposes bounded tasks; general-purpose AI's unbounded, fluent outputs break those assumptions, so policing frameworks risk failing unless rebuilt around GPAI-specific evidence.","lead":"This policy paper argues that current public-service AI governance frameworks, built for narrow task-specific AI, cannot assure the safety of general-purpose AI in policing because its unbounded outputs undermine accuracy, bias, explainability, and accountability. It recommends a taxonomic separation of narrow versus general-purpose AI, preference for narrow tools, a deployment pause, and a national safety body.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing premise is that GPAI model-level properties cannot be constrained by system design; the paper asserts this but does not test it. If a narrow interface can restore measurable accuracy, bias, explainability, and oversight, the blanket pause is overbroad.","rationale":"The reader's weakest_assumption accurately identifies the central empirical/structural claim: model-level dispositions are said to persist in every system built on the model (Section 1), and the collapse of safety conditions is said to be unrecoverable by system design (Section 4.5). My stress test agrees. The paper's own evidence supports risk, not impossibility. A concrete deployment study with a constrained-output GPAI pipeline would test whether task-level accuracy, demographic error-rate analysis, faithful explanation, and effective human oversight can be restored. If they can, the recommendations in Section 6 (especially 6.1's model-based taxonomy and 6.3's blanket pause) are overbroad. If they cannot, the paper's central claim is strengthened. Since the paper does not currently engage with this counterfactual, the conditional verdict stands: accept with revisions that scope the claims or provide this evidence.","tokens_in":18158,"tokens_out":4956,"duration_ms":52227,"concrete_test":"Conduct a deployment study on one concrete policing task, e.g., triage of incident reports into priority categories. Implement (A) a GPAI system constrained by a narrow interface: fixed output schema, retrieval-augmented grounding, deterministic rule-based post-validation; and (B) a narrow ML classifier on the same labelled data. On held-out data, compute task-specific accuracy/F1, demographic error-rate disparities, explanation faithfulness (do stated rationales track actual drivers, tested by counterfactual perturbation), and human-reviewer error-detection rate. If (A) matches or exceeds (B) on all four, model-level generality is not necessarily fatal to system-level safety assurance; if (A) fails, the paper's premise is supported for this task.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 1 makes the model the unit of analysis: 'Our concern lies with the model... They travel with the model into every system built upon it.' Sections 4.1–4.4 then convert model-level tendencies into claims that all GPAI-based systems have unmeasurable accuracy, structural bias, illusory explanation, and eroded oversight. Section 4.5 states the collapse is 'not as a side-effect that better design could restore, but as a direct consequence.' This is the load-bearing premise. It is not established. A deployment can bound the task: constrained output schemas, closed label sets, deterministic post-processing, retrieval-augmented generation with external verification, and task-specific evaluation protocols. Even if the underlying model remains general, the safety-relevant behavior of the deployed system is a property of the interface, prompt, and workflow, not only of the model. The cited evidence (unfaithful chain-of-thought, implicit bias, persuasive-but-inaccurate outputs) shows model-level risks exist; it does not show system-level mitigations cannot re-establish measurable accuracy, demographic error-rate comparisons, faithful explanations, or meaningful human oversight for specific tasks. Without demonstrating this, Recommendation 6.3's blanket pause and Recommendation 6.1's model-based taxonomy are overbroad. The paper could be charitably read as arguing only that current frameworks are inadequate; the stronger structural-impossibility claim is the part that needs support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This policy paper argues that general-purpose AI (GPAI) built on large language models undermines the four safety properties that public-service governance frameworks for narrow AI rely upon: accuracy, bias, explainability, and accountability. It develops this through policing, contending that GPAI's unbounded task scope, open-ended natural-language output, learned bias, and fluent rationalization invert these properties; that current frameworks' prescriptions (expert evaluation, human-in-the-loop, high-stakes distinctions) presuppose narrow-AI conditions; and that safety assurance therefore shifts from intrinsic to optional. It recommends a taxonomy separating narrow and general-purpose AI, technological parsimony, a pause on operational deployment of GPAI in policing until adequate evidence exists, and a centralized independent safety authority.","tokens_in":18455,"tokens_out":4560,"duration_ms":52402,"significance":"The paper addresses a genuinely important and timely policy question, and its core observation—that governance frameworks designed around narrow, task-bounded AI are poorly matched to GPAI—is well grounded in a careful review of recent empirical literature. It is transparent in its argument structure, engages seriously with the socio-technical context of policing, and makes concrete, falsifiable recommendations. Its main strength is the systematic articulation of how accuracy measurement, bias evaluation, and explainability become qualitatively harder when output is open-ended natural language rather than a bounded classification or score. If the structural-inversion thesis survives scrutiny, the paper would be a significant contribution to public-sector AI governance. However, the strongest form of the thesis—that GPAI's safety-relevant properties cannot be re-bounded by system-level design—is asserted rather than demonstrated, and the recommendations depend on that strong form.","major_comments":[{"comment":"The load-bearing premise is that GPAI model-level properties 'travel with the model into every system built upon it' (§1) and that the collapse of safety assurances is 'not as a side-effect that better design could restore, but as a direct consequence' (§4.5). This is asserted, not tested. The paper does not engage with system-level mitigations that can constrain task scope: constrained output schemas, closed label sets, deterministic post-processing, retrieval-augmented generation with external verification, task-specific evaluation protocols, or selective/exception-based human audit. If any of these can re-bound task scope and restore measurable accuracy, bias disaggregation, or meaningful oversight, then Recommendation 6.3 (blanket pause) and the 'model, not the wrapper' taxonomy of §6.1 are overbroad. Either supply evidence that such controls cannot restore assurability, or weaken th","section":"§1, §4.5, §6.1"},{"comment":"The claim that for general-purpose AI 'neither exists'—no agreed criteria, no established methodology, no accumulated body of findings, and no practitioners—is too absolute and conflicts with the paper's own citations. The paper cites evaluation frameworks (Weidinger et al. 2023), social-science measurement approaches (Wallach et al. 2025), safety-evaluation statements (AI Safety Institute 2024), and emerging domain benchmarks. The correct point is that these are immature and not yet adequate for high-stakes deployment, which is different from their wholesale absence. This overstatement makes the critique of expert evaluation vulnerable to a strawman and undercuts the paper's own recommendation for building evaluation infrastructure.","section":"§5.1"},{"comment":"The explainability argument correctly notes that chain-of-thought rationalizations are often unfaithful, but it overgeneralizes from 'explanations are frequently unreliable' to 'explanations are indistinguishable from genuine reasoning' and from this to the claim that human-in-the-loop oversight is 'conceptually incoherent' for analytical tasks. Oversight designs that treat AI output as a hypothesis to be independently checked for key facts, or that use selective audit on a sample of high-impact outputs, can be meaningful even when full replication is impossible. The paper does not consider such designs, which are directly relevant to the claim in §5.2 that the division of labor collapses for all GPAI analytical tasks.","section":"§4.3, §5.2"},{"comment":"The causal claim that the collapse of safety assurances follows 'as a direct consequence' of generality, accessibility, and low cost conflates intrinsic model properties with contingent socio-technical conditions. Accessibility and low deployment cost are shaped by regulatory, procurement, and interface decisions; they are not immutable properties of the model class. The paper's own recommendation 6.4, which would regulate access and require evaluation, implicitly concedes that these conditions can be altered. This tension weakens the structural-impossibility framing and should be resolved in favor of a more precise distinction between model-level properties and deployment conditions that are amenable to policy.","section":"§4.5"}],"minor_comments":[{"comment":"The model/system distinction is defined but then deliberately blurred ('use \"system,\" \"tool,\" and \"application\" in their ordinary sense throughout'), which can confuse the reader when later sections claim properties of the model apply to all systems built on it. The terminology should be used more consistently.","section":"§1"},{"comment":"The statement 'To the best of our knowledge, no adequate mitigations for these challenges have yet emerged' is an absence claim that would be stronger with a systematic search protocol or a more precise scoping of the mitigation literature reviewed.","section":"§4 intro"},{"comment":"The four safety properties are introduced as 'consistently understood' across governance frameworks, but only some of the bullet points carry citations. Adding references for the explainability and accountability definitions would strengthen the framing.","section":"§2.1"},{"comment":"The claim that feature-attribution methods 'can be sidestepped entirely by choosing architectures interpretable by design' (Rudin 2019) is presented as if it applies readily to narrow policing systems; the feasibility of fully interpretable models for complex policing tasks deserves more nuance.","section":"§2.2"},{"comment":"The phrase 'accuracy cannot be quantified over unbounded outputs' should be qualified: for specific bounded sub-tasks (e.g., structured field extraction or constrained classification) accuracy can be quantified even with a GPAI backend. The opening of §4.1 is stronger than the paper's own later nuances about construct validity.","section":"§4.1"},{"comment":"The reference to 'AI Security Institute' after 'AI Safety Institute' should be checked for accuracy and consistency, since the renaming is the kind of factual detail that reviewers cannot verify from the manuscript alone.","section":"§4.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious, well-written policy analysis with a clear thesis and a strong evidence base. My main concern is that the strong structural-impossibility claim in §4.5—that system-level design cannot restore safety assurances—is the load-bearing element for the pause recommendation, and it is asserted rather than established. The manuscript would be acceptable for publication after the authors either provide evidence against system-level re-bounding or moderate the thesis and correspondingly adjust the recommendations. I would not reject the paper; the topic is important and the existing argument is already a useful contribution even in its weaker, conditional form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read.\n\nThe paper earns its keep on the strength of the framing: four properties (accuracy, bias, explainability, accountability) that narrow AI made tractable, and a clear argument for why each inverts under GPAI. The policing case grounds it, and the discussion of human-in-the-loop being incoherent for open-ended analytical work is the strongest part—if a reviewer can't reproduce the analysis, oversight is performative. The high-stakes critique is also good.\n\nThe citation base is solid and current: they engage with Turpin et al. on unfaithful chain-of-thought, Bai et al. on implicit bias, Hackenburg et al. on persuasion/accuracy tradeoff, and Rein's METR piece. No obvious cherry-picking; the paper is transparent about the open questions.\n\nThe soft spot is the load-bearing premise. The authors take the model as unit of analysis and assert that model-level properties 'travel with the model into every system built upon it.' That's asserted, not shown. They don't wrestle with the standard counterargument: a constrained interface—closed label sets, output schemas, deterministic post-processing, retrieval with external verification, task-specific evaluation—can re-bound an otherwise general model. If that's possible, accuracy and oversight can be restored for a specific task, and the 'structural inversion' becomes a contingent claim about current practice, not a consequence of the architecture. The cited studies show model risks exist; they don't show system-level mitigation is impossible. This matters because the pause recommendation (6.3) and the model-based taxonomy (6.1) inherit this premise. The paper would be stronger if it framed the claim as a burden-of-proof argument: current frameworks give no evidence that wrappers work, so pause until the evidence exists. That's defensible. The stronger impossibility claim needs support.\n\nAlso, the 'no adequate mitigations have yet emerged' line is a strong universal negative supported by selected citations, not a systematic review. Minor, but worth saying.\n\nOverall: this is a real contribution to the AI governance conversation and deserves peer review. It needs a revision that either softens the impossibility claim or defends it against system-level bounding. I'd send it to referees, and I'd bring it to a reading group.","headline":"Relins & Birks offers a sharp structural-inversion argument about GPAI governance in policing, but the blanket pause rests on an unproven claim that model risks cannot be bounded by system design.","tokens_in":18908,"tokens_out":3866,"would_cite":true,"duration_ms":43027,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that general-purpose AI undoes the four properties that made earlier policing AI governable—measurable accuracy, disaggregable bias, real explanation, clear accountability—so current governance frameworks and their mitigati","keywords":["general-purpose AI","large language models","AI governance","policing","safety assurance","algorithmic accountability","explainability","public services"],"falsifier":"A controlled field evaluation of a GPAI system in a policing task where the model is locked to a constrained output template, outputs are scored against verified ground truth, error rates are disaggregated by demographic group, and independent reviewers reproduce the analysis from source material—if the system met pre-specified accuracy and fairness thresholds and reviewers caught material errors at high rates, it would undercut the paper's claim that GPAI cannot be assured. Conversely, finding that reviewers routinely fail to detect fabricated details in fluent summaries would support it.","tokens_in":18036,"feed_emoji":"🚔","tokens_out":6824,"duration_ms":64597,"temperature":0.7,"pith_summary":"The paper is trying to establish that general-purpose AI (GPAI) is not just a more capable version of the AI policing has governed for years, but a different category whose defining features—generality, accessibility, and low deployment cost—invert the four safety properties that made narrow AI manageable: accuracy, bias, explainability, and accountability. Because those properties cannot be measured or enforced for open-ended language models, the mitigations that governance frameworks recommend (expert evaluation, human-in-the-loop, high-stakes thresholds) fail. The paper develops this through policing, where a hallucinated detail or biased summary can lead to wrongful intervention, and argues the same failure will recur across other public services. It concludes with four recommendations: distinguish narrow from general-purpose AI in governance, prefer narrow tools where viable, pause operational deployment until domain-specific evaluation methods and evidence exist, and build a centralized independent safety authority. If the paper is right, current policy promises to 'do AI safely' in policing are largely ceremonial.","feed_headline":"General-purpose AI collapses the safety pillars policing built on","feed_subtitle":"Policing can't measure accuracy, bias, explanation, or accountability, so the paper urges a pause.","key_machinery":"The load-bearing distinction is between the GPAI model (the trained artefact) and the GPAI system (the model wrapped in an interface); the paper argues the risk properties belong to the model class and travel with the model into every system built upon it, so a tailored interface cannot convert a GPAI deployment into a narrow, assureable tool. The argument is carried by four paired inversions—accuracy becomes unmeasurable, bias structural, explanation illusory, accountability eroded—and by the claim that the two standard mitigations (expert evaluation, human-in-the-loop) each presuppose a bounded task and a known standard of correctness. The unit of analysis is the model, and the conclusion","core_discovery":"The paper's central claim is that the very properties that make GPAI transformational—its generality (any task describable in natural language), accessibility (no specialised engineering), and near-zero marginal deployment cost—undermine the conditions under which AI safety has historically been pursued. For narrow AI, safety was tractable because a task was fixed in advance, so accuracy was measurable, bias was a quantifiable disparity, explainability could be approximated by feature attribution, and accountability had a clear boundary at the tool's output. GPAI reverses each: accuracy lacks a settled standard of correctness for open-ended text; bias becomes a structural property absorbed f","pith_inferences":["The paper's taxonomic distinction suggests a practical test for governance: any system that uses a GPAI model for even one step in a processing chain should be governed as GPAI—this can be extended to procurement rules requiring disclosure of model provenance.","The argument likely generalizes beyond policing to other public services that already use generative AI for drafting, summarization, or triage; the same inversion of safety properties would apply to social care, healthcare administration, and benefits decisions.","A testable extension would be an empirical study measuring whether a GPAI system with a constrained output schema (fixed fields, no free text) can restore measurable accuracy and disaggregable bias—if it can, the paper's model-level claim would need to be narrowed to open-ended applications.","The paper's emphasis on the 'appearance of explanation' implies that transparency regulations requiring users to be told when an output is AI-generated may be insufficient; the harder requirement would be to show the output can be independently verified against a standard, not merely flagged."],"forward_implications":["If accuracy cannot be quantified for open-ended outputs, no pre-deployment evaluation can certify a GPAI system as fit for a specific policing task.","Because hidden bias can persist even when explicit bias benchmarks pass, standard fairness audits for GPAI are insufficient evidence of non-discrimination.","The fluent explanations produced by GPAI models are not reliable windows into their reasoning, so human reviewers cannot meaningfully verify AI-generated analyses without reproducing the work themselves.","A high-stakes threshold offers false comfort, since 'low-stakes' applications can cumulatively reshape institutional knowledge and decisions in ways that are difficult to audit.","Deployment before evaluation infrastructure exists is likely irreversible: once embedded, sunk costs and dependence make withdrawal unlikely, so a pause and a centralized evidence-gathering authority are necessary preconditions."],"fun_headline_variants":["GPAI shatters policing's safety benchmarks, paper urges pause","General-purpose AI breaks policing's safety assumptions","Why GPAI makes police AI governance impossible","Policing can't measure AI safety once AI goes general-purpose"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the risky properties of general-purpose AI—unbounded task scope, open-ended output, fluent rationalization—belong to the model itself and persist in every system built around it, so interface constraints or guardrails cannot restore the conditions for safety.","fun_headline_variants_meta":{"raw":{"variants":["GPAI shatters policing's safety benchmarks, paper urges pause","General-purpose AI breaks policing's safety assumptions","Why GPAI makes police AI governance impossible","Policing can't measure AI safety once AI goes general-purpose"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000473,"raw_usage":{"total_tokens":2224,"prompt_tokens":820,"completion_tokens":1404,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":1339}},"tokens_in":564,"tokens_out":1404,"duration_ms":11035,"temperature":1.0,"reasoning_tokens":1339,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:48:40.742479+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled field evaluation of a GPAI system in a policing task where the model is locked to a constrained output template, outputs are scored against verified ground truth, error rates are disaggregated by demographic group, and independent reviewers reproduce the analysis from source material—if the system met pre-specified accuracy and fairness thresholds and reviewers caught material errors at high rates, it would undercut the paper's claim that GPAI cannot be assured. Conversely, finding that reviewers routinely fail to detect fabricated details in fluent summaries would support it.","supporting_citations":[],"review_version":1}