{"id":"6845e957-8c86-4127-8dce-901d776e3d75","arxiv_id":"2608.05418","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A mixed-stakeholder UK workshop rated 13 AI policing use cases, rejecting recidivism risk assessment outright while accepting most others conditionally, and found that a racial-equity focus broadened, not narrowed, the questions asked.","lead":"A one-day workshop brought 30 police officers, community representatives, and academics together to classify the risk of 13 AI policing tools, with racial bias as the stated focus. The study reports that participants accepted most tools conditionally but rejected recidivism risk assessment, and that the racial-equity framing broadened rather than narrowed the debate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'did not narrow' claim is comparative, but the study has no comparison condition, and the process reconstruction rests on unrecorded observation; the claim should be softened or replicated with a framed-control arm.","rationale":"The reader's conditional verdict is appropriate. The paper's descriptive contributions are solid: 76 worksheets, transparent method, explicit limitations, and risk classifications; the result that recidivism risk assessment drew premise-level objections is anchored in worksheet quotes. The load-bearing weakness is the comparative framing of the central finding. Section 5 says the explicit framing was racial bias yet reasoning was broader; the abstract and conclusion convert this into a claim that foregrounding racial equity 'did not narrow' and 'deepened' deliberation. That is a causal/comparative claim. With only one framing condition, the observation that groups discussed broader questions cannot establish that the racial-equity framing caused, or failed to restrict, the scope of reasoning. A group asked to assess whether a tool works, delivers genuine benefit, and benefits everyone would likely produce similar questions without any equity framing; the study provides no comparison. The process reconstruction is also memory-based: discussions were not recorded, and the back-engineered questions were produced by authors present and cross-checked by other authors, not against transcripts. This makes the central process finding inherently unverifiable from the available data. The authors should soften the claim to 'deliberation was not confined to race-specific concerns' or conduct a framed-control replication. Neither concern undermines the feasibility demonstration or the risk-level findings, so the verdict stays conditional rather than moving to reject or accept.","tokens_in":16729,"tokens_out":4772,"duration_ms":44974,"concrete_test":"Run a two-arm replication with matched stakeholder composition: six groups receive the current racial-equity framing and six receive a generic risk/benefit framing, reviewing the same 13 use cases in counterbalanced order, with sessions audio-recorded and transcribed. Pre-register a codebook tagging each reasoning episode as race-specific, general efficacy, benefit, or equity/distribution. Compare the prevalence of the three Section 5 questions across arms. If the generic arm shows the same three-question pattern at similar rates, the 'did not narrow' claim is unsupported; if the racial-equity arm shows broader integration, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (Section 5: 'our main finding is that although the explicit framing was racial bias, the reasoning process that emerged was considerably broader') requires a counterfactual reading: 'foregrounding racial equity did not narrow the deliberation' asserts something about what would have happened absent that framing. Every group received the same racial-equity framing, and no baseline or comparison arm was included. The same observed pattern of questions about efficacy, benefit, and equitable distribution could have arisen under any risk/benefit framing; the design cannot distinguish 'the framing broadened reasoning' from 'these are the generic questions stakeholders ask about any AI tool.' The process reconstruction also depends on authors' memory of unrecorded discussions (Sections 3.4 and 3.5), so the back-engineered questions in Section 5 cannot be independently checked against primary dialogue. This is a scope-of-claim problem, not a data-integrity one: the descriptive worksheet findings and risk classifications are transparently reported, but the comparative 'did not narrow / deepened' language in the abstract and conclusion exceeds what the design supports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a one-day mixed-stakeholder deliberation workshop in which 30 participants (community representatives, police officers, and academics) assessed 13 AI use cases in policing using a green/amber/red risk classification, with an explicit framing of racial bias. The authors find that participants rejected three use cases as unacceptable, most notably recidivism risk assessment, and that the reasoning process centered on three broad questions: whether the tool works, whether it delivers genuine benefit, and whether the benefit extends to everyone. They interpret this as evidence that foregrounding racial equity did not narrow deliberation, likening it to the curb-cut effect in inclusive design.","tokens_in":16875,"tokens_out":7820,"duration_ms":66716,"significance":"The paper addresses a genuinely understudied question—how affected communities reason about AI use cases in policing—and its descriptive results are valuable. The authors are transparent about their limitations (single-occasion workshop, convenience sample, unrecorded discussions, order effects), and Table 2 provides per-use-case classifications that could be reused by other researchers. The distinction between premise-based and implementation-based objections, particularly for recidivism risk assessment, is a useful analytic contribution. However, the headline claim that the racial-equity framing 'did not narrow' deliberation is comparative and cannot be established by the single-arm design; the paper's value lies in the descriptive pattern rather than in the counterfactual claim.","major_comments":[{"comment":"The central claim that 'foregrounding racial equity did not narrow the deliberation' (Section 5, echoed in the Abstract and Conclusion) is a counterfactual, comparative statement. All workshop groups received the same explicit racial-bias framing, and there is no baseline or comparison condition in which participants assessed the same use cases under a neutral or generic risk-benefit framing. The observed pattern of questions (does it work, will it deliver genuine benefit, will benefit extend to everyone) could plausibly arise in any stakeholder deliberation about any AI tool, so the design cannot distinguish 'the equity framing broadened reasoning' from 'these are the generic questions stakeholders ask about AI.' Please reframe this as a descriptive finding about the content of reasoning, or explicitly acknowledge in the main text that a comparative conclusion is not supported by the single-arm design.","section":"Section 5; Abstract; Conclusion"},{"comment":"The reconstruction of participants' reasoning in Section 5 rests substantially on the authors' non-recorded observations of group discussions, as described in Sections 3.4 and 3.5. Because no primary audio or video record exists, the 'back-engineered questions' cannot be independently checked against the actual dialogue, and the assertions in Section 5 (e.g., 'the reasoning process that emerged was considerably broader') present an interpretative synthesis as if it were a direct description. The limitations paragraph 3.5 acknowledges that sessions were not recorded but does not temper the strength of the process-level claims elsewhere. I recommend presenting the three-question framework as an analyst-constructed interpretation, supported by worksheet quotations, and reporting any steps taken to validate the interpretation (e.g., member checking or inter-rater agreement) or making the interpretive status explicit wherever these claims appear.","section":"Sections 3.4, 3.5, and 5"},{"comment":"Table 2 reports means and standard deviations for ordinal risk categories (1=green, 3=amber, 5=red) and assigns the intermediate value 2 when groups marked two categories. This imposes an interval scale on ordinal data and is the basis for the claim in Section 6.3 that 'only 3 out of 13 use cases ... with an average risk above medium.' Because the mapping is arbitrary (e.g., one could equally code green=2, amber=4, red=5), the numerical averages should not be used for quantitative comparisons. Report the full distribution of green/amber/red classifications (including the number of groups that declined to classify) and use medians/modes or contingency tables; at minimum, clearly label the means as a crude descriptive summary rather than an interval-scale statistic. The small and non-independent group sizes (n=3-5, with participants remixed across morning and afternoon sessions) further limit the interpretability of standard deviations.","section":"Table 2; Section 6.3"}],"minor_comments":[{"comment":"Several words are missing spaces in the typeset text (e.g., 'wherecrime', 'whereand', 'whencrimes', 'PredPolT M'); please fix these formatting errors.","section":"Section 2.1"},{"comment":"The phenomenon described in the footnote (models trained on historical recruitment data learning existing norms) is not usually called the 'credit assignment problem' in machine learning; that term refers to apportioning credit across multiple decisions. Please correct or rename the reference.","section":"Section 4.1, footnote 1"},{"comment":"The text cites 'CPIA 2' in reference to the Criminal Procedure and Investigations Act; the Act is from 1996 and no '2' appears in its short title, so this appears to be a typo.","section":"Section 4.4"},{"comment":"The citation style for Moore (2015) is inconsistent: '(Moore 2015)' appears in Box 1 while the reference list uses 'Moore, R. 2015'; please unify in-text citations.","section":"References"},{"comment":"In the sentence 'we asked the PRAP team to highlight any use cases they wanted include,' the word 'to' appears to be missing before 'include'; please correct.","section":"Section 3.2"},{"comment":"The term 'back-engineered questions' is used without a definition at its first appearance in Section 5; either define it or refer readers to the analytic procedure described in Section 3.4.","section":"Section 5"},{"comment":"The concluding sentence 'Foregrounding racial bias did not narrow or politicise the deliberation; it deepened it' repeats the comparative claim that the design cannot support; please align the conclusion with the descriptive finding (e.g., 'participants raised a broader range of questions than racial bias alone').","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely to be of interest to the AI-and-society community, and the workshop data appear to be reported in good faith. My main concern is that the headline claim is presented in the abstract and conclusion with more strength than the single-arm, unrecorded design supports. If the authors reframe the central finding as descriptive and add appropriate hedging, I would be willing to accept. No issues of citation or novelty that I detected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading. The workshop itself is well run, the reporting is transparent, and the limitations section is above average. Two things you should know before you get to it. First, the descriptive core—how 30 mixed stakeholders rated 13 policing AI use cases, with recidivism risk assessment rejected on the premise rather than the implementation—is the real substance, and it holds up. Second, the process claim in the abstract and Section 5, that foregrounding racial equity 'did not narrow the deliberation,' is not supported by the design. Every group received the same racial-equity framing; there is no baseline, comparison arm, or counterfactual that would let you infer what the deliberation would have looked like absent that framing. The three questions they distil—does it work, will it deliver genuine benefit, will that benefit extend to everyone—may well be the generic questions any mixed stakeholder group asks about any AI tool. The paper cannot distinguish 'racial framing broadened the reasoning' from 'these are just the questions people ask.'  That is a scope-of-claim problem, not a data-integrity one, but it is the load-bearing interpretive move and it needs to be softened or empirically supported.\n\nThe genuinely new thing here is the specific PRAP-aligned mixed-stakeholder deliberation on a broad UK policing use-case set, and the striking result that participants rejected recidivism risk assessment on the premise, not its implementation. That distinction is hard to surface any other way, and the authors make a careful case for early community consultation. The curb-cut framing is an interpretive analogy, not a method, but it is a memorable way to motivate the point.\n\nSoft spots beyond the main design issue: the ordinal RAG categories are averaged as numbers (1, 3, 5) and reported with standard deviations. They disclose this, so it is a presentational choice rather than a statistical sin, but it invites a false precision—better to report the raw distributions. The n's are small (3–5 groups per use case) and groups are not independent, which is fine for qualitative description but not for any claim about generalizability. The process reconstruction also leans on the authors' unrecorded observations; they acknowledge this and cross-checked among four authors, but the back-engineered questions in Section 5 cannot be independently verified. I'd want the worksheets or an anonymized coding record if that claim stays. The comparison with police-only surveys (Kearney et al.) is not matched on protocol or instruments, so the 'more accepting than police alone' line should be read as suggestive only.\n\nWho is this for? People working on participatory governance of police AI, and practitioners engaging with the Police Race Action Plan. It deserves a serious referee. A good referee should push on the 'did not narrow' claim and ask for either softer language or a framed-control replication. I would support conditional acceptance with that revision.","headline":"A genuinely useful and honest participatory workshop study whose headline 'racial equity did not narrow deliberation' claim outruns the design.","tokens_in":17444,"tokens_out":2067,"would_cite":true,"duration_ms":20713,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A mixed-stakeholder deliberation on AI in policing, framed explicitly around racial bias, produced broader—not narrower—reasoning about whether tools work, benefit anyone, and benefit everyone.","keywords":["AI in policing","racial bias","mixed-stakeholder deliberation","risk assessment","recidivism risk prediction","live facial recognition","participatory design","curb-cut effect"],"falsifier":"Run the same workshop with sessions audio-recorded and independently coded: if transcripts show that participants' justifications for risk ratings mostly cite harms to racial minority communities and rarely invoke whether a tool works, whether it delivers genuine benefit, or whether that benefit reaches everyone, the paper's central process finding would be contradicted.","tokens_in":16490,"feed_emoji":"⚖️","tokens_out":10728,"duration_ms":86148,"temperature":0.7,"pith_summary":"The paper reports a one-day workshop in which 30 community representatives, police officers, and academics used a red, amber, and green risk framework to assess 13 AI use cases in policing, with an explicit instruction to focus on racial bias. The central finding is that the racial-bias framing did not narrow the discussion: participants consistently reasoned through broader questions about whether a tool works, whether it delivers genuine benefit, and whether that benefit extends to everyone. The paper argues this integrated reasoning resembles the curb-cut effect—the way a design intended for one marginalised group, like kerb ramps for wheelchair users, ends up helping everyone—and so treating racial equity as a starting lens rather than an add-on checklist improves the whole risk-benefit analysis. If correct, the finding implies that community consultation can surface objections to a tool's premise, as with the outright rejection of recidivism risk assessment, that technical review alone would likely miss.","feed_headline":"Foregrounding racial bias did not narrow police-AI risk talks","feed_subtitle":"Mixed stakeholders asked whether a tool works, delivers genuine benefit, and extends that benefit to everyone","key_machinery":"The central mechanism is the one-day mixed-stakeholder deliberation itself: 30 participants in six groups, each group reviewing AI use cases organised by the Police Race Action Plan's four themes and classifying them as low risk (green), medium risk (amber), or unacceptable risk (red) on a worksheet that asked for justifications, permitted uses, and monitoring needs. The analytic device that carries the argument is the authors' post-hoc reconstruction of what they call back-engineered questions—the implicit questions participants appeared to be asking as they assigned risk—based on 76 completed worksheets and the authors' own observations of unrecorded group conversations. These questions (does it work, will it deliver genuine benefit, will that benefit reach everyone, and can success be measured) are what the paper claims actually structured the deliberation, and they are the evidence that the racial-equity framing did not narrow the discussion. The interpretive lens is the curb-cut effect, the known phenomenon in which a design constraint introduced for a marginalised group, such as kerb ramps for wheelchair users, ends up producing better outcomes for everyone.","core_discovery":"The paper's central claim, stated as its main finding, is that although the explicit framing of the workshop was racial bias, the reasoning process that emerged was considerably broader. Across 13 use cases, participants did not focus only on harms to racial minority communities; their justifications repeatedly organised around three back-engineered questions: does the tool actually work, will it deliver genuine benefit, and will that benefit extend to everyone? Recidivism risk assessment is the pivotal case: it drew the strongest objections of any use case, and the objection was to the premise of predicting reoffending from arrest records rather than to failures in how a particular system was implemented. The authors present this pattern as evidence that an equity-centred lens can act like a curb-cut, producing reasoning and conditions that serve everyone, and they propose mixed-stakeholder deliberation at the ideation stage as a practical way to negotiate acceptable risk boundaries before adoption.","pith_inferences":["The paper's design includes no counterfactual, so a natural next step would be to run the same deliberation with and without an explicit racial-bias prompt to test whether the broad three-question pattern is caused by the framing or by the mixed-stakeholder composition.","The premise-versus-implementation distinction could be used predictively: tools rejected on premise, like recidivism risk assessment, should be resistant to technical fixes, while conditionally accepted tools, like live facial recognition, should shift with new performance and governance evidence.","The same three-question structure may transfer to other public-sector AI domains, such as health, housing, social care, or education, where equity-focused stakeholder deliberation could surface similarly broad questions.","Because participants were recruited through a police race-action network and had existing engagement with racial bias issues, the results describe one already-engaged stakeholder group; extrapolating to the general public or to police forces as a whole would require a differently sampled study."],"forward_implications":["If the finding is correct, policing bodies should put racial-equity deliberation at the start of the AI risk-benefit process rather than treating bias as a separate checklist item.","Community consultation can separate objections to a tool's premise from objections to its implementation: recidivism risk assessment was rejected on the premise, while hotspot mapping and live facial recognition were accepted only under conditions.","Mixed-stakeholder groups were broadly open to AI, rejecting only 3 of 13 use cases outright, so involving communities does not amount to a blanket obstacle to technological adoption.","The conditions attached to acceptable use cases—human oversight, transparency, monitoring, and restrictions such as using facial recognition only against high-harm offenders—can serve as concrete pre-deployment requirements.","The curb-cut analogy implies that tools scrutinised through a racial-equity lens from the outset are more likely to work for everyone, which makes the equity lens a design resource rather than a constraint."],"supporting_citations":[{"why":"Documents how predictive-policing feedback loops compound over-policing; grounds the implementation-level objections to hotspot mapping in the workshop discussions.","marker":"Ensign et al. 2018"},{"why":"Shows how arrest-record proxies encode racial disparities in risk-assessment instruments; supplies the technical basis for the premise-level rejection of recidivism prediction.","marker":"Zilka et al. 2023b"},{"why":"Establishes higher error rates for darker-skinned individuals in commercial facial recognition; anchors the racial-bias concerns in the live facial recognition deliberation.","marker":"Buolamwini and Gebru 2018"},{"why":"Documents racial inequity in recidivism risk assessment tools; provides background for the strongest objections raised in the workshop.","marker":"Angwin et al. 2016"},{"why":"Shows how predictive policing can reproduce racial bias; supports the hotspot-mapping discussion and the conditions participants attached to it.","marker":"Lum and Isaac 2016"},{"why":"Prior participatory study finding community members questioned the core motivation of a policing AI tool rather than its technical execution; supplies the comparison for the premise-versus-implementation distinction.","marker":"Haque et al. 2024"},{"why":"Interview study of police practitioners' aversion to predictive risk assessment and facial recognition; provides the contrast showing this mixed-stakeholder group was more accepting of AI.","marker":"Kearney et al. 2024"},{"why":"Names the curb-cut effect linking human-centred design to better outcomes for everyone; supplies the paper's interpretive lens for its main finding.","marker":"Shneiderman 2020"}],"fun_headline_variants":["Racial bias lens widens police-AI risk debate","Equity focus broadens police-AI deliberation","Curb-cut effect: bias lens opens police-AI questions","Mixed stakeholders ask if police AI works for everyone","Recidivism AI rejected on premise, not implementation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the authors' reconstruction of group reasoning, based only on worksheets and their own unrecorded observations, faithfully captures what participants actually reasoned, since the workshop sessions were intentionally not recorded and cannot be independently checked.","fun_headline_variants_meta":{"raw":{"variants":["Racial bias lens widens police-AI risk debate","Equity focus broadens police-AI deliberation","Curb-cut effect: bias lens opens police-AI questions","Mixed stakeholders ask if police AI works for everyone","Recidivism AI rejected on premise, not implementation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000262,"raw_usage":{"total_tokens":1573,"prompt_tokens":899,"completion_tokens":674,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":597}},"tokens_in":515,"tokens_out":674,"duration_ms":6865,"temperature":1.0,"reasoning_tokens":597,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T13:31:47.679119+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same workshop with sessions audio-recorded and independently coded: if transcripts show that participants' justifications for risk ratings mostly cite harms to racial minority communities and rarely invoke whether a tool works, whether it delivers genuine benefit, or whether that benefit reaches everyone, the paper's central process finding would be contradicted.","supporting_citations":[],"review_version":1}