{"id":"7d3bc221-a4c5-444b-9d87-1c44c54c56d8","arxiv_id":"2608.06831","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Based on 25 interviews, AI risk taxonomies are poorly integrated into governance because their design choices are invisible to users and they rarely link harms to responsible actors.","lead":"Interviews with 25 AI governance practitioners and researchers show that risk taxonomies (SOT) are weakly integrated into real governance decisions, and that their reductive design choices are often invisible to users. The paper proposes design fixes and argues that shared governance infrastructure is needed before taxonomies can meaningfully support AI accountability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'weak integration' finding equates developers' lack of visibility with absence of use; the paper's own codebook suggests concrete uses were reported.","rationale":"The reader's weakest assumption focused on sample representativeness: 25 interviewees recruited through prior SOT engagement, geographically concentrated, with no non-users and no Global South perspectives. That is a real limitation, but the more load-bearing problem is internal to the evidence: the study's own data show that developers have little visibility into downstream use, and the paper treats that absent visibility as evidence of weak integration. This conflates absence of evidence with evidence of absence. The Appendix C codebook and several participant quotes suggest concrete uses of SOT were described, including in pre-deployment review and product-team deliberation. If those uses exist in the transcripts, the central finding as worded overstates what the interviews support. This does not invalidate the paper's design critique or its recommendations; it does mean the empirical foundation for 'weakly integrated' is thinner than the abstract and Section 4 suggest. The reader's CONDITIONAL verdict remains appropriate, and no verdict change is needed beyond the condition already stated.","tokens_in":26694,"tokens_out":5492,"duration_ms":58530,"concrete_test":"Re-code all 25 transcripts into three mutually exclusive categories: (A) first-hand report that a named SOT informed a specific deployment decision; (B) first-hand report that SOT-derived categories were used in post-deployment monitoring; (C) statement that the speaker has no visibility into downstream use. Count participants in each category and list representative quotes. If (A) or (B) contains even a few first-hand reports, the Section 4 phrasing 'little evidence' should be revised to 'unsystematic, mostly invisible use,' and the paper should explain how the two explanatory mechanisms survive those reports. If (A) and (B) are empty, the current wording is supported but only for the engaged sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 4 — 'SOT are not well integrated into AI governance processes, with little evidence of SOT being used to inform AI deployment decisions or the monitoring of AI impacts' — is supported mainly by participants' reports that they lack visibility into how their taxonomies are used. For example, I20: 'I unfortunately have no evidence whether people have found it useful'; I14: little 'insight into how [developers] actually do, or if they actually do, use them.' These quotes are statements of ignorance, not observations of non-use. The existence of visibility gaps cannot by itself establish the prevalence of SOT use in deployment or monitoring. The paper's own Appendix C codebook (T4) lists 42 low-level codes under 'Specific use cases and applications,' including pre-deployment ethics review, evaluation design, and guiding product team deliberation, indicating that some participants did describe concrete uses. The inference from 'developers cannot see use' to 'SOT are weakly integrated' is therefore the weakest load-bearing step. A representative sample of deployment decisions or direct observation would be needed to measure integration; the retrospective, SOT-engaged sample cannot establish prevalence. This concern is partially acknowledged in the Limitations section, but it applies directly to the wording of the central finding, not only to generalizability.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study of how sociotechnical outcome taxonomies (SOT) for AI risks are developed and used, based on 25 semi-structured interviews with researchers and practitioners in industry, academia, civil society, and government. The central claim is that SOT are weakly integrated into AI governance processes, with little evidence that they inform deployment decisions or post-deployment monitoring. The authors identify two design/use features that explain this: (1) the reductive design choices behind SOT are invisible to downstream users, leading categories to be treated as exhaustive descriptions of risk rather than interpretive aids, and (2) SOT enumerate harms without linking them to decision points or responsible actors, making accountability difficult to assign. The paper closes with design recommendations (interoperability, extensibility, traceability) and an ecosystem-level proposal for a registry of SOT that would support monitoring and standardisation only after stronger infrastructural substrate exists.","tokens_in":26878,"tokens_out":2985,"duration_ms":34624,"significance":"If the findings hold, the paper makes a useful empirical contribution to an area that currently lacks direct qualitative evidence. It is one of the first studies to examine SOT development and use in practice, and it offers concrete, design-oriented recommendations with practical relevance to AI governance, RAI tooling, and standardisation efforts. The authors are transparent about their inter-sectoral positions, provide a detailed codebook, and include participant quotes that ground the analysis. The two explanatory mechanisms--invisible reduction and missing causal linkage to decision points--are plausible and connect well to existing STS scholarship on classification and infrastructure. The central evidentiary claim about 'weak integration' is, however, not as strongly supported as its general wording suggests: the interview data largely show that taxonomy developers lack visibility into downstream use, not that use is absent. The paper's value would be strengthened considerably if this distinction were made explicit and if the central finding were reframed accordingly.","major_comments":[{"comment":"The central claim that 'SOT are not well integrated into AI governance processes, with little evidence of SOT being used to inform AI deployment decisions or the monitoring of AI impacts' conflates participants' lack of visibility into downstream use with an observed absence of use. The quoted evidence--I20's 'I unfortunately have no evidence whether people have found it useful' and I14's 'insight into how [developers] actually do, or if they actually do, use them'--are statements of ignorance about use, not observations of non-use. The paper's own codebook (Appendix C, T4, 'Specific use cases and applications') contains 42 low-level codes documenting concrete reported uses, including pre-deployment ethics review, evaluation design, and guiding product team deliberation. The current wording of the finding therefore overstates the support in the data. I recommend reframing the central finding as, for example, 'SOT developers lack systematic visibility into downstream use, and no tracking infrastructure exists to assess adoption or effects'--a claim that the interview data do support--or, alternatively, providing direct observational or prevalence data that would justify the original 'weak integration' claim.","section":"3.1 Recruitment, Table 1, and 6 Limitations"},{"comment":"The generalisation 'SOT are weakly integrated into AI governance processes' is stated without geographic or sampling qualification, even though the sample is heavily concentrated in North America (18/25), includes no participants from the Global South, and was recruited entirely through prior engagement with SOT, thereby excluding practitioners who encountered SOT and chose not to use them. The Limitations section acknowledges this, but the abstract, findings, and conclusion all repeat the broad claim. Since the paper's contribution is partly a general empirical statement about the state of SOT adoption, this mismatch between claim and evidence is load-bearing. Please qualify the central claim (e.g., 'among SOT-engaged actors in North America and Europe/UK') or present additional evidence that the pattern extends beyond the sampled population.","section":"4 Findings, first paragraph"}],"minor_comments":[{"comment":"There is a typo: 'We refer to these artefacts associotechnical outcome taxonomies' should read 'as sociotechnical outcome taxonomies'.","section":"1 Introduction"},{"comment":"The codebook appendix states the full codebook is 'available at https://doi.org/10.5281/zenodo.21830185' but the final line says it 'will be made available as supplementary material'; please make these statements consistent.","section":"Appendix C, final line"},{"comment":"The row labelled '2025' for AGORA cites Arnold et al. 2024, while the reference list gives 2024; please align the year in the table with the actual publication year.","section":"Appendix A, Table 2"},{"comment":"The same paper by Lee et al. is listed twice (2024a and 2024b) with identical titles; please merge into a single reference or distinguish the versions if they are genuinely different.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a competent qualitative study with a transparent method and a useful codebook. My main concern is that the headline finding, as worded, claims more than the evidence supports. The distinction between 'developers cannot see whether their taxonomy is used' and 'taxonomies are not used' is crucial, and the paper already contains data (Appendix C, T4) that point in the direction of actual uses. I believe the authors can fix this with reframing and more careful language; I do not see it as a fatal flaw. The geographic and sampling limitations are acknowledged but should be reflected in the abstract and conclusion. I have no concerns about circularity: the findings are drawn from the interviews, and the authors' prior involvement in SOT development is disclosed. The recommendation for major revision reflects the load-bearing nature of the central claim, not the overall quality of the study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing: this is a careful, honest interview study, and the authors know their limits. It's one of the first empirical looks at how AI risk taxonomies (SOT) are actually built and used, based on 25 interviews. The methods are described well: semi-structured protocols, reflexive thematic analysis, a published codebook, and a transparent limitations section. The positionality statement is refreshingly direct, and self-citations are acknowledged rather than hidden. If you read it for the recommendations — interoperability, extensibility, traceability, plus the ecosystem registry — there's real practical value for RAI tool builders.\n\nThe central claim, though, is where I'd push back. The abstract says SOT are 'weakly integrated into AI governance processes, with little evidence of SOT being used to inform AI deployment decisions or monitoring.' But the support for that is largely participants saying they have no visibility into whether their taxonomies are used — 'I unfortunately have no evidence whether people have found it useful' (I20). Ignorance is not non-use. The paper's own codebook (T4) contains 42 low-level codes under specific use cases, including pre-deployment ethics review and evaluation design, which means some concrete uses were reported. So the leap from 'developers can't see use' to 'SOT are weakly integrated' is weaker than the categorical wording suggests. A prevalence claim like that needs non-user perspectives, direct observation of deployment decisions, or a baseline. The authors acknowledge the retrospective, self-selected, geographically concentrated sample (no Global South) in limitations, but the limitation applies directly to the main finding, not just to generalizability.\n\nThat said, the softer version of the finding — that integration is uneven, unmeasured, and unstructured, and that this is a problem — is well supported. The two design/use features (invisible design choices and missing causal linkage) are plausible and illustrated with quotes. And the discussion of premature standardization entrenching incumbents is thoughtful.\n\nWho it's for: RAI practitioners, SOT developers, and anyone working on AI governance infrastructure. It's a useful empirical contribution, not a definitive measure of SOT uptake. It deserves a real peer review; I'd send it out, but I'd ask the authors to either soften the abstract's categorical claim or add evidence on actual use.","headline":"Solid qualitative study of AI risk taxonomy practices, with a central claim that overreaches from developers' lack of visibility to weak integration.","tokens_in":27444,"tokens_out":1708,"would_cite":true,"duration_ms":16746,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI risk taxonomies are weakly integrated into AI governance, with two design–use features explaining why: invisible reductive choices and missing causal links to decision points.","keywords":["AI governance","sociotechnical outcome taxonomies","harm classification","responsible AI","accountability","qualitative interviews","standardisation","risk assessment"],"falsifier":"A reader could audit a sample of published SOT (for instance those catalogued by the AI Risk Repository) and count whether most include explicit causal pathways naming decision points and responsible actors; if a majority did, the paper's claim that SOT lack causal linkage would be falsified. Alternatively, direct observation of product teams during pre-deployment review could check whether risk categories from SOT are routinely traced to named owners and post-launch monitoring plans.","tokens_in":26464,"feed_emoji":"🗂️","tokens_out":4421,"duration_ms":41421,"temperature":0.7,"pith_summary":"AI risk classification schemes — sociotechnical outcome taxonomies (SOT) — are meant to give regulators, firms, and policymakers a structured account of AI harms to act on. Based on 25 interviews with developers and users across industry, academia, civil society, and government, the paper argues that SOT are only weakly integrated into actual AI governance: there is little evidence they inform deployment decisions or post-deployment monitoring. It identifies two design–use features behind this gap: the reductive choices that shape a taxonomy are invisible to downstream users, who treat categories as exhaustive accounts of risk rather than interpretive aids; and SOT enumerate harms without linking them to decision points or responsible actors, so accountability is hard to assign. The paper's recommendations — interoperability, extensibility, and traceability in SOT design, plus ecosystem-level infrastructure such as a shared registry — follow from these findings.","feed_headline":"AI risk taxonomies are barely used in governance","feed_subtitle":"25 interviews show risk lists that look exhaustive but leave out who decides and who monitors.","key_machinery":"The central object is the sociotechnical outcome taxonomy (SOT): a structured classification of AI risks, harms, and impacts intended to serve as a boundary object and proto-standard for AI governance. The argument is carried by two identified design–use features that explain weak integration: the invisibility of the reductive choices made during taxonomy construction, which leads users to treat categories as exhaustive; and the absence of causal pathways linking categories to decision points and actors, which makes accountability unassignable. The paper also introduces the 'altitude problem' — choosing the right level of abstraction for categories — as the central design tension SOT developers face, and proposes three design desiderata (interoperability, extensibility, traceability) plus an ecosystem-level registry as the remediation.","core_discovery":"The paper's central finding is that sociotechnical outcome taxonomies are not well integrated into AI governance processes: participants reported proliferation without coordination, no visibility into adoption, and little evidence that enumerated risks are monitored after deployment. The authors argue that two features explain this. First, the design choices through which an SOT reduces the complexity of AI risks are invisible to users, so frameworks intended as interpretive aids are misused as exhaustive descriptions of risk. Second, SOT generally describe harms without analysing how they arise, leaving out the decision points and actors that could be held responsible, which lets classification substitute for substantive action. These are presented as empirically grounded findings from reflexive thematic analysis of interviews, not as a formal evaluation of SOT effectiveness.","pith_inferences":["If the invisibility hypothesis is right, a low-cost intervention is to require published SOT to carry explicit design-rationale metadata (definitions, boundary conditions, evidence base), which would also make the 'altitude' choices contestable.","The missing causal-linkage result suggests SOT-based evaluations will keep tracking measurable proxies unless regulators require documented links from categories to decisions; the EU AI Act's risk classification could be a natural testbed.","The paper's registry proposal could double as a monitoring infrastructure: attaching reporting expectations to categories at deposit time would turn classification into an enforceable governance relation.","The geographic skew of the sample implies the proposed fixes should be stress-tested against how SOT are used in the Global South, where cultural erasure and diffuse societal harms may be more salient than the corporate reputational dynamics described."],"forward_implications":["If the paper is right, SOT currently provide limited governance value; simply producing more taxonomies will not fix the integration gap.","SOT design should make the reductive choices visible through metadata, boundary conditions, and documented mappings so users treat categories as interpretive aids rather than exhaustive descriptions.","SOT should include lightweight causal pathways and explicit observability requirements for classified risks, so accountability can be assigned and monitoring expectations attached.","Premature standardisation across SOT risks freezing the definitional asymmetries and entrenched interests; a shared registry that aggregates SOT along with their documentation and monitoring expectations is a more tractable first step.","Future empirical research should directly observe SOT in use across industry, civil society, and regulatory agencies rather than relying on retrospective accounts."],"supporting_citations":[{"why":"Supplies the STS frame that classification is never neutral and that design choices become invisible, motivating the paper's 'infrastructural inversion' and the opacity findings.","marker":"Bowker and Star 2000"},{"why":"Documents that there are over 70 SOT and provides the Causal Taxonomy example that the paper uses to illustrate traceability and to ground the proliferation finding.","marker":"Slattery et al. 2026"},{"why":"Canonical SOT example whose observed-versus-anticipated-risk distinction is used to illustrate the extensibility design desideratum.","marker":"Weidinger et al. 2022"},{"why":"Foundational SOT synthesizing harms by type and level; its modular architecture is cited as an example of extensibility across levels of abstraction.","marker":"Shelby et al. 2023"},{"why":"Shows how structured artefacts such as checklists provide organizational infrastructure, the model the paper likens SOT to and extends.","marker":"Madaio et al. 2020, 2022"},{"why":"Supports the claim that audit standards and comparability are underdeveloped, so classification can substitute for substantive action.","marker":"Costanza-Chock, Raji, and Buolamwini 2022"}],"fun_headline_variants":["AI risk lists look complete but dodge accountability","25 interviews: AI risk taxonomies fail to assign blame","Taxonomies of AI harms: exhaustive but not used for oversight","Risk taxonomies are not shaping AI governance, study finds","AI risk categories: designed to inform, used to absolve"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that 25 interviewees, all recruited through prior engagement with SOT and almost all based in North America or Europe, with no Global South representation, give a reliable picture of how SOT are used across AI governance worldwide.","fun_headline_variants_meta":{"raw":{"variants":["AI risk lists look complete but dodge accountability","25 interviews: AI risk taxonomies fail to assign blame","Taxonomies of AI harms: exhaustive but not used for oversight","Risk taxonomies are not shaping AI governance, study finds","AI risk categories: designed to inform, used to absolve"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000671,"raw_usage":{"total_tokens":3028,"prompt_tokens":887,"completion_tokens":2141,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":2057}},"tokens_in":503,"tokens_out":2141,"duration_ms":14602,"temperature":1.0,"reasoning_tokens":2057,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:01:00.900108+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could audit a sample of published SOT (for instance those catalogued by the AI Risk Repository) and count whether most include explicit causal pathways naming decision points and responsible actors; if a majority did, the paper's claim that SOT lack causal linkage would be falsified. Alternatively, direct observation of product teams during pre-deployment review could check whether risk categories from SOT are routinely traced to named owners and post-launch monitoring plans.","supporting_citations":[],"review_version":1}