{"id":"46590d52-a3f3-4932-aa1e-38d983561f7f","arxiv_id":"2506.09873","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Industry stakeholder involvement in AI development is driven by customer value and compliance, and currently contributes little to the responsible AI benefits that guidance documents promise.","lead":"The authors compared 56 responsible AI guidance documents with surveys and interviews of 140 AI practitioners, and found that companies mainly involve stakeholders to serve commercial goals, not to advance responsible AI aims like rebalancing power or public oversight. The study maps the gap between guidance and practice, which can help regulators and guidance writers design interventions that actually shift industry behavior.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'largely not able to contribute' claim rests on unvalidated Table 2 contribution ratings that lack a coding protocol or reliability check; the qualitative evidence supports direction but not the stated magnitude.","rationale":"The paper is transparent and well-conducted qualitative research. The guidance-document analysis is systematic; the survey and interview findings are internally consistent (commercial drivers lead to revenue-critical stakeholders and late-stage consultative methods); and prior work by Groves et al. and Corbett et al. points in the same direction. The authors also disclose the sample skew in §5. That is why I do not think the paper should be rejected. The load-bearing weakness is the jump from descriptive practice patterns to the quantified-sounding conclusion that current practices are 'largely not able to contribute.' That jump is made in Table 2, where the contribution estimates are presented as findings but are in fact author judgments. Because the ratings are not derived from a transparent, reliability-checked coding procedure, a skeptical reader cannot distinguish a genuine measurement from a restatement of the authors' interpretive stance. The proposed re-coding test would settle this: if independent coders reproduce the ratings, the claim stands; if not, the appropriate fix is to soften the claim and to present Table 2 as an interpretive mapping rather than an estimate. This supports the reader's conditional verdict rather than changing it.","tokens_in":28673,"tokens_out":4536,"duration_ms":56571,"concrete_test":"Provide the anonymized participant-level survey and interview data (or a random 30% subsample). Have two independent coders, blind to the paper's conclusions and to the five benefit labels, rate each participant's practices on a five-point scale from 'actively undermines' to 'fully realizes' each benefit, using a pre-specified rubric derived only from the guidance-document definitions in Table 1. Compute Cohen's kappa and compare the aggregate ratings with Table 2. If kappa is below 0.6 or any benefit's aggregate rating shifts by more than one category, the paper should soften 'largely not able to contribute' to a directional qualitative finding and relabel Table 2 as an interpretive synthesis instead of a measured estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (§4.3) is that current SHI practices 'are largely beyond what current SHI practices can achieve' relative to the five rAI benefits. The survey and interview evidence directly establishes that practitioners' reported drivers, stakeholder selection, and methods are commercially oriented (Figs. 1–3), but it does not directly measure whether those practices produce any of the five rAI benefits. The bridge is Table 2, whose italicized 'Estimated Contribution of Current Practices' entries (Very Low, Medium with Limited Scope, Low with Limited Scope) are the authors' interpretive judgments. No coding protocol, rubric, or inter-rater reliability is reported for these ratings, and the survey items ask about motivations, involvement, and methods, not about realized outcomes. The sampling is also convenience-based and self-selected toward SHI-attuned practitioners (99/130 via Prolific; recruitment text mentioned SHI; §3.2.2, §5), so the size of the disconnect is not measurable. The direction is likely robust—if anything, a SHI-attuned sample would overstate rAI-aligned practice—but the strength 'largely not able to contribute' exceeds what the instrument can certify.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether stakeholder involvement (SHI) as currently practiced in commercial AI development can deliver the benefits that responsible AI (rAI) guidance attributes to SHI. The authors thematically analyze 56 guidance documents from 29 organizations, identify five intended benefits, and then compare them with data from an online survey (n=130) and semi-structured interviews (n=10) with AI practitioners. They find that SHI in practice is driven mainly by commercial and compliance concerns, concentrates on revenue-critical stakeholders, uses late-stage low-agency methods, and is discouraged from expanding by internal and commercial agendas. On this basis they argue that current SHI practices are 'largely not able to contribute' to rAI efforts, and they propose guidance, terminology, regulatory, and research interventions.","tokens_in":29035,"tokens_out":5841,"duration_ms":67893,"significance":"If the directional finding is accepted, this is a valuable contribution: it challenges the assumption that familiar SHI practices automatically advance rAI, and it provides a concrete corpus of guidance benefits plus practitioner-reported patterns that can inform regulation and future participatory AI research. The authors transparently report demographic characteristics, statistical tests, and limitations, and the mixed-method design is appropriate for the research questions. The main weakness is that the headline claim about the extent of the disconnect goes beyond what the measurement instruments can certify; the evidence supports direction more strongly than magnitude.","major_comments":[{"comment":"The central claim that current SHI practices 'are largely not able to contribute' hinges on the 'Estimated Contribution of Current Practices' ratings in Table 2, but these ratings are presented without a coding protocol, rubric anchors, or inter-rater reliability check. The survey items measure drivers, stakeholder groups, methods, and barriers; they do not directly measure whether any of the five rAI benefits are realized. For example, 'Improved Risk Anticipation' is rated 'Medium With Limited Scope' based on inferences about whose harms are considered, not on any reported outcome. I recommend either softening the central claim to a directional statement ('evidence suggests that current practices are unlikely to realize the benefits...') or adding a transparent scoring procedure with independent raters to support the ordinal ratings.","section":"§4.3, Table 2"},{"comment":"The sample is a convenience sample recruited through the researchers' networks and Prolific, with 99 of 130 participants from Prolific and recruitment materials that mentioned stakeholder involvement. The paper acknowledges a possible SHI-attuned skew and argues that such a sample is useful for exploring bottlenecks. That argument is fair for the direction of the finding—if anything, a SHI-attuned sample would be expected to overstate rAI-aligned practice—but it cannot support a precise population-level magnitude. Phrases such as 'largely not able to contribute' therefore overstate what the sampling design can establish. I suggest framing the result as 'even among practitioners interested in SHI, current practices are misaligned with rAI guidance.'","section":"§3.2.2, §5"},{"comment":"The inference from 'practices are commercially driven and narrow in scope' to 'practices are not able to contribute to the five benefits' is not fully operationalized. Some commercially motivated activities could partially deliver benefits for a limited group of stakeholders (e.g., usability testing with representative users contributing to understanding of socio-technical context, or compliance-driven expert review contributing to risk anticipation). The survey and interviews establish that such benefits are not the goal and are not pursued for affected non-users or the public, but the data do not establish that the benefits are entirely absent even for the stakeholders who are involved. The mapping in Table 2 is a reasonable analytical synthesis, but it should be labeled as such, and the conclusion should be expressed in terms of 'not targeted' or 'narrowly realized' rather than 'not able to contribute.'","section":"§4.2, §4.3"}],"minor_comments":[{"comment":"The arrow '→' before 'Estimated Contribution of Current Practices' and the italic formatting may be lost in some renderings; a legend or verbal label would make the status of these ratings clearer.","section":"Table 2"},{"comment":"The headings read 'T-Tests'; these should be lowercase 't-tests,' and the phrase 'impact' should be 'impacted' in the column headings.","section":"Appendix C, Tables 6 and 7"},{"comment":"The text refers to blue, orange, and green categories; please ensure the color scheme is distinguishable for color-blind readers or add pattern labels.","section":"Figure 1"},{"comment":"The sentence beginning 'Thus, using law to tie rAI-advancing SHI more directly to commercial interests seems a powerful lever' reads as a conclusion from the present data, but the connection between regulation and changed SHI practice is plausible rather than empirically demonstrated here; consider marking it as a hypothesis for future work.","section":"§6.3"}],"recommendation":"major_revision","confidential_remarks":"The paper fits FAccT and makes a useful empirical contribution. The main issue is calibration of the headline claim; if the authors revise to separate direction from magnitude and make the Table 2 ratings auditable, I would be comfortable with acceptance. No concerns about citation or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. This is the first systematic comparison I know of between what rAI guidance promises from stakeholder involvement and what practitioners actually do. The authors analyzed 56 guidance documents from 29 organizations and derived five benefits of SHI (rebalancing power, understanding socio-technical context, anticipating risks, building public trust, enabling scrutiny). Then they surveyed 130 practitioners and interviewed 10, asking about concrete experiences rather than attitudes. That design is a real step up from Groves et al., which mostly captured theoretical knowledge from a select sample.\n\nThe core findings are well-evidenced: commercial drivers dominate (customer value, compliance), affected non-users and the public are rarely involved even when the system is expected to impact them, and methods are late-stage and consultative rather than formative. The t-tests for involvement-by-expected-impact are a nice touch, and the authors also document internal tensions—practitioners caught between management agendas and stakeholder insight. The limitations section is refreshingly honest about the SHI-attuned recruitment and the UK/male skew; they even argue the bias would likely understate the disconnect, which is credible.\n\nWhere I agree with the stress-test: Table 2's \"Estimated Contribution\" ratings (Very Low, Medium with Limited Scope) are presented as summary judgments without a coding protocol or reliability check. The raw survey/interview data support the direction—practices are not aimed at the five benefits—but they do not directly measure realized outcomes. The phrase \"current SHI practices are largely not able to contribute\" is stronger than the evidence certifies. The paper would be on firmer ground saying \"currently contribute little toward\" or \"rarely advance\" the guidance's benefits. That is a revision-level issue, not a fatal one: the qualitative data consistently point the same way, and the estimated ratings align with what practitioners directly reported.\n\nMinor quibbles: the n=10 interviewee table is useful but small, and the demographic homogeneity is acknowledged. The self-citations to the authors' earlier framework papers are appropriate background, not padding.\n\nWho this is for: people working on AI governance, participatory design, or responsible AI in industry. It gives regulators a concrete evidence base for why incentives need to change. Deserves a serious referee and likely publication after modest softening of the central claim. I would engage with it and would bring it to a reading group.","headline":"A transparent, well-scoped empirical study that will be useful to the rAI governance community; the headline claim is plausible but slightly overreaches what the instrument can measure.","tokens_in":29368,"tokens_out":1311,"would_cite":true,"duration_ms":19596,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The stakeholder engagement that AI companies already do is largely disconnected from the participation that responsible-AI guidance asks for, so it is not advancing responsible AI.","keywords":["stakeholder involvement","responsible AI","participatory design","AI development","AI policy","multi-stakeholder governance","public participation","co-design"],"falsifier":"A representative audit of commercial AI projects that measures whether affected non-users and the public are involved before system objectives are fixed, and whether their input changes those objectives, would settle the claim; if a substantial fraction of projects show such early, decision-relevant participation, the claimed disconnect is overstated.","tokens_in":28486,"feed_emoji":"🤝","tokens_out":11086,"duration_ms":117028,"temperature":0.7,"pith_summary":"Responsible-AI guidance increasingly tells developers to involve stakeholders, but this study suggests that the participation already happening in industry is mostly the wrong kind. The authors reviewed 56 guidance documents from 29 organisations and identified five expected benefits of stakeholder involvement: rebalancing decision power, understanding the social context of a system, anticipating harms, building public trust, and enabling outside scrutiny. They then surveyed 130 AI practitioners and interviewed 10, finding that real-world involvement is driven by customer value, usability, and compliance, centres on end-users, domain experts, and legal teams, and usually happens late with little decision authority. Comparing the two, the paper concludes that current practices can barely deliver the benefits guidance promises, and that commercial incentives actively discourage more responsible participation. The result matters because regulation is starting to require participation, and guidance that names it generically can be satisfied without meaningful change.","feed_headline":"AI firms consult users, not the people their systems affect","feed_subtitle":"Companies consult for customer value and compliance; guidance wants power shifted to affected communities.","key_machinery":"The analytical engine is a two-sided comparison: a thematic analysis of 56 guidance documents from 29 organisations that yields five benefit themes for stakeholder involvement, and a mixed-method picture of industry practice built from an online survey of 130 AI practitioners and 10 semi-structured interviews. The bridge between the sides is Table 2, which maps each benefit theme—rebalancing decision power, detailed understanding of the socio-technical context, improved risk anticipation, increased public understanding and trust, and enabling public scrutiny and monitoring—to observed practitioner behaviour and assigns an estimated contribution level. Stakeholder involvement (SHI) is defined broadly as engaging people with an interest in, or affected by, an AI system at points in its lifecycle; the paper contrasts this with the 'traditional' SHI of agile and user-centred design, which focuses on customers and usability.","core_discovery":"The central claim, stated in §4.3, is that current stakeholder-involvement practices in commercial AI development are not able to contribute to responsible-AI efforts: the benefits that rAI guidance associates with participation are largely beyond what current practices can achieve. Guidance treats participation as a way to shift agency toward affected communities and to let them shape system objectives; practice treats it as a way to find out what customers want, to make an interface usable, and to check legal boxes. As a result, the people most likely to be harmed—affected non-users, marginalised groups, the general public—are rarely involved even when developers know they will be impacted, and when they are involved it is late and consultative rather than early and decision-making. The paper maps each of the five guidance benefits against its practitioner findings and estimates that current practice contributes 'very low' to two of them, 'medium with limited scope' to two, and 'low with limited scope' to one.","pith_inferences":["This inference goes beyond the paper: the five benefit themes could double as an audit rubric, letting an organisation score its participation practice against the guidance and track whether changes move the score.","This inference goes beyond the paper: because the sample skewed toward practitioners already interested in participation, the gap in the broader industry is likely at least as large as the paper measures, not smaller; the paper's own limitations point in the same direction.","This inference goes beyond the paper: if the paper is right that legal requirements are the lever, a testable prediction follows—jurisdictions that impose participation duties should show a measurable rise in early involvement of affected non-users, while purely voluntary guidance should not."],"forward_implications":["Harms to people outside the customer base will keep going unanticipated, because the groups least involved are exactly the groups most likely to be harmed, even when developers know those groups are affected.","Guidance that mentions participation generically is insufficient; it must specify early involvement, who is included, and who holds decision power, otherwise usability testing can satisfy it.","Legal and regulatory pressure is the most promising lever, because practitioners already prioritise compliance-driven involvement; tying participation to legal requirements could shift the practice.","A clearer vocabulary that separates participatory development from public participation, expert consultation, and public oversight would help prevent customer testing from being labelled as responsible participation.","Tools and concrete methods are underused, so guidance with actionable techniques could raise the quality of participation."],"supporting_citations":[{"why":"Supplies the ladder-of-participation idea the paper uses to judge whether involvement grants real decision authority.","marker":"[8]"},{"why":"Documents that academic AI projects often involve stakeholders without empowering them, the comparison point for practice's low decision-authority.","marker":"[22]"},{"why":"Defines agile's customer-collaboration principle, the traditional, commercially focused form of involvement that practice reflects.","marker":"[24]"},{"why":"Shows that most AI-development studies involve stakeholders without influence over core objectives, supporting the claim that current practice falls short.","marker":"[25]"},{"why":"Represents the user-centred design standard behind the usability-focused involvement that dominates industry practice.","marker":"[34]"},{"why":"Provides the distinction between engagement that advances ethical goals and engagement that does not, which underlies the paper's interpretation.","marker":"[40]"},{"why":"Earlier study of public participation in commercial AI labs, used as the baseline for how rare and unsupported public involvement is in industry.","marker":"[43]"},{"why":"Supplies the system-lifecycle definition of stakeholder that frames who counts as a stakeholder in the study.","marker":"[56]"},{"why":"The AI regulation that the paper argues creates compliance pressure that could be redirected toward responsible participation.","marker":"[104]"}],"fun_headline_variants":["AI developer participation practices miss responsible-AI aims","Stakeholder input in AI: customer value over community power","Commercial AI participation doesn't reach responsible-AI goals","Guidance promises power shift, practice delivers compliance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a self-selected sample of 130 survey respondents and 10 interviewees, recruited through the authors' networks and an online participant platform with descriptions mentioning stakeholder involvement, stands in for commercial AI development generally, and that the authors' Table 2 ratings validly measure how much current practice supports each benefit.","fun_headline_variants_meta":{"raw":{"variants":["AI developer participation practices miss responsible-AI aims","Stakeholder input in AI: customer value over community power","Commercial AI participation doesn't reach responsible-AI goals","Guidance promises power shift, practice delivers compliance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000432,"raw_usage":{"total_tokens":2209,"prompt_tokens":954,"completion_tokens":1255,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":1192}},"tokens_in":570,"tokens_out":1255,"duration_ms":9382,"temperature":1.0,"reasoning_tokens":1192,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:37:20.005000+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A representative audit of commercial AI projects that measures whether affected non-users and the public are involved before system objectives are fixed, and whether their input changes those objectives, would settle the claim; if a substantial fraction of projects show such early, decision-relevant participation, the claimed disconnect is overstated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The AI regulation that the paper argues creates compliance pressure that could be redirected toward responsible participation."}],"review_version":1}