{"id":"0343afde-fad6-41f1-9a32-f98c7ff78fa8","arxiv_id":"2608.10431","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A synthesis of 161 empirical studies finds that industry responsible AI practices have professionalized since 2019, yet persistent gaps in training, organizational support, and tailored interventions remain.","lead":"Researchers reviewed 161 studies that interviewed or surveyed people who build AI in companies. They found that awareness of responsible AI has grown but workers still lack training, support, and tools tailored to their daily jobs.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5's temporal progress claims may reflect changing sample composition: later studies disproportionately recruit designated RAI practitioners, whose awareness and formalized practices are expected to be higher, so the pre/post contrasts are not clean evidence of industry-wide change.","rationale":"The reader's weakest assumption already identified temporal comparability and corpus representativeness; my review agrees and sharpens the mechanism. The paper is a valuable, transparently reported synthesis with an open database, and the persistent-challenges findings are supported by a large number of studies across years. However, the headline progress narrative is load-bearing: it is the main reason the abstract claims 'meaningful progress,' and it is inferred from cross-sectional comparisons of studies that differ in recruitment strategy, target roles, and instrument design. Because later studies disproportionately sample practitioners who are already designated RAI specialists or who self-select into RAI work, the observed increases in awareness, formalized fairness testing, and toolkit adoption may reflect who is studied rather than how industry practice has changed. This does not invalidate the review, but it means the strongest claim is not yet established at the confidence level claimed. A stratified re-analysis of the released database, or a revised interpretation that limits the progress claim to the studied subpopulations, would resolve the concern. The proposed test is feasible and directly targets the confound.","tokens_in":40147,"tokens_out":3417,"duration_ms":36499,"concrete_test":"Using the released database of 161 studies, code each study by recruitment population (designated RAI role-holders or self-selected RAI practitioners vs. general AI/ML/software practitioners) and by outcome type (self-reported awareness vs. observed practices). Then recompute the Section 5.1–5.3 pre/post contrasts within the general-practitioner stratum only, and also within studies matched on geographic region and company size. If the pre/post differences in awareness, formalized fairness testing, and toolkit adoption shrink substantially or reverse in this stratum, the progress narrative is not robust to sample composition; if they persist, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim—that RAI awareness, professionalization, and intervention adoption have increased—rests on comparing studies published before and after roughly 2022 (Sections 5.1.1, 5.2, 5.3). The early period contains only 34 papers (2019–2022) versus 107 from 2023 onward, and more importantly the later corpus is compositionally different. Section 5.2 itself documents that recent studies increasingly report dedicated RAI roles, and many post-2023 studies deliberately recruit RAI specialists (e.g., Madaio et al. 2024, Smith et al. 2025) or practitioners who opted into RAI work; such samples would report higher awareness and more formalized fairness and documentation processes regardless of any real sector-wide shift. The review does not stratify its temporal comparisons by participant recruitment criteria, role, organization size, or geography, so the apparent progress could be a selection artifact rather than a change in the underlying population. This is the load-bearing assumption behind the abstract's 'meaningful progress' conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a systematic literature review of 161 empirical studies (published 2019 through August 2025) that engage industry practitioners in responsible AI (RAI) work. The authors searched computer science and specialized databases with keyword queries and citation snowballing, then manually annotated each paper for methods, findings, and implications. The synthesis is organized around four areas of reported progress (increased RAI awareness, professionalization of RAI roles and practices, wider adoption of policies/processes/toolkits, and initial external stakeholder engagement) and four persistent challenges (insufficient RAI knowledge and training, organizational resource constraints and misalignment, lack of tailored interventions, and structural barriers to external engagement). The paper concludes with implications for researchers, practitioners, and policymakers, and releases an open-source database of the annotated corpus.","tokens_in":40304,"tokens_out":6997,"duration_ms":62169,"significance":"If the synthesis is correct, this is the most comprehensive and practice-grounded account of empirical RAI-in-industry research to date, consolidating a dispersed body of work across HCI, software engineering, and AI ethics venues. The paper's strengths include a transparent search and annotation protocol, an open-source database that supports verification and reuse, explicit acknowledgment of search truncation (Section 3.1) and reliance on self-reported data, and a clear separation between synthesized findings and forward-looking recommendations. The review also identifies concrete research gaps (e.g., understudied generative-AI contexts, limited geographic diversity, narrow role coverage) that are well supported by the corpus. The central claims about 'meaningful progress' versus persistent challenges are plausible and broadly consistent with the cited literature, but the temporal progress claims rest on a comparability assumption that the manuscript does not adequately defend, and several quantitative prevalence statements lack the supporting breakdowns needed to assess their robustness.","major_comments":[{"comment":"The claim that awareness, professionalization, and adoption of RAI interventions have increased over time is inferred by comparing studies published before and after roughly 2022, but the earlier and later corpora differ systematically in composition. The early period contains 34 papers (2019–2022) versus 107 from 2023 onward, and Section 5.2 itself documents that recent studies increasingly report dedicated RAI roles, while many post-2023 studies deliberately recruit RAI specialists (e.g., Madaio et al., 2024; Smith et al., 2025) or practitioners who opted into RAI work. Such samples would report higher awareness and more formalized practices regardless of any sector-wide shift. The review does not stratify its temporal comparisons by participant recruitment criteria, role, organization size, or geography, nor does it test whether the observed pattern survives such stratification. Because the abstract's headline 'meaningful progress' conclusion depends on these temporal contrasts, the authors should either provide a compositional analysis (e.g., a table showing role and recruitment distributions per period) or substantially soften the progress claims and explicitly frame them as suggestive evidence conditional on sample comparability. The current Section 6.5 limitations paragraph does not mention this threat.","section":"Sections 5.1.1, 5.2.1, 5.3, 5.4; Abstract"},{"comment":"The review repeatedly anchors its 'persistent challenges' narrative in prevalence counts (e.g., 'more than 120 papers' for lack of expertise, 'more than 120 papers' for lack of prioritization, 'more than 100 papers' for untailored interventions, 'nearly all papers' for organizational support), but it never reports the exact numerators or a table of thematic frequencies broken down by year, role, or method. Since the open-source database could support such counts, the manuscript should include a supplementary table with these numbers and the criteria used to define each theme. Without this, the reader cannot assess the robustness of the prevalence claims or the proportionality of the synthesis relative to the underlying evidence.","section":"Sections 6.1, 6.2, 6.3"}],"minor_comments":[{"comment":"The text 'around 80% fo the paper' contains a typo; it should read 'of the papers.'","section":"Section 4"},{"comment":"The interview participant statistics (mean = 27, SD = 25.79, range = 12–55) are mathematically inconsistent: for a bounded variable with range 12–55, the maximum possible standard deviation is (55−12)/2 = 21.5. Please verify all descriptive statistics and correct any data entry or reporting errors.","section":"Section 4"},{"comment":"In the list of four motivators, the third item reads 'attending to the demand on RAI from end users []' with an empty citation bracket; a supporting reference should be supplied or the bracket removed.","section":"Section 6.2.1"},{"comment":"The citation placeholder '[87?]' should be resolved to the intended reference (e.g., Hurst et al., GPT-4o system card) or removed.","section":"Section 7.2.2"},{"comment":"The sentence 'Figure 2 provides an overview of the identified challenges and opportunities, mapped to the practices we observed' refers to Figure 2, but Figure 2 is the paper review process diagram in Section 3.1; the overview described here appears to correspond to Figure 1.","section":"Section 6 (introductory paragraph)"},{"comment":"The first-300-results truncation is acknowledged, but the authors do not estimate how many potentially relevant papers might have been missed by this procedure; a brief sensitivity note would help readers calibrate the corpus's completeness.","section":"Section 3.1"},{"comment":"The seed set was assembled from the authors' own knowledge, and all authors have previously published papers meeting the inclusion criteria. The manuscript does not report how many of the 45 seed papers or of the final 161 are authored or co-authored by the review team. Given the visible presence of the authors' own work in the corpus (e.g., refs [55,57,83,119,120,121,174]), a disclosure of the self-citation count would increase transparency.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"This is a strong and useful literature review, and the open-source database is a valuable contribution. The main risk to the paper's central narrative is the unaddressed compositional shift in the temporal comparisons; I would urge the editor to require either a stratified analysis or explicit and prominent caveats before publication. The authors' own prior work constitutes a noticeable fraction of the corpus; this is not inherently disqualifying, but a disclosure of self-citation counts would be appropriate and would also serve the paper's own methodological transparency goals."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a systematic review of 161 empirical studies on responsible AI in industry, and the main thing to know is that it's a useful, well-organized consolidation with an open database, but the temporal 'progress' claims in Section 5 rest on a sampling comparison that doesn't hold up as cleanly as the abstract suggests.\n\nWhat's genuinely useful: the authors read all 161 papers, annotated them, and are shipping the database. The synthesis organizes the literature into awareness, professionalization, tool adoption, external engagement, plus persistent barriers. For a researcher or policymaker wanting a map of what we actually know from studies of practitioners, this is the most comprehensive single source I know. The method is transparent: they state the seed set, the keyword searches, the first-300-results truncation, and the snowball process. That alone is worth something.\n\nWhere it gets softer: the abstract's 'meaningful progress' is mainly supported by comparing studies before and after roughly 2022. But the earlier period is only 34 papers, the later period 107, and the later papers increasingly recruit people in designated RAI roles—so you'd expect those samples to report more awareness and more formalized processes even if nothing changed in the underlying population. The review doesn't stratify by role, recruitment method, organization size, or region. So the progress claim is suggestive, not demonstrated. The rest of the synthesis—barriers like training gaps, organizational support, untailored interventions—is less dependent on that temporal contrast and holds up as a description of the corpus.\n\nThere's also a mild circularity issue: the seed set comes from the authors' own prior work, which then gets cited heavily. That's not fatal for a review, but it's worth asking how much the corpus overrepresents MSR/Carnegie Mellon adjacent findings. And some prevalence numbers ('more than 120 papers') are deliberately vague; they don't give exact counts.\n\nBottom line: this is a solid, careful review with a reproducible artifact. The central descriptive claims about persistent barriers are well supported. The temporal progress narrative needs a caveat, but doesn't sink the paper. Who benefits: anyone doing empirical RAI research, and policy people who want a grounded overview. I'd send it to a serious referee.","headline":"A solid, comprehensive review of 161 empirical RAI studies with an open database, but the abstract's 'meaningful progress' conclusion rests on a temporal comparison that is weaker than it looks.","tokens_in":40853,"tokens_out":1974,"would_cite":true,"duration_ms":17432,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A synthesis of 161 empirical studies finds that responsible AI practice in industry has matured in awareness, professionalization, and tool adoption, while gaps in training, support, and tailored interventions persist.","keywords":["responsible AI","industry practice","empirical literature review","AI fairness","AI governance","practitioner challenges","RAI interventions","thematic synthesis"],"falsifier":"A matched-sample replication: take the same survey or interview protocol used in a 2019-2021 study (for example, the fairness-related practitioner survey), recruit a comparable sample of practitioners from the same roles and regions in 2024-2025, and check whether awareness and formalized RAI practices are actually higher. If awareness and practice levels match the earlier results, the paper's temporal progress claim would be refuted. A second check would compare the pre-2022 and post-2022 subsets of the corpus after excluding studies from EU-regulated contexts; if the progress signal disappears, regulation-driven geography, not general maturation, explains the trend.","tokens_in":39941,"feed_emoji":"⚖️","tokens_out":4701,"duration_ms":41316,"temperature":0.7,"pith_summary":"This paper synthesizes six years (2019-2025) of empirical research on how industry practitioners actually do responsible AI (RAI) work, drawing on 161 studies that used interviews, surveys, workshops, and other methods. It claims that the literature as a whole shows two simultaneous trends: practitioners' awareness of RAI has grown, RAI roles are becoming professionalized, and toolkits, guidelines, and policies are being adopted; yet the same practitioners still face limited training, uneven organizational support, and interventions poorly matched to their workflows. The authors argue that the field now has enough accumulated evidence to move from cataloging problems toward designing and field-testing RAI interventions as end-to-end sociotechnical systems embedded in real workflows.","feed_headline":"Responsible AI is maturing in industry, but support lags","feed_subtitle":"Synthesis of 161 studies finds progress in awareness and tools, yet gaps in training and support persist.","key_machinery":"The load-bearing mechanism is a systematically assembled and manually annotated corpus of 161 empirical papers, combined with a temporal comparison. The authors marked each paper for descriptive features (methods, participants) and findings, then clustered findings thematically into four progress themes (awareness, professionalization, intervention adoption, external engagement) and four challenge themes (knowledge gaps, organizational dynamics, tailoring of interventions, engagement barriers). The temporal comparison between studies published before and after roughly 2022 is what carries the 'progress' claims; the prevalence counts across the corpus (e.g., more than 120 papers reporting knowledge gaps) carry the 'persistent challenges' claims.","core_discovery":"The central claim is that, across 161 empirical studies, the evidence supports a reading of industry RAI as both progressing and persistently challenged. Earlier studies (roughly 2019-2022) describe ad hoc, advocacy-driven fairness work, practitioners unaware of RAI concepts, and tools rarely used; later studies (2023 onward) report practitioners who can articulate RAI's importance, dedicated RAI roles with formal responsibilities, routine fairness testing and documentation, and active adoption and customization of toolkits and guidelines. The same corpus, however, consistently documents that more than 120 papers mention inadequate RAI knowledge and training, that organizational resource constraints remain the most salient barrier, and that RAI interventions are often too abstract or output-focused for the domains, applications, and pipeline stages where practitioners work. The paper's contribution is a consolidated, practice-grounded account of this literature, organized into progress areas and persistent challenges, with implications for researchers, practitioners, and policymakers.","pith_inferences":["If the temporal progress is real, a testable prediction follows: matched-sample replications of early studies (for example, the 2019 fairness survey) should show higher baseline awareness and more formalized practices; as of this review, only one such direct comparison exists.","The review's geographic skew implies that its 'progress' story may be specific to Western, often US/UK/European, technology hubs; the same literature may not yet support claims about RAI maturation in Asian or Global South contexts.","The emphasis on generative AI as under-studied suggests an extension: empirical studies of frontier labs and LLM supply chains may reveal that awareness and professionalization claims do not carry over to foundation-model development, where accountability is fragmented across organizations.","A practical extension would be to treat the open-source database as a living resource, re-running the temporal analysis at later cutoff dates to see whether the post-2022 trends continue or plateau."],"forward_implications":["If the synthesis is correct, future RAI research should shift from documenting known problems to co-designing and evaluating interventions with practitioners in real workflows.","Organizations should treat RAI capability as infrastructure: dedicated roles, ongoing cross-role training, and integration into existing pipelines rather than parallel compliance work.","Regulators should prioritize implementability and substantive accountability over procedural check-box compliance, and use anticipated regulation as a signal that shapes organizational capacity-building.","Practitioners' external engagement with users, domain experts, and annotators needs formal infrastructures, not ad hoc outreach, to be meaningful and sustainable.","The persistence of the same challenges across roles and years suggests that isolated tool design will not solve RAI; attention must move to organizational incentives and supply-chain accountability."],"supporting_citations":[{"why":"Baseline survey of industry practitioners' fairness needs, used to anchor the early-period claims of low awareness and ad hoc practices.","marker":"[83]"},{"why":"Recent survey evidence of broad practitioner agreement on RAI importance and formalized trustworthy AI frameworks, supporting the progress claim.","marker":"[7]"},{"why":"Large-scale survey of executives showing adoption of RAI policies but persistent operationalization struggles, used for both progress and challenge claims.","marker":"[40]"},{"why":"Direct comparison with an earlier fairness survey, interpreted as evidence that practitioners increasingly prioritize fairness.","marker":"[137]"},{"why":"Documents learning about RAI on the job and the growing professionalization of RAI roles.","marker":"[119]"},{"why":"Shows practitioners actively customizing fairness checklists, supporting the claim of increased intervention adoption and adaptation.","marker":"[120]"},{"why":"Describes dedicated fairness-testing pipelines in industry, evidence of professionalized fairness work.","marker":"[174]"},{"why":"Documents increasing practitioner awareness of annotator diversity, supporting the external-stakeholder progress claim.","marker":"[94]"},{"why":"Introduces the accountability horizon concept in algorithmic supply chains, central to the supply-chain fragmentation challenge.","marker":"[46]"},{"why":"Provides practitioner perspectives on organizational enablers and persistent barriers, supporting the organizational dynamics challenge.","marker":"[151]"}],"fun_headline_variants":["The responsible AI gap: industry knows more, supports less","161 studies: Responsible AI advances, but barriers remain","From ad hoc to formal: Responsible AI's six-year evolution","Toolkits and roles spread, but RAI training stays thin","Industry RAI: awareness up, organizational support down"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's 'progress' conclusions depend on comparing studies published before and after about 2022 as if they measure the same population; if the later studies recruited different roles, company sizes, or regions, the apparent improvement could be an artifact of sampling rather than a real change.","fun_headline_variants_meta":{"raw":{"variants":["The responsible AI gap: industry knows more, supports less","161 studies: Responsible AI advances, but barriers remain","From ad hoc to formal: Responsible AI's six-year evolution","Toolkits and roles spread, but RAI training stays thin","Industry RAI: awareness up, organizational support down"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000326,"raw_usage":{"total_tokens":1829,"prompt_tokens":953,"completion_tokens":876,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":794}},"tokens_in":569,"tokens_out":876,"duration_ms":8353,"temperature":1.0,"reasoning_tokens":794,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:19:49.666799+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched-sample replication: take the same survey or interview protocol used in a 2019-2021 study (for example, the fairness-related practitioner survey), recruit a comparable sample of practitioners from the same roles and regions in 2024-2025, and check whether awareness and formalized RAI practices are actually higher. If awareness and practice levels match the earlier results, the paper's temporal progress claim would be refuted. A second check would compare the pre-2022 and post-2022 subsets of the corpus after excluding studies from EU-regulated contexts; if the progress signal disappears, regulation-driven geography, not general maturation, explains the trend.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Direct comparison with an earlier fairness survey, interpreted as evidence that practitioners increasingly prioritize fairness."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides practitioner perspectives on organizational enablers and persistent barriers, supporting the organizational dynamics challenge."}],"review_version":1}