{"id":"d07a2b5c-e5ff-457e-9244-c5539e5cc10f","arxiv_id":"2505.18523","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 61 AI/ML practitioners finds that while most believe diverse teams and data reduce bias, actual practices like bias audits and post-development D&I checks are inconsistent and often missing.","lead":"The paper reports a survey of 61 AI and machine learning practitioners about how diversity and inclusion (D&I) is perceived and applied in their organizations. It documents a clear gap between widespread belief that D&I reduces AI bias and the inconsistent practices organizations actually follow.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Subgroup and barrier claims rely on tiny cell counts and no inferential tests; e.g., 10% sceptics ≈ 6 respondents and 13% no-policy ≈ 8, so the abstract's 'major barriers' may be sampling noise.","rationale":"The reader's weakest assumption is that the self-selected, network-recruited sample (n=61, skewed demographics) is treated as sufficient for cross-demographic comparative claims. My concern is related but distinct: even under the study's own descriptive logic, the key barrier claims are based on extremely small conditional subgroups with no inferential statistics, so the patterns may be internal noise rather than real demographic differences. This is a correctness risk that can be settled by re-analyzing raw responses with exact counts and simple significance tests. It does not change the overall verdict: the paper remains a transparent, useful descriptive study with acknowledged limitations, and conditional acceptance is still appropriate. The abstract should temper the barrier list if the re-analysis fails to support it. I therefore mark agreement as partial: the reader identified generalizability as the weak spot, while I focus on the internal statistical fragility of the subgroup inferences that feed the abstract's central claim.","tokens_in":19846,"tokens_out":7094,"duration_ms":63016,"concrete_test":"Obtain the anonymized raw survey responses from the authors; recompute every percentage in Figures 2–9 and Section 5 with exact numerators and denominators. Run two-sided Fisher's exact tests on the key comparative claims (early-career vs experienced 'Not Sure' rates, male vs female scepticism, small-company 'None' policy rate, and the §6.2 claim that women face greater career barriers despite only 26% of women vs 70% of men citing barriers). Also compute Wilson 95% confidence intervals for the headline percentages (67%, 87%, 41%). If any confidence interval overlaps the comparison group or any Fisher test yields p>0.05, the abstract's 'major barriers' and the strength of the 'disconnect' should be explicitly qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('disconnect between perceived benefits and current practices, with major barriers including ... limited awareness among early-career professionals') depends on comparative percentages computed from very small denominators and reported without confidence intervals or significance tests. For instance, the 10% who believe D&I has minimal or no impact is about 6 respondents, and the 13% reporting no governance policies is about 8 respondents. The paper then cross-tabulates these tiny groups by gender, ethnicity, role, and organization size, yielding single-digit cell counts. The claim that early-career professionals have limited awareness is inferred from the fact that 75% of the 'Not Sure' respondents are early-career, but this conditional percentage does not establish that early-career respondents are more likely to be unsure than experienced ones. With n=61, such differences could easily arise by chance. Section 7.2 acknowledges the sample's external-validity limits, but the internal fragility of the subgroup analysis is not addressed. The abstract's 'major barriers' therefore rest on statistically ungrounded subgroup patterns.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a survey-based study of 61 AI/ML practitioners recruited through convenience and network sampling, examining how diversity and inclusion (D&I) principles are perceived, implemented, and challenged across the AI lifecycle. The survey is organized around five pillars (humans, data, process, system, governance) and collects both multiple-choice and open-ended responses. The central descriptive finding is a disconnect between practitioners' widespread agreement on the value of D&I (e.g., 67% say diverse teams reduce bias) and actual organizational practices (e.g., only 41% report bias audits, and D&I consideration declines from 59% in pre-development/development to 41% in post-development). The authors also identify three major barriers: under-representation of marginalized groups, lack of organizational transparency, and limited awareness among early-career professionals. The paper concludes by discussing implications for researchers, practitioners, and policymakers and acknowledges external validity limitations in Section 7.2.","tokens_in":19994,"tokens_out":6455,"duration_ms":52047,"significance":"If the results are credited, the paper contributes useful empirical evidence on a perception-practice gap in D&I within AI organizations, an area where prior work has been largely theoretical. The manuscript's strengths include a pilot-tested instrument, a mixed-methods design with open-ended questions, and an unusually candid threats-to-validity section that acknowledges self-selection bias and demographic imbalances. However, the small, self-selected sample (n=61; 70% Asian, 67% male, 52% early-career) severely limits generalizability, and the paper's more granular comparative claims about demographic subgroups and about 'limited awareness among early-career professionals' are not statistically supported. Because these specific claims appear in the abstract as major barriers, they are load-bearing for the paper's headline message. With appropriate qualification of those claims, the descriptive core of the paper remains a useful snapshot of practitioner perspectives.","major_comments":[{"comment":"The abstract's claim that 'limited awareness among early-career professionals' is a major barrier rests on subgroup analyses with very small denominators and no uncertainty quantification. For example, the 10% of respondents who believe diverse teams have minimal or no impact is about 6 of 61 respondents (Fig. 2, §5.2), yet this group is cross-tabulated by gender, ethnicity, age, and organization size, yielding single-digit cell counts. Similarly, §5.3 says 'uncertainty about D&I initiatives is more common among early-career professionals (75%)'; this is a conditional proportion of the 'Not Sure' group and does not demonstrate that early-career respondents are more likely to be unsure than more experienced respondents. With n=61, such differences could easily arise by chance. Please either report confidence intervals or other measures of uncertainty and explicitly temper these comparative claims, or remove them from the abstract.","section":"§5.2, §5.3, and Abstract"},{"comment":"The paper equates 'Not Sure' responses with limited awareness, as in §6.1: 'Many early-career professionals also reported uncertainty about their organisation's D&I in AI efforts, pointing to possible communication and visibility gaps.' The survey instrument (Appendix A) contains no direct measure of awareness, and 'Not Sure' could equally reflect question ambiguity, lack of direct knowledge due to role specialization (as the authors themselves note for data governance teams), or genuine uncertainty. Section 7.3 discusses construct validity but does not address this specific conflation. Please justify the interpretation of 'Not Sure' as awareness, or soften the wording to avoid overreach.","section":"§6.1 and §7.3"},{"comment":"The governance statistics are internally inconsistent across sections. Section 5.3 and Fig. 7 report 51% 'somewhat established', 18% 'fully implemented', 20% 'I don't know', and 13% 'None', which implies 69% of respondents report some form of policy. Section 6.3, however, states that '55% of respondents reported some presence of structured D&I governance policies'—a figure that is not presented in Section 5 or in Fig. 7. In addition, the sentence 'Among respondents who do not know, 13% of the respondents confirm that no such policies exist' is self-contradictory, since respondents who do not know cannot confirm the absence of policies. Please reconcile these numbers and clarify the exact denominator and response options used.","section":"§6.3 and Fig. 7"},{"comment":"The claim that 'only 26% of women in our survey citing barrier to career progression, compared to 70% of men' is not traceable to any item in the survey instrument in Appendix A; none of the Q17 response options refers to a 'barrier to career progression'. The same subsection states that '70% of early-career professionals (1–5 years of experience) believe that diverse teams improve AI fairness,' but the supporting cross-tabulation is not shown in the results. Please provide the underlying data or remove these unsupported comparative statements.","section":"§6.2"}],"minor_comments":[{"comment":"The manuscript contains several typographical errors that should be corrected, for example 'citepd' (§5.3), 'demograhics' (§5.3), 'hey' (§5.3), 'lifescyle' (§5.3), and 'om D&I' in the caption of Fig. 2.","section":"Throughout"},{"comment":"The text refers to 'Figure 9 in Appendix' when discussing challenges in creating inclusive AI systems, but Figure 9 appears in the main body of the paper; the additional figures in Appendix B are numbered 10–16. Please correct the cross-reference.","section":"§5.4 and Appendix B"},{"comment":"The survey instrument is not numbered consecutively: the 'Systems' question appears as 'Q: What challenges are faced...' without a question number, so the subsequent governance question is numbered Q21 but no Q20 exists. Please renumber the questions consistently.","section":"Appendix A"},{"comment":"The abstract writes 'Machine Learning(ML)' without a space before the parenthesis; please fix the spacing for consistency with journal style.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper's core descriptive finding—a perception-practice gap in D&I—is defensible and worth publishing after revision. The main concern is that the abstract and several discussion claims overstate the strength of subgroup findings from a convenience sample of 61 respondents. The authors should either add appropriate statistical qualifications and uncertainty measures or soften the specific claims. The internal inconsistency in the governance percentages (55% vs. 69%) also needs to be fixed. No issues of research integrity are apparent; the authors are transparent about limitations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Malik et al. give us a small, honestly reported survey of 61 AI/ML practitioners. The thing to know: the central claim — a gap between practitioners' stated belief in D&I and what their organizations actually do — is supported by whole-sample descriptive numbers and is worth taking seriously. The thing to be careful about: every cross-demographic comparison, and especially the \"limited awareness among early-career professionals\" line in the abstract, is built on single-digit cell counts and no inferential statistics.\n\nWhat's new and good: they ran a genuine mixed-methods survey, include the full instrument in an appendix, and report validity threats candidly. The five-pillar framework (humans, data, process, system, governance) from their earlier work is used as an organizing lens, and it does not appear to force the results. The finding that D&I consideration drops from pre-development/development (59%) to post-development (41%), and the tension between data anonymisation (62%) and bias audits (41%), are useful, credible observations from the whole sample.\n\nThe soft spots. With n=61, a convenience, network-recruited sample that is 70% Asian and 67% male, the external generalizability is limited — they admit this in Section 7.2, and they should also say it in the abstract. More important: the subgroup analysis. The 10% skeptics are about six people; the 13% without governance policies are about eight. Stating that most of the \"Not Sure\" respondents are early-career (75%) is not the same as showing early-career professionals are less aware; the paper drifts into that stronger reading. If the early-career-awareness barrier is kept, it needs proper conditioning (e.g., comparing awareness rates by experience group with confidence intervals) or it should be dropped to a hypothesis, not a finding. The rest of the abstract's barriers — under-representation of marginalised groups and lack of organisational transparency — are directly supported by whole-sample percentages (38% report under-representation; 23% are unsure about their organisation's D&I efforts) and are fine as descriptive claims. Minor: several typos, a \"Government Structures\" heading, and no raw data or aggregate cross-tabulations, which limits reanalysis.\n\nWho should read it: people working on D&I in AI or responsible AI practice who want a recent snapshot and a reusable survey. It is not a landmark, but it is honest progress.\n\nMy call: send it to peer review. A competent referee will ask for tempering of the subgroup language, some confidence intervals or raw data, and an abstract that matches the sample's limits. Conditional acceptance, leaning revise.","headline":"Small, transparent practitioner survey; the perception-practice gap claim holds, but the subgroup barrier claims need statistical tempering.","tokens_in":20526,"tokens_out":3063,"would_cite":true,"duration_ms":27507,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tries to establish that AI practitioners' belief in diversity and inclusion is not matched by their organisations' practices.","keywords":["Diversity and Inclusion","AI ethics","AI/ML practitioners","survey","AI lifecycle","bias mitigation","governance","five pillars of AI ecosystem"],"falsifier":"A representative survey or audit of AI/ML organisations that finds most conduct regular bias audits and apply D&I principles after deployment would falsify the paper's central claim of a systematic perception-practice disconnect.","tokens_in":19626,"feed_emoji":"📊","tokens_out":5746,"duration_ms":42434,"temperature":0.7,"pith_summary":"This paper tries to establish that AI and machine-learning practitioners broadly believe diversity and inclusion (D&I) make AI fairer and more trustworthy, but their organisations implement these principles only partially and unevenly. Using a survey of 61 industry professionals, it shows that most respondents credit diverse teams and datasets with reducing bias, while practices such as bias audits, inclusive governance, and post-deployment monitoring lag far behind. The authors identify under-representation of marginalised groups, lack of organisational transparency, and limited awareness among early-career professionals as the main barriers. If the finding is right, the gap between ethical commitment and everyday AI development is a structural feature of current practice rather than a lack of goodwill.","feed_headline":"AI practitioners value diversity more than they practise it","feed_subtitle":"A 61-person survey finds D&I is recognised but inconsistently applied across the AI lifecycle, especially after deployment.","key_machinery":"The organising device is the five-pillar model of the AI ecosystem — humans, data, system, process, and governance — used to structure the survey and its analysis. Each of the three research questions (perceived impact, current practices, challenges) is asked across those pillars, so the results show where in the lifecycle D&I is strongest and where it drops away. The second mechanism is the mixed-methods analysis: descriptive statistics and cross-tabulations identify demographic patterns, while thematic coding of open-ended responses captures the reasoning behind the numbers.","core_discovery":"The paper's central claim is that there is a systematic disconnect between what AI/ML practitioners believe about D&I and what their organisations actually do. For example, 87% of respondents say diverse and inclusive data reduces biased or discriminatory outcomes, but only 41% report that their organisations perform bias audits; 67% believe diverse teams reduce bias, yet only about half report active recruitment practices. D&I principles are incorporated at the pre-development and development stages by 59% of respondents but at the post-development stage by just 41%. Only 18% say their organisations have fully implemented D&I governance policies, and 20% do not know whether such policies exist. The authors read these patterns as evidence that D&I is treated as an early-stage or compliance matter rather than embedded across the AI lifecycle, and they link the gap to the under-representation of marginalised groups, weak transparency, and uneven awareness among early-career staff.","pith_inferences":["We infer that the perceived-benefit percentages are probably upper bounds, because people who volunteer for a D&I survey are likely to care more about D&I than the average practitioner; a representative sample might show less consensus and an even wider practice gap.","We infer that the sceptical minority (10%, mostly male developers, often in small organisations) is a target population for developer-facing education, since policy-level documents seem not to reach them.","The five-pillar structure implies a testable prediction: organisations with strong early-stage D&I but weak post-development oversight will show more bias drift in deployed systems, so linking survey answers to audit outcomes could validate the disconnect.","The 2024–2025 rollback of corporate DEI programs described in the paper would be expected to widen this perception-practice gap; repeating the same survey over time could measure that effect."],"forward_implications":["Organisations that already have D&I policies should expect a drop-off after the development phase and should target monitoring, feedback, and audit processes specifically.","Because data anonymisation is the most commonly cited data-diversity method yet can hide the demographic detail needed for bias detection, fairness work needs privacy-preserving ways to retain that detail.","Early-career professionals' uncertainty about their organisations' D&I efforts points to a communication gap that more visible governance and training could close.","Regulated industries such as banking, finance, and healthcare may need tailored guidance, since compliance pressures can crowd out proactive fairness measures.","Fully implemented D&I governance is rare in the sample, so enforceable standards and transparency requirements are likely needed to make D&I routine rather than aspirational."],"supporting_citations":[{"why":"Supplies the five-pillar AI ecosystem model that structures the survey and its analysis.","marker":"(Zowghi and da Rimini, 2023)"},{"why":"Systematic review of D&I challenges and solutions whose categories the survey's results are compared against.","marker":"(Shams et al., 2023)"},{"why":"Interview study of organisational enablers and barriers used to interpret governance and culture findings.","marker":"(Rakova et al., 2021)"},{"why":"Taxonomy of data and algorithmic bias sources used to frame the bias-mitigation practice results.","marker":"(Mehrabi et al., 2019)"},{"why":"Healthcare algorithm bias example used to explain why anonymisation can conflict with fairness auditing.","marker":"(Obermeyer et al., 2019)"},{"why":"Global survey showing a policy-practice divide in responsible AI, cited as corroborating the paper's governance gap.","marker":"(Workday, 2024)"},{"why":"Vision for operationalising D&I in AI that motivates the paper's focus on practical implementation.","marker":"(Bano et al., 2024a)"}],"fun_headline_variants":["AI's diversity gap: 87% endorse, 41% implement","Most AI pros value diversity, but bias audits lag","Belief vs practice: AI teams fail to embed diversity","Diverse AI data praised, but only 41% check for bias","AI diversity: widespread support, inconsistent action"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey treats its self-selected, network-recruited sample of 61 respondents — 70% Asian, 67% male, 52% with 1–5 years' experience — as sufficient evidence for comparisons across gender, ethnicity, experience, and organisation size.","fun_headline_variants_meta":{"raw":{"variants":["AI's diversity gap: 87% endorse, 41% implement","Most AI pros value diversity, but bias audits lag","Belief vs practice: AI teams fail to embed diversity","Diverse AI data praised, but only 41% check for bias","AI diversity: widespread support, inconsistent action"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000307,"raw_usage":{"total_tokens":2010,"prompt_tokens":947,"completion_tokens":1063,"prompt_tokens_details":{"cached_tokens":896},"prompt_cache_hit_tokens":896,"prompt_cache_miss_tokens":51,"completion_tokens_details":{"reasoning_tokens":980}},"tokens_in":51,"tokens_out":1063,"duration_ms":255662,"temperature":1.0,"reasoning_tokens":980,"cache_read_input_tokens":896,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:28:49.286508+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A representative survey or audit of AI/ML organisations that finds most conduct regular bias audits and apply D&I principles after deployment would falsify the paper's central claim of a systematic perception-practice disconnect.","supporting_citations":[{"cited_title":"and da Rimini, F","cited_arxiv_id":null,"evidence_quote":"Supplies the five-pillar AI ecosystem model that structures the survey and its analysis."},{"cited_title":"A., Zowghi, D., and Bano, M","cited_arxiv_id":null,"evidence_quote":"Systematic review of D&I challenges and solutions whose categories the survey's results are compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Interview study of organisational enablers and barriers used to interpret governance and culture findings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Healthcare algorithm bias example used to explain why anonymisation can conflict with fairness auditing."},{"cited_title":"Global Study : Closing the AI Trust Gap","cited_arxiv_id":null,"evidence_quote":"Global survey showing a policy-practice divide in responsible AI, cited as corroborating the paper's governance gap."}],"review_version":1}