{"id":"85b418dc-caf9-415a-ac52-91af8c97dde6","arxiv_id":"2506.18538","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A 253-question bank for assessing AI inclusivity, organized into five pillars, built from guidelines, literature, and LLM assistance, but validated only through AI-generated personas.","lead":"This paper introduces a 253-question bank for evaluating how well AI systems address diversity and inclusion across five pillars: humans, data, process, system, and governance. It offers this as a standardized tool for AI teams and policymakers, but the validation rests entirely on AI-simulated personas rather than real users.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The simulated user study cannot carry the validation claim: Section IV uses GPT-4o to generate both the 70 personas and their answers, while Section VII.A concedes the QB has not been tested in real-world projects; the abstract's 'rigorous analysis' thus rests on an unvalidated same-model…","rationale":"Read in good faith, the paper constructs a structured artifact from multiple credible sources: the 46 D&I guidelines, the SLR-derived challenges, and the RAI question bank, with human-in-the-loop review at each iteration. That construction process is a genuine contribution, and a checklist organized around humans/data/process/system/governance could be practically useful. The problem is not the idea but the epistemic weight placed on the simulated user study. The authors do not claim the simulation is exploratory; the introduction and abstract present it as part of a 'rigorous analysis' that validates the question bank's relevance and effectiveness. The load-bearing premise is that GPT-4o personas are faithful proxies for human practitioners across roles and domains. Nothing in the paper establishes that premise; in fact, the paper's own limitations section acknowledges no real-world testing and that simulated feedback may not capture deployment complexity. The fact that the same model generated questions, personas, and answers compounds the issue, because the evaluation can at most measure the model's internal consistency, not the usefulness of the QB to humans. The cited prior work on simulated user studies is not a defense: Breuer et al. validate simulations against real user queries, which the authors do not do. Thus the central validation claim is unsupported, and the reader's REJECT verdict is appropriate. If the authors reframe the simulated study as a design-time pilot for wording and clarity, and release the full QB for independent real-user testing, the contribution could be salvageable; the current framing overclaims.","tokens_in":18742,"tokens_out":4323,"duration_ms":50656,"concrete_test":"Recruit at least 20 real AI practitioners covering the 14 roles in Section IV Step 1, administer a stratified sample of roughly 40 questions from the QB (eight per pillar) plus the five research questions from Section IV Step 3, and compare their answers with the GPT-4o persona responses summarized in Table I and Section VI.B. If real practitioners' role-relevance frequencies or usefulness themes differ substantially (for example, roles that GPT-4o reported as having zero system/governance relevance instead identify relevant questions, or more than a small fraction of questions are judged unclear), the simulated user study cannot support the validation claim. A cheaper complementary check is to regenerate personas and answers with a different LLM; if conclusions change materially, the original findings are model artifacts rather than properties of the question bank.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, that the 253-question bank was validated and provides an actionable tool, depends on the assumption that the 70 GPT-4o-generated personas in Section IV Steps 2 and 4 give responses that faithfully represent real AI practitioners. That assumption is not supported. GPT-4o also generated many of the questions (Section III.B V2, V4, V6), so the simulated evaluation is essentially a same-model self-assessment: the model answers the authors' research questions about questions it helped create. The authors' manual review only checks that responses are coherent and aligned with study objectives; it cannot establish fidelity to human judgment. The paper itself states in Section VII.A: 'it has not yet been tested in real-world AI development projects. Simulated feedback may not fully capture the complexities of actual AI deployment environments.' That is exactly the load-bearing condition. Consequently, the abstract's phrase 'rigorous analysis of literature, D&I guidelines, Responsible AI frameworks, and a simulated user study' overstates the evidence: the literature synthesis is real, but the simulated user study does not validate relevance or effectiveness as claimed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a 253-question bank for assessing the inclusivity of AI systems, organized into five pillars (Humans, Data, Process, System, Governance). The authors describe an iterative construction process that draws on D&I guidelines, a systematic literature review of D&I challenges, an existing Responsible AI question bank, and GPT-4o-generated question prompts. They then report a \"simulated user study\" in which 70 GPT-4o-generated personas answer five research questions about the relevance and usefulness of the question bank. The paper presents descriptive insights from this simulated evaluation, a comparison with prior AI question banks, and a discussion of threats to validity. The central claim is that the question bank, validated through this process, is an actionable tool for researchers, practitioners, and policymakers.","tokens_in":18948,"tokens_out":6685,"duration_ms":72074,"significance":"If the validation were sound, the question bank could be a genuinely useful practical resource for integrating D&I considerations into AI development and governance. The authors are transparent about their construction process, make a dataset publicly available ([37]), and address an underexplored niche relative to existing XAI and RAI question banks. The literature synthesis and the articulation of five assessment pillars have value. However, the paper's main evidential claim rests on a simulated evaluation in which the same large language model generated both many of the questions and all of the evaluator personas and responses. That is a self-referential loop that cannot support the claimed validity or effectiveness of the instrument. The paper's own limitation statement (Section VII.A) concedes that the question bank has not been tested in real-world projects, which directly undercuts the abstract's and conclusion's assertions of affirmation. The contribution is therefore contingent on a substantial reframing of the validation claim and on making the full instrument available for inspection.","major_comments":[{"comment":"The simulated user study is a same-model self-assessment. GPT-4o generated many of the questions (as described in Section III.B for V2, V4, and V6) and also generated the 70 personas and their answers to the research questions. The authors' manual review checks coherence and alignment with study objectives only; it cannot establish that the responses represent real AI practitioners' judgments. Section VII.A concedes that the question bank \"has not yet been tested in real-world AI development projects\" and that \"simulated feedback may not fully capture the complexities.\" These admissions directly contradict the abstract's characterization of the validation as rigorous and the Conclusion's statement (Section VIII) that the simulated user study findings \"affirm its relevance and effectiveness.\" The simulation is not a valid proxy for human evaluation, so the central validation claim is unsupported.","section":"Section IV, Steps 1-4; Section VII.A"},{"comment":"All reported findings—role-relevance frequencies in Table I, domain applicability themes, usefulness categories, and educational value—are derived from GPT-4o-generated persona responses. These are not empirical observations about AI practitioners; they are outputs of the same model that helped generate the question bank. The \"insights from Q1-Q4\" therefore describe the behavior of a language model, not the professional reasoning of data scientists, policy advisors, or UX designers. Consequently, the Conclusion's claim that the question bank is \"relevant and effective\" is not supported by the evidence presented in Section VI.B.","section":"Section VI.B, Table I and Conclusion"},{"comment":"The full 253-question instrument is not included in the manuscript. Section V provides only aggregate counts per pillar and a handful of illustrative examples, while the cited dataset ([37]) is described as the simulated user study data rather than the complete question bank. Since the central contribution is an actionable assessment tool, the complete question bank should be included in the paper or in a clearly labeled supplementary appendix. Without it, readers cannot evaluate the content, assess pillar coverage, or use the tool in practice, which undermines the paper's stated purpose.","section":"Section V; Section III.B"}],"minor_comments":[{"comment":"The manuscript is internally inconsistent about the final version number: Section III.B says the simulated study led to \"the development of V8,\" while Section IV Step 5 states \"we arrived at the final version (V9) of our question bank.\" Please reconcile the version numbering and clarify whether the 253-question count refers to V8, V9, or both.","section":"Section III.B V8 and Section IV Step 5"},{"comment":"The statement that using GPT \"ensur[ed] ... minimizing human bias in the question formation process\" is not substantiated; LLM-generated questions can readily encode biases from training data, and the authors' human-in-the-loop review is the actual bias-mitigation mechanism and should be credited as such.","section":"Section III.A"},{"comment":"The claims that \"our research is unique\" and that no existing study has proposed a structured question bank for inclusive AI are stronger than the later comparison in Section VI.D supports; please soften the introduction to acknowledge the adjacent XAI and RAI question banks while noting the distinct D&I focus.","section":"Section I"},{"comment":"Please add a note explaining how the \"frequency of relevant questions\" was computed from the persona responses, and include a total row or column so readers can interpret the pillar-wise counts in context.","section":"Section VI.B, Table I"},{"comment":"There are typographical and phrasing issues, such as \"that allows for a structured\" in Section III.B V1 and the repeated \"why do you think so?\" in the research questions; a careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The decisive issue is the circular validation loop: GPT-4o generated both the questions and the personas that evaluate them. This cannot support the paper's central validation claim. However, the underlying artifact—a structured question bank derived from D&I guidelines and prior RAI frameworks—may still be a useful contribution if the authors reframe the simulated study as an illustrative exercise rather than validation, remove the unsupported 'affirm' language, and make the full 253-question instrument available as a supplementary appendix. If the authors are unwilling to temper the validation claims, the paper would not be suitable for publication in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this paper presents a genuinely new artifact — a 253-question question bank for assessing AI inclusivity across five pillars (humans, data, process, system, governance). The gap is real; existing question banks (XAI, QB4AIRA, RAI) don't focus on D&I. The authors did real work synthesizing guidelines, a systematic review, and an RAI question bank, and the evolution through eight versions shows care. The five-pillar structure is sensible, and the examples read as concrete and useful.\n\nThe problem is the validation. The simulated user study uses GPT-4o to generate both the 70 personas and their responses to the research questions about the question bank — and GPT-4o also generated many of the questions in V2, V4, and V6. The manual review of the AI-generated responses checks coherence, but cannot establish that real practitioners would find the questions relevant or effective. The paper itself admits in Section VII.A that the QB has not been tested in real-world projects. So the abstract's phrase \"rigorous analysis... and a simulated user study\" overstates the evidence: the literature synthesis is fine, but the simulated study is a same-model self-assessment, not validation. The authors are transparent about this limitation in the threats to validity, which I respect, but it doesn't fix the problem.\n\nThe other soft spot is that the full question bank is not released in the paper or via a repository; only examples are shown. So as a reader I can't apply or audit the central artifact. The Zenodo dataset for the simulated personas and responses is available, but the QB itself is not.\n\nWhat's genuinely good: the problem matters, the synthesis of prior sources is real, and the question bank, once released, could be a useful checklist for AI teams. The circular validation is a load-bearing flaw in the current claims, but it's fixable: release the QB and run a small pilot with real AI practitioners.\n\nFor peer review: this deserves a serious referee, not because the validation holds, but because the artifact and the gap are worth engaging. A referee should push for real-world validation and QB release.","headline":"Genuinely useful question bank for AI inclusivity, but the same-model simulated validation can't support the paper's validation claims.","tokens_in":19503,"tokens_out":1851,"would_cite":false,"duration_ms":19014,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes a 253-question bank, organized into five pillars, for assessing AI inclusivity.","keywords":["diversity and inclusion","inclusive AI","AI question bank","responsible AI","bias mitigation","AI governance","AI ethics","simulated user study"],"falsifier":"Recruit real AI practitioners from the same 14 roles and several of the same domains, ask them the five research questions about the 253-question bank, and compare their judgments of relevance, clarity, and usefulness with the simulated personas' responses; a large divergence would show the simulated study does not by itself validate the bank.","tokens_in":18505,"feed_emoji":"📋","tokens_out":6381,"duration_ms":66881,"temperature":0.7,"pith_summary":"The paper aims to fill a gap it identifies in AI assessment tools: risk and explainability checklists exist, but none are designed specifically to measure inclusivity. To close that gap, it builds a structured bank of 253 questions organized under five pillars — Humans, Data, Process, System, and Governance — and argues this bank can be used to evaluate AI systems for diversity and inclusion throughout the lifecycle. The bank was assembled over eight versions from D&I guidelines, challenges found in a systematic literature review, an existing responsible-AI question bank, and questions generated by a large language model, then refined after a simulated study with 70 AI-generated personas. A sympathetic reader would care because the paper offers a concrete, standardized starting point for teams and regulators that currently lack one.","feed_headline":"A 253-question checklist can test AI inclusivity","feed_subtitle":"The bank spans humans, data, process, system, and governance, giving teams a concrete way to spot exclusion before deployment.","key_machinery":"The load-bearing object is the question bank itself: 253 questions organized under the five pillars of Humans, Data, Process, System, and Governance. The construction history is the argument — eight sequential versions that merge manual drafting from D&I guidelines, large-language-model prompts based on those guidelines, questions derived from a systematic review of D&I challenges, and fifteen questions from an existing responsible-AI question bank related to bias and fairness. The validation machinery is a simulated user study in which a large language model generated 70 personas over 14 AI-related roles, answered five research questions about the bank's relevance and usefulness, and produced feedback that the authors manually analyzed to refine the bank.","core_discovery":"The central claim is that AI inclusivity can be assessed with a dedicated question bank rather than left to general ethical principles or fairness metrics. The paper proposes that its 253 questions, mapped to the five pillars of humans, data, process, system, and governance, capture the relevant D&I considerations across the AI lifecycle, and that the simulated user study with 70 personas across 14 AI roles supports the bank's relevance, usefulness, educational value, and domain applicability. Feedback from the simulated personas led to refinements in seven questions, and the authors present this as validation that the bank is ready for adoption while acknowledging it has not yet been tested in real projects.","pith_inferences":["My inference: the real test of the bank is empirical; a deployment study with actual AI teams would reveal whether the questions are actionable or only sensible on paper.","My inference: because much of the content derives from existing guidelines and a large language model's expansion of them, the bank may be blind to exclusion patterns that are not yet documented in those sources; mining AI-incident reports could feed new questions.","My inference: the five-pillar structure invites aggregation into a maturity score or dashboard, which the paper lists only as future work but could be built directly from the current questions and would make the tool more useful to executives and regulators."],"forward_implications":["Organizations can use the 253 questions as a pre-deployment checklist to surface exclusion risks in people, data, development processes, system behavior, and governance.","Regulators and internal auditors get a common reference point for asking D&I questions that existing risk and explainability assessments do not cover.","The five-pillar organization lets different roles, such as data scientists, product managers, UX designers, and policy advisors, see which inclusivity concerns fall in their lane.","The bank can serve as an awareness and training tool, particularly for entry-level practitioners, by turning D&I principles into concrete yes/no questions."],"supporting_citations":[{"why":"Supplies the 46 D&I guidelines and the five-pillar structure that seed the question bank.","marker":"[13]"},{"why":"Provides the catalog of D&I challenges in AI that the authors translate into additional questions.","marker":"[1]"},{"why":"Supplies the responsible-AI question bank from which 15 bias and fairness questions are integrated for cross-validation.","marker":"[17]"},{"why":"Is the dataset of simulated user responses generated by the 70 personas that the paper analyzes as its validation evidence.","marker":"[37]"},{"why":"Establishes the earlier question-bank approach for explainability that the paper positions as covering a different need.","marker":"[14]"},{"why":"Is the AI risk assessment question bank used as the main comparison to show the new bank's inclusivity focus is distinct.","marker":"[16]"}],"fun_headline_variants":["253 questions to check AI inclusivity across five pillars","Five-pillar question bank targets AI inclusivity gaps","New question bank scores AI systems on inclusion","Assess AI inclusivity with 253 questions and five pillars"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that responses from 70 AI-generated personas accurately stand in for how real AI practitioners in those roles and domains would judge the question bank; if the simulated voices are not representative, the study's validation claim loses its support.","fun_headline_variants_meta":{"raw":{"variants":["253 questions to check AI inclusivity across five pillars","Five-pillar question bank targets AI inclusivity gaps","New question bank scores AI systems on inclusion","Assess AI inclusivity with 253 questions and five pillars"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000493,"raw_usage":{"total_tokens":2384,"prompt_tokens":868,"completion_tokens":1516,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":1456}},"tokens_in":484,"tokens_out":1516,"duration_ms":12876,"temperature":1.0,"reasoning_tokens":1456,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:15:25.858542+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit real AI practitioners from the same 14 roles and several of the same domains, ask them the five research questions about the 253-question bank, and compare their judgments of relevance, clarity, and usefulness with the simulated personas' responses; a large divergence would show the simulated study does not by itself validate the bank.","supporting_citations":[{"cited_title":"Diversity and Inclusion in Artificial Intelligence","cited_arxiv_id":"2305.12728","evidence_quote":"Supplies the 46 D&I guidelines and the five-pillar structure that seed the question bank."},{"cited_title":"Ai and the quest for diversity and inclusion: A systematic literature review,","cited_arxiv_id":null,"evidence_quote":"Provides the catalog of D&I challenges in AI that the authors translate into additional questions."},{"cited_title":"Responsible AI Question Bank: A Comprehensive Tool for AI Risk Assessment","cited_arxiv_id":"2408.11820","evidence_quote":"Supplies the responsible-AI question bank from which 15 bias and fairness questions are integrated for cross-validation."},{"cited_title":"The Simulates User Study Dataset of the Paper Titled “A Question Bank to Assess Inclusive AI","cited_arxiv_id":null,"evidence_quote":"Is the dataset of simulated user responses generated by the 70 personas that the paper analyzes as its validation evidence."},{"cited_title":"Questioning the ai: informing design practices for explainable ai user experiences,","cited_arxiv_id":null,"evidence_quote":"Establishes the earlier question-bank approach for explainability that the paper positions as covering a different need."},{"cited_title":"QB4AIRA: A Question Bank for AI Risk Assessment","cited_arxiv_id":"2305.09300","evidence_quote":"Is the AI risk assessment question bank used as the main comparison to show the new bank's inclusivity focus is distinct."}],"review_version":1}