{"id":"d6c02af0-3f2b-4ce8-ae5a-5434fd1f14c6","arxiv_id":"2508.11067","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A ten-year review of 189 AI bias papers finds most do not define bias, most center gender, and few propose real-world fixes.","lead":"This paper reviews 189 AI and LLM bias papers from four top venues and finds that most never define bias and most focus on gender. It gives researchers and funders a map of what fairness research studies, what it ignores, and why it rarely reaches real systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 82% 'no working definition' statistic is undermined by the abstract's own wording: the 155 papers are described as establishing 'only mathematical and technical definition of bias,' which is a definition, not an absence of one.","rationale":"This concern is load-bearing because the central claim of a conceptually thin, gender-centric field rests on the 82% 'no working definition' figure. I read the abstract in good faith: the gender and implementation numbers are plausible, and the corpus size is reasonable. But the contradiction in the definitional claim means the central statistic is not trustworthy as reported. This is not a disagreement with the field's consensus; it is an internal inconsistency between the stated category and the evidence cited for it. The reader's weakest assumption already flagged coding reliability; my check would settle whether the 82% is a finding or an artifact of the rubric. If recoding confirms most of the 155 have no definition at all, the claim stands. If most have technical definitions, the paper needs revision. Given this, I would not accept the central claim as is; the verdict should be conditional on the availability of a codebook and a recoded breakdown.","tokens_in":20230,"tokens_out":7256,"duration_ms":70983,"concrete_test":"Take a random sample of 30 papers from the 155 classified as 'no working definition.' Have two independent annotators, blind to the paper's categories, classify each paper's treatment of bias as: (a) no explicit definition; (b) an explicit mathematical/technical definition only; (c) a social/contextual definition with or without a technical definition. Compute Cohen's kappa on the three-way classification and report the distribution. Also check, for each sample paper coded (b), whether the abstract or method section contains a formal definition of bias, such as a bias metric. If a substantial fraction are (b), the 82% headline must be restated as 'papers rarely offer broader, socially contextualized definitions,' and the claim that papers lack a working definition is not supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing step is the binary coding of 'working definition.' The abstract says 155/189 papers 'did not establish a working definition of bias' and immediately characterizes the same papers as 'establishing only mathematical and technical definition of bias.' A mathematical/technical definition is still a working definition unless the coding rubric explicitly requires a social/contextual definition. The appeal to Blodgett et al. (2020), described by the authors as recommending that researchers go 'beyond strictly technical definitions,' confirms that the intended contrast is technical versus non-technical, not defined versus undefined. As written, the 82% figure blends two different observations: (1) no definition of bias appears anywhere in the paper, and (2) a definition appears but is only technical/mathematical. The title and abstract use the first reading ('did not establish a working definition') while the evidence described supports only the second. Without a codebook distinguishing these categories, the headline prevalence and the temporal trend (94/155 in the last five years) are not neutral measurements; they encode the authors' preference for socially situated definitions. This is the weakest link because the other headline statistics (79.9% gender focus, 10.6% implementation recommendations) depend on more straightforward presence/absence coding, whereas 'working definition' is a judgment call. The supplied full text is unreadable, so the presence of a codebook or inter-annotator agreement cannot be confirmed from the material provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic literature review of 189 AI/LLM bias papers drawn from ACL, FAccT, NeurIPS, and AAAI over the past 10 years. The central claims are that 82% (155/189) of these papers did not establish a working definition of \"bias,\" that 79.9% (151/189) focus on gender bias, and that only 10.6% (20/189) include recommendations for real-world implementation. The authors further report lower coverage of race/ethnicity, age, religion, and nationality, and use these patterns to argue that the field has a narrow, gender-centric conception of bias and a weak academia-industry link. The abstract also situates the findings against Blodgett et al. (2020) and recommends stronger coverage of marginalized and non-Western populations.","tokens_in":20457,"tokens_out":3412,"duration_ms":39259,"significance":"If the quantitative findings were independently reproducible, this would be a useful descriptive contribution to the AI fairness literature: the identity-axis breakdown and the academia-industry gap are practically important, and the arithmetic in the abstract is internally consistent (155/189, 151/189, 20/189, 57/189, 39/189, 36/189, and 25/189 round to the stated percentages). The venue selection is appropriate, and the paper's call for broader demographic coverage is well motivated. However, the significance is conditional: the headline \"working definition\" statistic conflates two different observations, and no search protocol, coding manual, inter-annotator agreement, or corpus list is visible in the material supplied. As submitted, the results cannot be verified or reproduced, which limits the paper's contribution to a preliminary descriptive claim.","major_comments":[{"comment":"The 82% claim is internally ambiguous: the abstract first says 155/189 papers 'did not establish a working definition of bias,' but then characterizes the same papers as 'establishing only mathematical and technical definition of bias.' A mathematical/technical definition is still a definition, so the headline conflates 'no definition at all' with 'only a technical definition.' The paper must define the coding categories separately, report counts for each, and adjust the wording so the reader knows which claim is actually being made.","section":"Abstract"},{"comment":"No search strategy, inclusion/exclusion criteria, screening process, coding protocol, inter-annotator agreement, or uncertainty estimates are reported in the material supplied. The extracted full text is not readable in the version provided, so I could not locate a methods section, a corpus list, or a codebook. For a literature review whose entire contribution is a set of prevalence percentages, the absence of these methodological details is load-bearing; they must be added or the percentages cannot be assessed.","section":"Abstract / Methods"},{"comment":"The claim that only 10.6% of papers 'include recommendations for how to implement their findings or contributions in real-world AI systems or design processes' depends on an unspecified boundary for what counts as a recommendation. It is unclear whether a released debiasing method, a code repository, an evaluation benchmark, or a design guideline would qualify. The coding definition should be stated, and examples of both counted and not-counted cases should be provided.","section":"Abstract"},{"comment":"The temporal claim that 94 of the 155 papers appeared in the last five years is presented as evidence that the field persisted after Blodgett et al. (2020), but the abstract gives no denominator of total corpus papers per year. If publication volume grew over the period, the raw count of 94/155 would be expected even with no change in the rate of technical-only definitions. The paper should report per-year proportions or a comparison with the overall publication rate.","section":"Abstract"}],"minor_comments":[{"comment":"The final sentence contains an incomplete grammatical structure: 'especially since many of the biases that our corpus contains several successful mitigation methods that still persist within the outputs of AI systems' is not a complete clause and should be rewritten.","section":"Abstract"},{"comment":"The percentages are reported with varying decimal precision (82% vs. 79.9%). Please use a consistent number of decimal places for all reported proportions and confirm that each percentage matches the stated fraction.","section":"Abstract"},{"comment":"The text extract provided to the referee is heavily corrupted by character-encoding issues. Please ensure that the submitted version is a readable PDF or text file so that the methods, tables, and references can be independently checked.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The supplied full-text file for this manuscript is not readable due to encoding corruption; if this reflects what was submitted, the authors should be asked to provide a clean version before further review. The more substantive concern is the ambiguity in the 'working definition' coding, which affects the paper's central statistic and should be resolved before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading this one. The headline 82% figure is muddier than it looks: the abstract says 155/189 papers 'did not establish a working definition' and then immediately says those same papers offered 'only mathematical and technical definition of bias.' That is a definition, just not the kind the authors prefer. Unless the codebook explicitly separates 'no definition anywhere' from 'technical-only definition,' the statistic blends two different observations. The stress-test note is right about this, and it is the load-bearing claim of the paper.\n\nThe genuinely useful part is the quantitative sweep: 189 papers across four venues, with counts of identity axes and implementation recommendations. Extending Blodgett et al. to this scale is a legitimate contribution, and the gender-focus (79.9%) and implementation-gap (10.6%) findings are based on more straightforward presence/absence coding, so they are likely more robust. If those numbers hold up, the 'narrow conceptions' argument has real teeth.\n\nSoft spots, in order of severity. First, the definitional conflation I already mentioned; the title overstates the evidence ('Bias is a Math Problem' is not what the data actually show). Second, the supplied full text is unreadable encoding, so I cannot verify whether a coding protocol, inter-annotator agreement, or even a corpus list exists. That is not a flaw in the authors' work per se, but it means I cannot audit the methods from what we have. Third, the four-venue sample with unspecified search/screening is a narrow slice of 'AI/LLM bias research'; generalizing to the whole field is a stretch.\n\nWho this is for: AI fairness researchers, reviewers, and funders who want a compact picture of the field's focus. It deserves a serious referee, but only with major revision. The authors should make the codebook explicit, recode the definition variable into at least three categories (none, technical-only, contextual), release the corpus and the coding, and soften the title. If that happens, it could be a useful reference. As it stands, I would not cite the 82% number in my own work.\n\nRecommendation: send it to peer review with a clear request for methodological transparency. This is not a desk reject; it is a revision problem.","headline":"A useful corpus-level extension of Blodgett et al. that currently undermines its own 82% stat by conflating 'no definition' with 'technical-only definition'; deserves review after a coding overhaul.","tokens_in":21008,"tokens_out":3156,"would_cite":false,"duration_ms":33221,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A review of 189 AI bias papers finds most never define bias, most focus on gender, and few reach implementation.","keywords":["systematic literature review","AI bias","LLM bias","gender bias","fairness research","working definition","debiasing","academia-industry gap"],"falsifier":"A second coding team could re-run the review with a written screening protocol and inter-annotator agreement measures; if the shares of no-definition or gender-focused papers move substantially, the rates are not stable. A simpler check is to look at the 155 papers coded as having no working definition and ask whether any define bias operationally through their evaluation metric, since that would change the 82% count.","tokens_in":19994,"feed_emoji":"📊","tokens_out":8151,"duration_ms":85449,"temperature":0.7,"pith_summary":"This paper tries to establish that a decade of research on bias in AI and large language models is conceptually thin and lopsided. Reading 189 papers from four premier venues, the authors find that 82 percent never state a working definition of bias and 79.9 percent concentrate on gender bias, with race, age, religion, and nationality far behind. Only 10.6 percent of the papers suggest how their results could be put to use in real AI systems. If this picture is right, a large part of the fairness literature is measuring a narrow slice of bias without saying what bias means, which weakens both scientific comparison and practical impact.","feed_headline":"82% of AI bias papers skip a working definition","feed_subtitle":"Across 189 papers from four top venues, gender dominates and only one in ten gives implementation advice.","key_machinery":"The load-bearing device is a hand-coded corpus: 189 papers, each labeled for whether it gives a working definition of 'bias,' which identity axes it targets, and whether it offers implementation recommendations. The corpus labels are the machinery because every reported percentage is just a tally over those labels, and the paper's conclusions about gender-centrism and the academia-industry gap come directly from those tallies.","core_discovery":"The paper's central claim is that AI/LLM bias research, as sampled from four leading venues over the past ten years, is dominated by gender-centered work and lacks a shared conceptual foundation. The headline numbers are 82% (155/189) with no working definition of bias, 79.9% (151/189) focusing on gender, and 10.6% (20/189) including real-world implementation recommendations. The authors interpret these patterns as evidence that bias is treated as a mathematical or technical problem first, with social and cultural dimensions secondary, and they argue the field needs broader coverage of marginalized non-Western communities and stronger links to practice.","pith_inferences":["A testable extension would be to code the same 189 papers for whether they define bias operationally through a metric, such as an association test or a disparity score, even when they do not give a verbal definition; this could lower the 82% figure.","The gender-heavy focus may partly follow from the availability of standard word-embedding and occupation datasets, which create path dependence; the paper does not test this, but the pattern it documents is consistent with it.","The paper's coding categories could be re-run on a broader set of venues or on industry-facing practice reports; if the rates persist, the academia-industry gap would be confirmed outside the four selected venues."],"forward_implications":["If the rates hold, a paper on AI bias is far more likely to be counted as about gender than about race, age, religion, or nationality, so broad statements about 'AI bias' are generalizing from a narrow slice of identity.","The 82% figure implies that most papers state that stereotypes exist and then move to technical mitigation without saying what would count as bias in their setting, which makes results hard to compare across papers.","The 10.6% implementation figure implies that even papers that successfully debias a model often stop at the model, so persistent real-world harms may not be addressed by the research.","A direct corollary is that future reviewers and venues can require a working definition of bias and an implementation section, which would shift the field's composition."],"supporting_citations":[{"why":"The prior literature review the authors build on, whose finding that NLP research often lacks substantive definitions of bias frames the 82% result and the post-2020 comparison.","marker":"Blodgett et al. (2020)"}],"fun_headline_variants":["Bias research blind spot: 82% lack definition, 80% fixate on gender","AI bias literature: gender-centric, definition-light, implementation-poor","10-year AI bias review: 4 in 5 papers target gender, 1 in 10 actionable","AI bias studies: 80% gender-focused, 82% undefined, 10% applied"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The percentages assume that the 189 papers chosen from four venues, under an unstated screening process, stand for AI/LLM bias research as a whole, and that the authors' binary coding of 'working definition' and 'gender focus' is reliable and reproducible.","fun_headline_variants_meta":{"raw":{"variants":["Bias research blind spot: 82% lack definition, 80% fixate on gender","AI bias literature: gender-centric, definition-light, implementation-poor","10-year AI bias review: 4 in 5 papers target gender, 1 in 10 actionable","AI bias studies: 80% gender-focused, 82% undefined, 10% applied"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000846,"raw_usage":{"total_tokens":3811,"prompt_tokens":1205,"completion_tokens":2606,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":821,"completion_tokens_details":{"reasoning_tokens":2511}},"tokens_in":821,"tokens_out":2606,"duration_ms":17366,"temperature":1.0,"reasoning_tokens":2511,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:28:20.520221+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A second coding team could re-run the review with a written screening protocol and inter-annotator agreement measures; if the shares of no-definition or gender-focused papers move substantially, the rates are not stable. A simpler check is to look at the 155 papers coded as having no working definition and ask whether any define bias operationally through their evaluation metric, since that would change the 82% count.","supporting_citations":[],"review_version":2}