{"id":"9a366d96-35f3-47ae-bae6-6bbe34968c64","arxiv_id":"2506.17095","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A single-company interview study of 22 practitioners shows AI fairness testing is ad hoc, data-scientist-driven, and hampered by data quality, time pressure, and missing tools.","lead":"Researchers interviewed 22 software professionals at one large South American company about how they test AI systems for fairness. They found that fairness testing is mostly ad hoc, led by data scientists, and limited by data quality, time pressure, and a lack of dedicated tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-company sample cannot support the industry-level framing; acknowledged non-generalizability conflicts with the broad claims in the abstract and findings.","rationale":"The paper is a carefully conducted single-case qualitative study, and the data collection, coding, and triangulation are described in sufficient detail. The problem is not internal validity but external validity: the title, abstract, and findings make claims about 'industry' practice while the evidence comes from 22 practitioners at one company. The reader's weakest_assumption names exactly this gap. This is a framing issue rather than a fatal flaw; the method supports claims about the case, and the authors do acknowledge transferability limits in Section V-D. The conditional verdict is therefore appropriate: the paper should either narrow its claims to the studied organization or add comparative evidence across companies and regions. Reporting theme-prevalence counts would also make the 'mostly ad hoc' claim more transparent. Because my concern confirms the reader's identified weakest assumption without moving to a different verdict, no adjustment is needed.","tokens_in":14705,"tokens_out":3558,"duration_ms":42908,"concrete_test":"Apply the same semi-structured interview guide (Table I) to 20-30 practitioners distributed across at least five companies in different countries and sectors, including at least one organization with an AI-governance or regulatory-compliance function. Pre-register the criterion: if two or more of the external companies describe documented, standardized fairness-testing processes or dedicated fairness-testing roles, then the 'ad hoc / data-scientist-driven' finding does not generalize and the paper should be re-scoped to the original company; if no external company reports such standardized practice, the external-validity concern is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central assertion is that fairness testing in industry relies on four key strategies and is mostly ad hoc because formalized guidelines are lacking (Findings 2, Section IV-B). That claim requires the sample to represent industry practice. All 22 interviews come from one large South American software company (Section III-A), across four projects inside that single organization. Saturation was reached within that one company, but within-case saturation does not establish saturation across companies, sectors, or regulatory environments. Section V-D concedes the findings are 'not intended for statistical generalization' and are only 'transferable to similar settings,' yet the abstract, Findings summaries, and conclusions repeatedly assert industry-wide practice ('Fairness testing in industry relies on four key strategies...'). If other organizations have formalized fairness-testing workflows, dedicated QA roles, or compliance-driven checklists, the headline finding would not hold there; the single case cannot rule that out. The paper also reports no prevalence counts, so 'mostly developed on an ad-hoc basis' cannot be verified from the presented evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative case study of fairness testing practice: the authors conducted 22 semi-structured interviews with practitioners across four AI/ML projects inside one large South American software company, supplemented by project artifacts. Thematic analysis, assisted by GPT-4 Omni for open coding, yielded three fairness-testing targets (diversity/inclusivity, model consistency, data representation/balancing), four strategies (continuous iteration, data manipulation/simulation, well-defined metrics, specific tools), and five challenges (data quality/diversity, time constraints, missing tools/metrics, knowledge gaps, black-box behavior). The paper claims that fairness testing in industry is largely ad hoc and practitioner-driven, and that a gap exists between academic fairness-testing research and industrial practice. The manuscript includes an interview guide, the data-extraction prompt, demographic tables, a triangulation description, a threats-to-validity section, and a data-availability link for anonymized quotations.","tokens_in":14837,"tokens_out":2867,"duration_ms":35318,"significance":"If the findings are taken as a case study, the paper is a useful and reasonably detailed empirical contribution: it documents how one organization's teams approach fairness testing, provides concrete practitioner quotes, and offers actionable implications for tools and guidelines. The method is transparent in several respects that strengthen trust in the data: the interview guide is included, the GPT-assisted coding prompt is reproduced, the sampling strategy is described as convenience plus snowball plus theoretical sampling, and a sample of extracted quotes was manually verified. However, the paper's central contribution is weakened by a mismatch between the evidence base—one company, 22 self-selected participants—and the repeated 'industry' framing in the abstract, key findings, and conclusions. The load-bearing claims about how industry at large tests fairness would need either broader evidence or a systematic reframing to the level of a case study.","major_comments":[{"comment":"The central claim, stated as 'Fairness testing in industry relies on four key strategies, mostly developed on an ad-hoc basis due to the lack of formalized guidelines' (Section IV-B, Findings 2 Summary), is not supported by the evidence described in the paper. Section III-A identifies the case as a single large South American company, and Section V-D explicitly states that the findings are 'not intended for statistical generalization' and are only 'transferable to similar settings.' Yet the abstract, the Findings summaries, and the Conclusions repeatedly generalize to 'industry.' This mismatch is load-bearing because the paper's main contribution is an empirical statement about industrial practice. I recommend reframing the abstract, findings, and conclusions to refer to 'the studied organization' or 'the case,' and adding an explicit sentence that the study is a single-company case study whose transferability to other companies, sectors, and regulatory environments remains open.","section":"IV-B / Abstract / VI"},{"comment":"The paper uses quantifier-like language such as 'mostly developed on an ad-hoc basis,' 'primarily led by data scientists,' and 'recurring challenges' without reporting how many of the 22 participants reported each strategy, role, or challenge. Tables V and VI provide representative quotes but not prevalence counts. As a result, the reader cannot verify whether 'mostly' or 'primarily' is a faithful summary of the dataset. The manuscript should either report theme prevalence (e.g., number or percentage of participants mentioning each strategy/challenge) or replace these quantifiers with hedged formulations such as 'in the reported experiences of participants.'","section":"IV-B / Tables V and VI"},{"comment":"The paper states that data saturation was reached, but saturation was reached within one organization. Within-case saturation does not establish saturation across companies, sectors, or regulatory contexts, so it does not strengthen the external validity of the 'industry' claims. Section III-C should clarify that saturation was reached within the studied case only, and Section V-D should acknowledge that the transferability argument rests on analytic generalization rather than on saturation across contexts.","section":"III-C / V-D"}],"minor_comments":[{"comment":"Reference [25] has the same title as [24] but appears intended to refer to the smart lipstick project mentioned in Project D; please correct the title and URL so each reference matches its citation.","section":"III-A / References [24] and [25]"},{"comment":"The email address for Cleyton Magalhaes appears to contain a typo ('cleyton.vanut.ufrpe.br'); please verify and correct it.","section":"Author affiliation"},{"comment":"The Figshare data-availability URL is broken across a line break in the text; provide a single, clickable URL.","section":"VII"},{"comment":"The phrase 'The interviews about testing, held between June 1 and July 5, 2024' is awkward; consider 'The interviews, held between June 1 and July 5, 2024, lasted between 15 and 25 minutes.'","section":"III-C"},{"comment":"In Section V-A, 'could provide valuable results into best practices' should be 'could provide valuable insights into best practices.'","section":"V-A"},{"comment":"The headings 'Finding 1 – Summary' and 'Findings 2 – Summary' and 'Findings 3 – Summary' are inconsistent in number; use 'Finding 2' and 'Finding 3' for consistency.","section":"IV (Finding headings)"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely within scope for a software-engineering empirical venue, and the qualitative data are presented transparently. The main issue is a framing problem rather than a fundamental flaw in the study: the authors have a single-company case study but report it as if it described industry-wide practice. If the authors systematically reframe the claims and add prevalence information, the paper could become a solid case-study contribution. I would not recommend rejection, because the underlying empirical work is real and the limitations are acknowledged in the threats-to-validity section; the needed changes are substantial but achievable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Think of this as a well-executed single-company case study wearing an industry-wide jacket. The 22 interviews from one large South American software company are reported in enough detail to trust: you get the interview guide, the GPT coding prompt, a manual verification step, and a figshare link to anonymized quotes. That is real discipline for a qualitative paper, and the findings—three fairness-testing targets, four mostly ad hoc strategies, five recurring challenges—are a useful synthesis that matches what earlier surveys and smaller studies have suggested.\n\nThe genuinely new part is the practitioner-level texture: data scientists leading iteration loops, testers defining scenarios but not driving the process, GANs and GPTs used for data augmentation, and the honest admissions about concept drift and hallucinated model outputs. The thematic categories are original enough and arrived at transparently.\n\nThe soft spot is exactly what the stress-test flags: the paper uses 'industry' in the title, abstract, and finding summaries ('Fairness testing in industry relies on four key strategies...') while the evidence is one company with 22 self-selected participants. The authors do include a limitations paragraph saying the findings are 'not intended for statistical generalization' and only 'transferable to similar settings,' so the acknowledgement is there. But the framing should be tightened: either narrow the claim to 'the studied company' or add comparative evidence. I'd also like to see prevalence counts for the themes—when the text says 'mostly ad hoc,' a count or at least a quote-density indication would turn that from impression to evidence.\n\nNone of this is fatal. The paper is honest, the method is documented, and the data is shared. The circularity burden is low; this is not a paper that fits its own prior work into a self-confirming loop. The citation pattern is fine.\n\nWho should read this: people studying how fairness research does or does not reach practice, and tool builders looking for concrete requirements. It deserves a serious referee; a competent reviewer can ask for the frame to be adjusted and the prevalence data added, but the core empirical material stands.","headline":"A transparent single-company case study with a framing mismatch: the interview data are solid and the synthesis is useful, but the title and abstract overclaim industry-wide applicability.","tokens_in":15361,"tokens_out":2308,"would_cite":true,"duration_ms":28273,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fairness testing in AI is ad hoc and data-scientist-driven, a 22-practitioner case study finds.","keywords":["software fairness","fairness testing","AI bias","machine learning testing","case study","qualitative research","software engineering practice","ad hoc testing"],"falsifier":"A multi-company observational study or survey that finds dedicated fairness testing roles, formalized fairness guidelines, and standardized fairness testing procedures in routine use would contradict the claim that practice is ad hoc and primarily driven by data scientists. A narrower check: a single team outside this company that can show a documented, repeatable fairness testing process with defined oracles and coverage criteria applied across multiple projects already weakens the 'lack of formalized guidelines' generalization.","tokens_in":14503,"feed_emoji":"⚖️","tokens_out":6712,"duration_ms":71426,"temperature":0.7,"pith_summary":"This paper reports a case study of 22 practitioners at one large South American software company who build AI systems across four projects: sign-language translation, energy digital twins, LLM-based education, and facial recognition. It claims that fairness testing in industry is not standardized: teams test for diversity and inclusivity, model consistency, and data representation, and they work through continuous iteration, data manipulation, metrics, and a few generative-model tools, with decisions made mainly by data scientists rather than dedicated testers. The study also reports five recurring challenges, including poor data quality, time pressure, missing tools and metrics, team knowledge gaps, and black-box model behavior. The result is a map of the gap between fairness testing research and actual practice, intended to steer researchers toward tools and guidelines that fit existing workflows.","feed_headline":"AI fairness testing in industry is ad hoc, 22 practitioners reveal","feed_subtitle":"Teams check diversity, consistency, and data balance, but lack standardized guidelines and dedicated testing roles.","key_machinery":"The machinery is a qualitative case study: semi-structured interviews with 22 practitioners from four AI projects at one large South American software company, analyzed through three-phase thematic analysis with an LLM-assisted open coding pass, manual axial coding, and selective coding, plus cross-case and data triangulation. This design converts practitioner quotes into the paper's central organizing result: a taxonomy of three fairness testing targets, four ad hoc strategies, and five recurring challenges.","core_discovery":"On the paper's own terms, the central discovery is that fairness testing is happening in industry, but it is ad hoc and practitioner-driven rather than formalized. From interviews with 22 professionals across four AI projects, the authors identify three targets that teams actually test for—diversity and inclusivity, model consistency, and data representation and balancing—and four strategies: continuous iteration and adjustment, manipulation and simulation of data, use of well-defined metrics such as accuracy and A/B tests, and use of specialized tools such as GANs and GPT-based models. Fairness decisions were made mainly by data scientists and programmers, with limited involvement of dedicated testing professionals, and no project followed a formalized fairness testing process. The authors frame this as a gap between a rich academic literature on fairness testing and an industry that lacks clear guidelines, accessible tools, and dedicated roles.","pith_inferences":["A testable extension the paper does not run: applying the same interview protocol across multiple companies and regions; if the four-strategy taxonomy reproduces, the ad hoc characterization would hold beyond this single case.","The authors report use of GANs and GPT for data augmentation but do not claim these are dedicated fairness tools; one implication is that practitioners satisfy fairness needs with general ML tooling, so dedicated fairness tools may need to wrap existing pipelines to be adopted.","The finding that testers were minimally involved suggests a possible blind spot: if QA professionals are bypassed in AI fairness work, fairness coverage may depend on whatever data scientists happen to check, a dynamic the study describes but does not measure.","The paper names the absence of testing for extreme data distribution changes as a gap; a concrete follow-up would be developing test oracles for concept drift, not proposed in the study."],"forward_implications":["Academic fairness definitions and testing frameworks will not transfer to practice until they are translated into concrete workflow steps, because the studied teams did not use them.","Fairness testing tools should be designed for data scientists' existing workflows, since fairness decisions in the studied projects were made mostly by data scientists and programmers rather than dedicated testers.","Because time pressure was a recurring barrier, fairness testing that cannot be automated or performed quickly is likely to be dropped, so CI/CD integration is a natural target.","Practitioners already use GANs and GPT-based models for data augmentation; fairness tooling that builds on these familiar tools may be adopted more readily than standalone fairness libraries."],"supporting_citations":[{"why":"supplies the academic fairness-testing framework (definitions, steps, components, tools) that the study contrasts with observed industry practice","marker":"[11]"},{"why":"provides prior evidence that fairness testing tools have limited industry adoption, which the interview findings echo","marker":"[12]"},{"why":"supplies the case study method and reporting guidelines the study follows","marker":"[20]"},{"why":"grounds fairness testing as testing software for discrimination, the phenomenon the practitioners describe","marker":"[9]"},{"why":"defines software fairness, the core construct the interview guide probes","marker":"[6]"}],"fun_headline_variants":["AI fairness testing is ad hoc, 22 practitioners say","Fairness testing in industry lacks standards and roles","22 practitioners: fairness testing is informal and unguided","Fairness testing in AI: ad hoc, no formal process"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that fairness testing in industry is ad hoc rests on the assumption that 22 self-selected practitioners from one large South American company represent software professionals in general; the paper says its findings are not intended for statistical generalization, but its title and conclusions treat them as industry practice.","fun_headline_variants_meta":{"raw":{"variants":["AI fairness testing is ad hoc, 22 practitioners say","Fairness testing in industry lacks standards and roles","22 practitioners: fairness testing is informal and unguided","Fairness testing in AI: ad hoc, no formal process"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000215,"raw_usage":{"total_tokens":1413,"prompt_tokens":912,"completion_tokens":501,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":436}},"tokens_in":528,"tokens_out":501,"duration_ms":5782,"temperature":1.0,"reasoning_tokens":436,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:30:34.398124+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A multi-company observational study or survey that finds dedicated fairness testing roles, formalized fairness guidelines, and standardized fairness testing procedures in routine use would contradict the claim that practice is ad hoc and primarily driven by data scientists. A narrower check: a single team outside this company that can show a documented, repeatable fairness testing process with defined oracles and coverage criteria applied across multiple projects already weakens the 'lack of formalized guidelines' generalization.","supporting_citations":[{"cited_title":"Fairness testing: A comprehensive survey and analysis of trends,","cited_arxiv_id":null,"evidence_quote":"supplies the academic fairness-testing framework (definitions, steps, components, tools) that the study contrasts with observed industry practice"},{"cited_title":"From literature to practice: Exploring fairness testing tools for the software industry adoption,","cited_arxiv_id":null,"evidence_quote":"provides prior evidence that fairness testing tools have limited industry adoption, which the interview findings echo"},{"cited_title":"Guidelines for conducting and reporting case study research in software engineering,","cited_arxiv_id":null,"evidence_quote":"supplies the case study method and reporting guidelines the study follows"},{"cited_title":"Fairness testing: testing software for discrimination,","cited_arxiv_id":null,"evidence_quote":"grounds fairness testing as testing software for discrimination, the phenomenon the practitioners describe"},{"cited_title":"Software fairness,","cited_arxiv_id":null,"evidence_quote":"defines software fairness, the core construct the interview guide probes"}],"review_version":1}