REVIEW 3 major objections 6 minor 36 references
Software Fairness Testing in Practice
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Fairness testing in AI is ad hoc and data-scientist-driven, a 22-practitioner case study finds.
desk verdict A transparent single-company case study with a framing mismatch: the interview data are solid and the synthesis is useful, but the title and abstract overclaim industry-wide applicability. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a qualitative case study: semi-structured interviews with 22 practitioners from four AI projects at one large South American software company, analyzed through three-phase thematic analysis with an LLM-assisted open coding pass, manual axial coding, and selective coding, plus cross-case and data triangulation. This design converts practitioner quotes into the paper's central organizing result: a taxonomy of three fairness testing targets, four ad hoc strategies, and five recurring challenges.
What would settle it
A multi-company observational study or survey that finds dedicated fairness testing roles, formalized fairness guidelines, and standardized fairness testing procedures in routine use would contradict the claim that practice is ad hoc and primarily driven by data scientists. A narrower check: a single team outside this company that can show a documented, repeatable fairness testing process with defined oracles and coverage criteria applied across multiple projects already weakens the 'lack of formalized guidelines' generalization.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that fairness testing is happening in industry, but it is ad hoc and practitioner-driven rather than formalized. From interviews with 22 professionals across four AI projects, the authors identify three targets that teams actually test for—diversity and inclusivity, model consistency, and data representation and balancing—and four strategies: continuous iteration and adjustment, manipulation and simulation of data, use of well-defined metrics such as accuracy and A/B tests, and use of specialized tools such as GANs and GPT-based models. Fairness decisions were made mainly by data scientists and programmers, with limited involvement of dedicated testing professionals, and no project followed a formalized fairness testing process. The authors frame this as a gap between a rich academic literature on fairness testing and an industry that lacks clear guidelines, accessible tools, and dedicated roles.
Load-bearing premise
The claim that fairness testing in industry is ad hoc rests on the assumption that 22 self-selected practitioners from one large South American company represent software professionals in general; the paper says its findings are not intended for statistical generalization, but its title and conclusions treat them as industry practice.
Editorial extensions
If this is right
- Academic fairness definitions and testing frameworks will not transfer to practice until they are translated into concrete workflow steps, because the studied teams did not use them.
- Fairness testing tools should be designed for data scientists' existing workflows, since fairness decisions in the studied projects were made mostly by data scientists and programmers rather than dedicated testers.
- Because time pressure was a recurring barrier, fairness testing that cannot be automated or performed quickly is likely to be dropped, so CI/CD integration is a natural target.
- Practitioners already use GANs and GPT-based models for data augmentation; fairness tooling that builds on these familiar tools may be adopted more readily than standalone fairness libraries.
Reading between the lines
- A testable extension the paper does not run: applying the same interview protocol across multiple companies and regions; if the four-strategy taxonomy reproduces, the ad hoc characterization would hold beyond this single case.
- The authors report use of GANs and GPT for data augmentation but do not claim these are dedicated fairness tools; one implication is that practitioners satisfy fairness needs with general ML tooling, so dedicated fairness tools may need to wrap existing pipelines to be adopted.
- The finding that testers were minimally involved suggests a possible blind spot: if QA professionals are bypassed in AI fairness work, fairness coverage may depend on whatever data scientists happen to check, a dynamic the study describes but does not measure.
- The paper names the absence of testing for extreme data distribution changes as a gap; a concrete follow-up would be developing test oracles for concept drift, not proposed in the study.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a qualitative case study of fairness testing practice: the authors conducted 22 semi-structured interviews with practitioners across four AI/ML projects inside one large South American software company, supplemented by project artifacts. Thematic analysis, assisted by GPT-4 Omni for open coding, yielded three fairness-testing targets (diversity/inclusivity, model consistency, data representation/balancing), four strategies (continuous iteration, data manipulation/simulation, well-defined metrics, specific tools), and five challenges (data quality/diversity, time constraints, missing tools/metrics, knowledge gaps, black-box behavior). The paper claims that fairness testing in industry is largely ad hoc and practitioner-driven, and that a gap exists between academic fairness-testing research and industrial practice. The manuscript includes an interview guide, the data-extraction prompt, demographic tables, a triangulation description, a threats-to-validity section, and a data-availability link for anonymized quotations.
Significance. If the findings are taken as a case study, the paper is a useful and reasonably detailed empirical contribution: it documents how one organization's teams approach fairness testing, provides concrete practitioner quotes, and offers actionable implications for tools and guidelines. The method is transparent in several respects that strengthen trust in the data: the interview guide is included, the GPT-assisted coding prompt is reproduced, the sampling strategy is described as convenience plus snowball plus theoretical sampling, and a sample of extracted quotes was manually verified. However, the paper's central contribution is weakened by a mismatch between the evidence base—one company, 22 self-selected participants—and the repeated 'industry' framing in the abstract, key findings, and conclusions. The load-bearing claims about how industry at large tests fairness would need either broader evidence or a systematic reframing to the level of a case study.
major comments (3)
- [IV-B / Abstract / VI] The central claim, stated as 'Fairness testing in industry relies on four key strategies, mostly developed on an ad-hoc basis due to the lack of formalized guidelines' (Section IV-B, Findings 2 Summary), is not supported by the evidence described in the paper. Section III-A identifies the case as a single large South American company, and Section V-D explicitly states that the findings are 'not intended for statistical generalization' and are only 'transferable to similar settings.' Yet the abstract, the Findings summaries, and the Conclusions repeatedly generalize to 'industry.' This mismatch is load-bearing because the paper's main contribution is an empirical statement about industrial practice. I recommend reframing the abstract, findings, and conclusions to refer to 'the studied organization' or 'the case,' and adding an explicit sentence that the study is a single-company case study whose transferability to other companies, sectors, and regulatory environments remains open.
- [IV-B / Tables V and VI] The paper uses quantifier-like language such as 'mostly developed on an ad-hoc basis,' 'primarily led by data scientists,' and 'recurring challenges' without reporting how many of the 22 participants reported each strategy, role, or challenge. Tables V and VI provide representative quotes but not prevalence counts. As a result, the reader cannot verify whether 'mostly' or 'primarily' is a faithful summary of the dataset. The manuscript should either report theme prevalence (e.g., number or percentage of participants mentioning each strategy/challenge) or replace these quantifiers with hedged formulations such as 'in the reported experiences of participants.'
- [III-C / V-D] The paper states that data saturation was reached, but saturation was reached within one organization. Within-case saturation does not establish saturation across companies, sectors, or regulatory contexts, so it does not strengthen the external validity of the 'industry' claims. Section III-C should clarify that saturation was reached within the studied case only, and Section V-D should acknowledge that the transferability argument rests on analytic generalization rather than on saturation across contexts.
minor comments (6)
- [III-A / References [24] and [25]] Reference [25] has the same title as [24] but appears intended to refer to the smart lipstick project mentioned in Project D; please correct the title and URL so each reference matches its citation.
- [Author affiliation] The email address for Cleyton Magalhaes appears to contain a typo ('cleyton.vanut.ufrpe.br'); please verify and correct it.
- [VII] The Figshare data-availability URL is broken across a line break in the text; provide a single, clickable URL.
- [III-C] The phrase 'The interviews about testing, held between June 1 and July 5, 2024' is awkward; consider 'The interviews, held between June 1 and July 5, 2024, lasted between 15 and 25 minutes.'
- [V-A] In Section V-A, 'could provide valuable results into best practices' should be 'could provide valuable insights into best practices.'
- [IV (Finding headings)] The headings 'Finding 1 – Summary' and 'Findings 2 – Summary' and 'Findings 3 – Summary' are inconsistent in number; use 'Finding 2' and 'Finding 3' for consistency.
Circularity Check
No circularity: interview-based findings are self-contained; self-citations are background/methodological and not load-bearing.
full rationale
This paper is an empirical qualitative case study, not a derivation chain. The central findings—three fairness-testing targets (Finding 1), four ad-hoc strategies (Finding 2), and five challenges (Finding 3)—are induced from 22 practitioner interviews and are supported by direct participant quotations in Tables IV, V, and VI. No parameter is fitted and then renamed as a prediction; no result is imported from prior work to force a conclusion; and no definition is constructed in terms of the outcome it is said to explain. The self-citations that appear ([8], [12], [30]) support background framing and the choice of GPT-assisted coding, but the empirical content of the findings comes from the interviews, not from those citations. The acknowledged limitation in Section V-D that the findings are 'not intended for statistical generalization' and are best seen as 'transferable to similar settings' is a genuine external-validity caveat about the single-company sample versus the industry-level wording; that is a sampling and generalization concern, not a circularity concern. The GPT prompt and the interview guide are data-extraction instruments whose aspects (what is tested, how it is tested, challenges) mirror the eventual thematic organization, but the conclusions are grounded in specific participant statements rather than being equivalent to the prompt or guide by construction. Accordingly, no circular step meets the evidentiary standard of quoting a specific reduction to the paper's own inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption Self-reported interview accounts reflect actual fairness testing practices.
- domain assumption A single South American company with 22 participants is representative enough to support statements about 'software professionals' and 'industry' practice.
- domain assumption GPT-4 Omni-assisted open coding, with manual verification of 20% of quotes, produces reliable and complete theme extraction.
Cite this review
Pith. "Pith review of Software Fairness Testing in Practice." pith.science (2026). https://pith.science/paper/5YA7MKMR
@misc{pith2026250617095,
author = {Pith},
title = {Pith review of: Software Fairness Testing in Practice},
year = {2026},
howpublished = {\url{https://pith.science/paper/5YA7MKMR}},
note = {Machine review of arXiv:2506.17095}
}
read the original abstract
Software testing ensures that a system functions correctly, meets specified requirements, and maintains high quality. As artificial intelligence and machine learning (ML) technologies become integral to software systems, testing has evolved to address their unique complexities. A critical advancement in this space is fairness testing, which identifies and mitigates biases in AI applications to promote ethical and equitable outcomes. Despite extensive academic research on fairness testing, including test input generation, test oracle identification, and component testing, practical adoption remains limited. Industry practitioners often lack clear guidelines and effective tools to integrate fairness testing into real-world AI development. This study investigates how software professionals test AI-powered systems for fairness through interviews with 22 practitioners working on AI and ML projects. Our findings highlight a significant gap between theoretical fairness concepts and industry practice. While fairness definitions continue to evolve, they remain difficult for practitioners to interpret and apply. The absence of industry-aligned fairness testing tools further complicates adoption, necessitating research into practical, accessible solutions. Key challenges include data quality and diversity, time constraints, defining effective metrics, and ensuring model interoperability. These insights emphasize the need to bridge academic advancements with actionable strategies and tools, enabling practitioners to systematically address fairness in AI systems.
Figures
Reference graph
Works this paper leans on
-
[1]
A literature review of software test cases and future research,
M. L. Gillenson, X. Zhang, T. F. Stafford, and Y . Shi, “A literature review of software test cases and future research,” in 2018 IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW) . IEEE, 2018, pp. 252–256
work page 2018
-
[2]
A systematic literature review of litera- ture reviews in software testing,
V . Garousi and M. V . M¨antyl¨a, “A systematic literature review of litera- ture reviews in software testing,” Information and Software Technology, vol. 80, pp. 195–216, 2016
work page 2016
-
[3]
The growth of software testing,
D. Gelperin and B. Hetzel, “The growth of software testing,” Commu- nications of the ACM , vol. 31, no. 6, pp. 687–695, 1988
work page 1988
-
[4]
What is ai software testing? and why,
J. Gao, C. Tao, D. Jie, and S. Lu, “What is ai software testing? and why,” in 2019 IEEE International Conference on Service-Oriented System Engineering (SOSE). IEEE, 2019, pp. 27–2709
work page 2019
-
[5]
Testing and quality validation for ai software–perspectives, issues, and practices,
C. Tao, J. Gao, and T. Wang, “Testing and quality validation for ai software–perspectives, issues, and practices,” IEEE Access , vol. 7, pp. 120 164–120 175, 2019
work page 2019
-
[6]
Y . Brun and A. Meliou, “Software fairness,” in Proceedings of the 2018 26th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering , 2018, pp. 754–759
work page 2018
-
[7]
Fairness definitions explained,
S. Verma and J. Rubin, “Fairness definitions explained,” in Proceedings of the international workshop on software fairness , 2018, pp. 1–7
work page 2018
-
[8]
Software fairness debt: Building a research agenda for addressing bias in ai systems,
R. de Souza Santos, F. Fronchetti, S. Freire, and R. Spinola, “Software fairness debt: Building a research agenda for addressing bias in ai systems,” ACM Transactions on Software Engineering and Methodology, 2025
work page 2025
Show all 36 references
-
[9]
Fairness testing: testing software for discrimination,
S. Galhotra, Y . Brun, and A. Meliou, “Fairness testing: testing software for discrimination,” in Proceedings of the 2017 11th Joint meeting on foundations of software engineering , 2017, pp. 498–510
2017
-
[10]
Black box fairness testing of machine learning models,
A. Aggarwal, P. Lohia, S. Nagar, K. Dey, and D. Saha, “Black box fairness testing of machine learning models,” in Proceedings of the 2019 27th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering , 2019, pp. 625–635
2019
-
[11]
Fairness testing: A comprehensive survey and analysis of trends,
Z. Chen, J. M. Zhang, M. Hort, M. Harman, and F. Sarro, “Fairness testing: A comprehensive survey and analysis of trends,” ACM Trans- actions on Software Engineering and Methodology , vol. 33, no. 5, pp. 1–59, 2024
2024
-
[12]
From literature to practice: Exploring fairness testing tools for the software industry adoption,
T. Nguyen, M. T. Baldassarre, L. F. de Lima, and R. de Souza Santos, “From literature to practice: Exploring fairness testing tools for the software industry adoption,” in Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Mea- surem...
2024
-
[13]
Survey of fairness notions,
M. Z. Kwiatkowska, “Survey of fairness notions,” Information and Software Technology, vol. 31, no. 7, pp. 371–386, 1989
1989
-
[14]
Software fairness: An analysis and survey,
E. Soremekun, M. Papadakis, M. Cordy, and Y . L. Traon, “Software fairness: An analysis and survey,” arXiv preprint arXiv:2205.08809 , 2022
2022 arXiv
-
[15]
A survey on bias and fairness in machine learning,
N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM computing surveys (CSUR), vol. 54, no. 6, pp. 1–35, 2021
2021
-
[16]
Testing theories of fair- ness—intentions matter,
A. Falk, E. Fehr, and U. Fischbacher, “Testing theories of fair- ness—intentions matter,” Games and Economic Behavior, vol. 62, no. 1, pp. 287–303, 2008
2008
-
[17]
Amazon just showed us that’unbiased’algorithms can be inadvertently racist,
R. Letzter, “Amazon just showed us that’unbiased’algorithms can be inadvertently racist,” TECH Insider, 2016
2016
-
[18]
Machine bias,
J. Angwin, J. Larson, S. Mattu, and L. Kirchner, “Machine bias,” in Ethics of data and analytics . Auerbach Publications, 2022, pp. 254– 264
2022
-
[19]
Automated directed fairness testing,
S. Udeshi, P. Arora, and S. Chattopadhyay, “Automated directed fairness testing,” in Proceedings of the 33rd ACM/IEEE international conference on automated software engineering , 2018, pp. 98–108
2018
-
[20]
Guidelines for conducting and reporting case study research in software engineering,
P. Runeson and M. H ¨ost, “Guidelines for conducting and reporting case study research in software engineering,” Empirical software engineering, vol. 14, pp. 131–164, 2009
2009
-
[21]
Case study research in software engineering—it is a case, and it is a study, but is it a case study?
C. Wohlin, “Case study research in software engineering—it is a case, and it is a study, but is it a case study?” Information and Software Technology, vol. 133, p. 106514, 2021
2021
-
[22]
Empirical standards for software engineering research,
P. Ralph, N. b. Ali, S. Baltes, D. Bianculli, J. Diaz, Y . Dittrich, N. Ernst, M. Felderer, R. Feldt, A. Filieri et al., “Empirical standards for software engineering research,” arXiv preprint arXiv:2010.03525 , 2020
2010
-
[23]
Selecting em- pirical methods for software engineering research,
S. Easterbrook, J. Singer, M.-A. Storey, and D. Damian, “Selecting em- pirical methods for software engineering research,” Guide to advanced empirical software engineering , pp. 285–311, 2008
2008
-
[24]
translator
A. Mari. (2023) Lenovo and brazilian innovation hub cesar create sign language “translator” for hearing people with ai. [Online]. Available: https://www.forbes.com/sites/angelicamarideoliveira/2023/08/09/lenov o-and-brazilian-innovation-hub-cesar-create-sign-language-translato...
2023
-
[25]
translator
K. Shaikh. (2025) Lenovo and brazilian innovation hub cesar create sign language “translator” for hearing people with ai. [Online]. Available: https://www.prweb.com/releases/the-worlds-first-smart-lipstick-wins-t he-oscars-of-innovation-award-at-sxsw-2025-302399441.html
2025
-
[26]
Sampling in software engineering research: A critical review and guidelines,
S. Baltes and P. Ralph, “Sampling in software engineering research: A critical review and guidelines,” Empirical Software Engineering, vol. 27, no. 4, p. 94, 2022
2022
-
[27]
Charmaz, Constructing grounded theory
K. Charmaz, Constructing grounded theory . sage, 2014
2014
-
[28]
Recommended steps for thematic synthesis in software engineering,
D. S. Cruzes and T. Dyba, “Recommended steps for thematic synthesis in software engineering,” in 2011 international symposium on empirical software engineering and measurement . IEEE, 2011, pp. 275–284
2011
-
[29]
Supporting qualitative analysis with large language models: Combining codebook with gpt-3 for deductive coding,
Z. Xiao, X. Yuan, Q. V . Liao, R. Abdelghani, and P.-Y . Oudeyer, “Supporting qualitative analysis with large language models: Combining codebook with gpt-3 for deductive coding,” in Companion proceedings of the 28th international conference on intelligent user interfaces , 20...
2023
-
[30]
Applications and implications of large language models in qualitative analysis: A new frontier for empirical software engineering,
M. d. M. Lec ¸a, L. Valenc ¸a, R. Santos, and R. d. S. Santos, “Applications and implications of large language models in qualitative analysis: A new frontier for empirical software engineering,” 2nd International Workshop on Methodological Issues with Empirical Studies in Sof...
2024
-
[31]
Can rapid approaches to qualitative analysis deliver timely, valid findings to clinical leaders? a mixed methods study comparing rapid and thematic analysis,
B. Taylor, C. Henshall, S. Kenyon, I. Litchfield, and S. Greenfield, “Can rapid approaches to qualitative analysis deliver timely, valid findings to clinical leaders? a mixed methods study comparing rapid and thematic analysis,” BMJ open, vol. 8, no. 10, p. e019993, 2018
2018
-
[32]
Prompt engineering in medical education,
T. Heston and C. Khun, “Prompt engineering in medical education,” International Medical Education , vol. 2, pp. 198–205, 8 2023
2023
-
[33]
A prompt pattern catalog to enhance prompt engineering with chatgpt,
J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. El- nashar, J. Spencer-Smith, and D. C. Schmidt, “A prompt pattern catalog to enhance prompt engineering with chatgpt,” arXiv preprint arXiv:2302.11382, 2023
2023 arXiv
-
[34]
A systematic survey of prompt engineering in large language models: Techniques and applications,
P. Sahoo, A. K. Singh, S. Saha, V . Jain, S. Mondal, and A. Chadha, “A systematic survey of prompt engineering in large language models: Techniques and applications,” arXiv preprint arXiv:2402.07927 , 2024
2024 arXiv
-
[35]
The diversity crisis of software engineering for artificial intelligence,
B. Adams and F. Khomh, “The diversity crisis of software engineering for artificial intelligence,” IEEE Software, vol. 37, no. 5, pp. 104–108, 2020
2020
-
[36]
” ignorance and prejudice
J. M. Zhang and M. Harman, “” ignorance and prejudice” in software fairness,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 2021, pp. 1436–1447. 11
2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.