REVIEW 3 major objections 2 minor 7 references
Can LLMs Hire Fairly? Racial Bias in Resume Screening
T0 review · 3 major / 2 minor · reviewed 2026-06-30 · grok-4.3
Pith's one-line read Only the 2023 LLM shows pro-White hiring bias matching human experiments, while 2024 and later models show null or pro-Black gaps.
desk verdict The paper shows a reversal in LLM resume bias from pro-White in the 2023 model to null or pro-Black in 2024+ ones via paired resumes, but the direct link to real hiring discrimination remains unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Paired-resume methodology applied to LLM outputs, comparing callback rates on resumes identical except for names that signal race or gender.
What would settle it
Repeating the exact paired-resume tests on the same 14 models but with a fresh set of job postings or name variants and finding no generational difference in the direction of bias.
Extended reading notes
Core claim
The paper establishes that the direction of bias in LLM resume screening reversed across model generations. The 2023 model reproduces the pro-White callback gap of +2.12 pp documented in field experiments on labor market discrimination, significant at the 1% level. Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal of up to -3.01 pp. The same pattern holds on the gender axis. These results are based on 24,024 paired postings per model across the 14 models tested.
Load-bearing premise
The paired resumes that differ only in the racial or gender signal from names produce a valid measure of the LLM's hiring discrimination in the same way they measure human discrimination.
Editorial extensions
If this is right
- Newer LLMs could produce different or reversed demographic hiring outcomes when used for automated screening.
- The bias reversal coincides with model release year, indicating that training or alignment changes alter how demographic signals are processed.
- The generational shift appears on both the racial and gender axes.
- Large-scale testing across 24,024 pairs per model supports the claim of a consistent reversal after 2023.
Reading between the lines
- If the reversal stems from post-2023 alignment techniques, those techniques may have over-corrected on demographic cues.
- The pattern suggests that bias in LLM decision systems may require repeated audits with each new model release rather than a one-time check.
- The findings could connect to broader questions about how training data updates affect fairness in other high-stakes LLM applications.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper audits 14 mainstream LLMs for hiring discrimination on race and gender using the paired-resume callback methodology of Kline, Rose, and Walters (2022). It reports that the single 2023-vintage model reproduces the pro-White callback gap observed in human field experiments (+2.12 pp, p<0.01), while all 2024+ models exhibit either a null gap or a significant pro-Black reversal (up to -3.01 pp); an analogous generational reversal appears on the gender axis. Results rest on 24,024 paired evaluations per model.
Significance. If the mapping from LLM output differences to the field-experiment callback gap is valid, the documented reversal would indicate that post-2023 alignment and safety training have materially altered the direction of demographic bias in LLM-based resume screening. This would be a substantive empirical contribution to the literature on algorithmic fairness and the effects of RLHF-style training, with direct implications for deployment of LLMs in employment contexts.
major comments (3)
- [Methods] Methods section: the manuscript provides no prompt templates, system instructions, temperature settings, or refusal-handling protocol. Because the measured callback gap is extracted from model text completions, the absence of these details prevents evaluation of whether the reported gaps are robust to prompt variation or are artifacts of particular phrasings.
- [Results / Discussion] §3 (or equivalent results section) and the comparison to Kline et al. (2022): the central claim equates LLM output differences on name-swapped resumes with the employer callback gap. No ablation, human-rater calibration, or discussion of cost/legal constraints is supplied to support this equivalence; the generational reversal result therefore rests on an untested assumption that the two measures capture the same construct.
- [Results] Table 1 (or model-by-model results table): effect sizes and significance levels are reported, yet the paper does not report per-model variance across resume templates, occupation categories, or name sets. Without these breakdowns it is impossible to assess whether the reported reversal is driven by a subset of conditions or holds uniformly.
minor comments (2)
- [Abstract / Introduction] Abstract and §1: the phrase “algorithmic hiring bias” is used without distinguishing between bias in the LLM output distribution and bias in downstream hiring decisions; a brief clarifying sentence would improve precision.
- [Methods] The total of 24,024 paired postings per model is stated but the exact breakdown (number of occupations, name pairs, resume templates) is not summarized in a table; adding such a table would aid reproducibility.
Simulated Author's Rebuttal
We thank the referee for their constructive comments, which highlight important aspects of transparency and robustness. We address each major comment below and indicate planned revisions.
read point-by-point responses
-
Referee: [Methods] Methods section: the manuscript provides no prompt templates, system instructions, temperature settings, or refusal-handling protocol. Because the measured callback gap is extracted from model text completions, the absence of these details prevents evaluation of whether the reported gaps are robust to prompt variation or are artifacts of particular phrasings.
Authors: We agree that these details are essential for reproducibility. In the revised manuscript we will add a new subsection to the Methods that provides the complete prompt templates, system instructions, temperature settings (set to 0 wherever supported for determinism), and the exact protocol used to extract callback decisions and handle any refusals or non-callback outputs. revision: yes
-
Referee: [Results / Discussion] §3 (or equivalent results section) and the comparison to Kline et al. (2022): the central claim equates LLM output differences on name-swapped resumes with the employer callback gap. No ablation, human-rater calibration, or discussion of cost/legal constraints is supplied to support this equivalence; the generational reversal result therefore rests on an untested assumption that the two measures capture the same construct.
Authors: The paper applies the identical paired-resume callback definition and name-swapping procedure as Kline, Rose, and Walters (2022), so the measured gap is the same operational quantity. The contribution is to show how this quantity changes across LLM generations. We will expand the discussion to clarify the mapping, note that LLM outputs are not employer decisions, and address limitations including cost and legal constraints on deployment; however, new human-rater calibration or ablations are outside the scope of the current study design. revision: partial
-
Referee: [Results] Table 1 (or model-by-model results table): effect sizes and significance levels are reported, yet the paper does not report per-model variance across resume templates, occupation categories, or name sets. Without these breakdowns it is impossible to assess whether the reported reversal is driven by a subset of conditions or holds uniformly.
Authors: We agree that disaggregated results would strengthen the presentation. The revised manuscript will include supplementary tables reporting callback gaps broken down by resume template, occupation category, and name set for each model, confirming that the generational reversal pattern is not confined to particular subsets. revision: yes
Circularity Check
No circularity: results are direct empirical counts from fixed LLM prompts on resume pairs
full rationale
The paper applies the external Kline, Rose & Walters (2022) paired-resume audit to LLM outputs and reports raw differences in model recommendations across name signals. No equations, fitted parameters, self-citations, or ansatzes are used to derive the generational reversal result; the central findings are literal tallies of model responses to fixed inputs. The cited methodology is from independent authors and is treated as a measurement protocol rather than a self-referential definition. No load-bearing step reduces the reported gaps (+2.12 pp or -3.01 pp) to anything other than the observed LLM outputs.
Assumptions & free parameters
assumptions (2)
- domain assumption Paired resumes differ solely in the racial or gender signal as constructed in the 2022 methodology
- standard math Statistical significance testing assumptions hold for the callback rate differences
Cite this review
Pith. "Pith review of Can LLMs Hire Fairly? Racial Bias in Resume Screening." pith.science (2026). https://pith.science/paper/AUT65GFF
@misc{pith2026260628978,
author = {Pith},
title = {Pith review of: Can LLMs Hire Fairly? Racial Bias in Resume Screening},
year = {2026},
howpublished = {\url{https://pith.science/paper/AUT65GFF}},
note = {Machine review of arXiv:2606.28978}
}
abstract
We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback gap documented in field experiments on labor market discrimination ($+2.12$ pp, significant at the 1\% level). Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal (up to $-3.01$ pp). The same pattern holds on the gender axis. Based on 24,024 paired postings per model across 14 models, our results document a reversal in the direction of algorithmic hiring bias across model generations.
Reference graph
Works this paper leans on
-
[1]
American Economic Review , Volume =
Bertrand, Marianne and Mullainathan, Sendhil , Title =. American Economic Review , Volume =. 2004 , Month =
work page 2004
-
[2]
The Quarterly Journal of Economics , volume=
Systemic discrimination among large US employers , author=. The Quarterly Journal of Economics , volume=. 2022 , publisher=
work page 2022
- [3]
-
[4]
Blodgett, Su Lin and Barocas, Solon and Daum \'e III, Hal and Wallach, Hanna. Language (Technology) is Power: A Critical Survey of ``Bias'' in NLP. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.485
-
[5]
Computational Linguistics , author =
Gallegos, Isabel O. and Rossi, Ryan A. and Barrow, Joe and Tanjim, Md Mehrab and Kim, Sungchul and Dernoncourt, Franck and Yu, Tong and Zhang, Ruiyi and Ahmed, Nesreen K. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics. 2024. doi:10.1162/coli_a_00524
-
[6]
The Journal of Finance , volume=
Predictably unequal? The effects of machine learning on credit markets , author=. The Journal of Finance , volume=. 2022 , publisher=
work page 2022
-
[7]
Veldanda, Akshaj Kumar and Grob, Fabian and Thakur, Shailja and Pearce, Hammond and Tan, Benjamin and Karri, Ramesh and Garg, Siddharth , journal=. Are
Reviewed June 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.