Pith. sign in

REVIEW 3 major objections 2 minor 7 references

Can LLMs Hire Fairly? Racial Bias in Resume Screening

T0 review · 3 major / 2 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read Only the 2023 LLM shows pro-White hiring bias matching human experiments, while 2024 and later models show null or pro-Black gaps.

desk verdict The paper shows a reversal in LLM resume bias from pro-White in the 2023 model to null or pro-Black in 2024+ ones via paired resumes, but the direct link to real hiring discrimination remains unproven. read the letter →

arxiv 2606.28978 v1 pith:AUT65GFF submitted 2026-06-27 cs.CL cs.CY

classification cs.CLcs.CY
keywords largelanguagemodelshiringdiscriminationracialbiasgenderresumescreeningalgorithmicpairedaudit
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper audits fourteen mainstream LLMs for hiring discrimination by applying the paired-resume method, feeding each model thousands of resume pairs that differ only in names signaling race or gender. The single 2023-vintage model produces a statistically significant pro-White callback gap of 2.12 percentage points, replicating the direction and rough size of gaps found in real-world employer experiments. Every model released in 2024 or after instead shows either no detectable gap or a significant pro-Black reversal reaching 3.01 percentage points, with the identical generational pattern appearing on the gender axis. The tests cover 24,024 paired evaluations per model. A sympathetic reader would care because LLMs are increasingly used to screen job applications, so any shift in their bias direction would directly affect who receives callbacks.

What carries the argument

Paired-resume methodology applied to LLM outputs, comparing callback rates on resumes identical except for names that signal race or gender.

What would settle it

Repeating the exact paired-resume tests on the same 14 models but with a fresh set of job postings or name variants and finding no generational difference in the direction of bias.

Watch

Extended reading notes

Core claim

The paper establishes that the direction of bias in LLM resume screening reversed across model generations. The 2023 model reproduces the pro-White callback gap of +2.12 pp documented in field experiments on labor market discrimination, significant at the 1% level. Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal of up to -3.01 pp. The same pattern holds on the gender axis. These results are based on 24,024 paired postings per model across the 14 models tested.

Load-bearing premise

The paired resumes that differ only in the racial or gender signal from names produce a valid measure of the LLM's hiring discrimination in the same way they measure human discrimination.

Editorial extensions

If this is right

  • Newer LLMs could produce different or reversed demographic hiring outcomes when used for automated screening.
  • The bias reversal coincides with model release year, indicating that training or alignment changes alter how demographic signals are processed.
  • The generational shift appears on both the racial and gender axes.
  • Large-scale testing across 24,024 pairs per model supports the claim of a consistent reversal after 2023.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the reversal stems from post-2023 alignment techniques, those techniques may have over-corrected on demographic cues.
  • The pattern suggests that bias in LLM decision systems may require repeated audits with each new model release rather than a one-time check.
  • The findings could connect to broader questions about how training data updates affect fairness in other high-stakes LLM applications.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper audits 14 mainstream LLMs for hiring discrimination on race and gender using the paired-resume callback methodology of Kline, Rose, and Walters (2022). It reports that the single 2023-vintage model reproduces the pro-White callback gap observed in human field experiments (+2.12 pp, p<0.01), while all 2024+ models exhibit either a null gap or a significant pro-Black reversal (up to -3.01 pp); an analogous generational reversal appears on the gender axis. Results rest on 24,024 paired evaluations per model.

Significance. If the mapping from LLM output differences to the field-experiment callback gap is valid, the documented reversal would indicate that post-2023 alignment and safety training have materially altered the direction of demographic bias in LLM-based resume screening. This would be a substantive empirical contribution to the literature on algorithmic fairness and the effects of RLHF-style training, with direct implications for deployment of LLMs in employment contexts.

major comments (3)
  1. [Methods] Methods section: the manuscript provides no prompt templates, system instructions, temperature settings, or refusal-handling protocol. Because the measured callback gap is extracted from model text completions, the absence of these details prevents evaluation of whether the reported gaps are robust to prompt variation or are artifacts of particular phrasings.
  2. [Results / Discussion] §3 (or equivalent results section) and the comparison to Kline et al. (2022): the central claim equates LLM output differences on name-swapped resumes with the employer callback gap. No ablation, human-rater calibration, or discussion of cost/legal constraints is supplied to support this equivalence; the generational reversal result therefore rests on an untested assumption that the two measures capture the same construct.
  3. [Results] Table 1 (or model-by-model results table): effect sizes and significance levels are reported, yet the paper does not report per-model variance across resume templates, occupation categories, or name sets. Without these breakdowns it is impossible to assess whether the reported reversal is driven by a subset of conditions or holds uniformly.
minor comments (2)
  1. [Abstract / Introduction] Abstract and §1: the phrase “algorithmic hiring bias” is used without distinguishing between bias in the LLM output distribution and bias in downstream hiring decisions; a brief clarifying sentence would improve precision.
  2. [Methods] The total of 24,024 paired postings per model is stated but the exact breakdown (number of occupations, name pairs, resume templates) is not summarized in a table; adding such a table would aid reproducibility.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their constructive comments, which highlight important aspects of transparency and robustness. We address each major comment below and indicate planned revisions.

read point-by-point responses
  1. Referee: [Methods] Methods section: the manuscript provides no prompt templates, system instructions, temperature settings, or refusal-handling protocol. Because the measured callback gap is extracted from model text completions, the absence of these details prevents evaluation of whether the reported gaps are robust to prompt variation or are artifacts of particular phrasings.

    Authors: We agree that these details are essential for reproducibility. In the revised manuscript we will add a new subsection to the Methods that provides the complete prompt templates, system instructions, temperature settings (set to 0 wherever supported for determinism), and the exact protocol used to extract callback decisions and handle any refusals or non-callback outputs. revision: yes

  2. Referee: [Results / Discussion] §3 (or equivalent results section) and the comparison to Kline et al. (2022): the central claim equates LLM output differences on name-swapped resumes with the employer callback gap. No ablation, human-rater calibration, or discussion of cost/legal constraints is supplied to support this equivalence; the generational reversal result therefore rests on an untested assumption that the two measures capture the same construct.

    Authors: The paper applies the identical paired-resume callback definition and name-swapping procedure as Kline, Rose, and Walters (2022), so the measured gap is the same operational quantity. The contribution is to show how this quantity changes across LLM generations. We will expand the discussion to clarify the mapping, note that LLM outputs are not employer decisions, and address limitations including cost and legal constraints on deployment; however, new human-rater calibration or ablations are outside the scope of the current study design. revision: partial

  3. Referee: [Results] Table 1 (or model-by-model results table): effect sizes and significance levels are reported, yet the paper does not report per-model variance across resume templates, occupation categories, or name sets. Without these breakdowns it is impossible to assess whether the reported reversal is driven by a subset of conditions or holds uniformly.

    Authors: We agree that disaggregated results would strengthen the presentation. The revised manuscript will include supplementary tables reporting callback gaps broken down by resume template, occupation category, and name set for each model, confirming that the generational reversal pattern is not confined to particular subsets. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: results are direct empirical counts from fixed LLM prompts on resume pairs

full rationale

The paper applies the external Kline, Rose & Walters (2022) paired-resume audit to LLM outputs and reports raw differences in model recommendations across name signals. No equations, fitted parameters, self-citations, or ansatzes are used to derive the generational reversal result; the central findings are literal tallies of model responses to fixed inputs. The cited methodology is from independent authors and is treated as a measurement protocol rather than a self-referential definition. No load-bearing step reduces the reported gaps (+2.12 pp or -3.01 pp) to anything other than the observed LLM outputs.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on direct application of an existing paired-comparison method to LLM outputs with no free parameters fitted, no new entities postulated, and only standard statistical assumptions.

assumptions (2)
  • domain assumption Paired resumes differ solely in the racial or gender signal as constructed in the 2022 methodology
    Invoked by direct use of the Kline et al. paired-resume design on LLM inputs.
  • standard math Statistical significance testing assumptions hold for the callback rate differences
    Used to report 1% level significance on the percentage point gaps.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can LLMs Hire Fairly? Racial Bias in Resume Screening." pith.science (2026). https://pith.science/paper/AUT65GFF

@misc{pith2026260628978,
  author       = {Pith},
  title        = {Pith review of: Can LLMs Hire Fairly? Racial Bias in Resume Screening},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AUT65GFF}},
  note         = {Machine review of arXiv:2606.28978}
}
abstract

We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback gap documented in field experiments on labor market discrimination ($+2.12$ pp, significant at the 1\% level). Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal (up to $-3.01$ pp). The same pattern holds on the gender axis. Based on 24,024 paired postings per model across 14 models, our results document a reversal in the direction of algorithmic hiring bias across model generations.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

7 extracted references · 7 canonical work pages

  1. [1]

    American Economic Review , Volume =

    Bertrand, Marianne and Mullainathan, Sendhil , Title =. American Economic Review , Volume =. 2004 , Month =

  2. [2]

    The Quarterly Journal of Economics , volume=

    Systemic discrimination among large US employers , author=. The Quarterly Journal of Economics , volume=. 2022 , publisher=

  3. [3]

    2026 , month = jun, day =

    Daniel Wiessner , title =. 2026 , month = jun, day =

  4. [4]

    Language (

    Blodgett, Su Lin and Barocas, Solon and Daum \'e III, Hal and Wallach, Hanna. Language (Technology) is Power: A Critical Survey of ``Bias'' in NLP. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.485

  5. [5]

    Computational Linguistics , author =

    Gallegos, Isabel O. and Rossi, Ryan A. and Barrow, Joe and Tanjim, Md Mehrab and Kim, Sungchul and Dernoncourt, Franck and Yu, Tong and Zhang, Ruiyi and Ahmed, Nesreen K. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics. 2024. doi:10.1162/coli_a_00524

  6. [6]

    The Journal of Finance , volume=

    Predictably unequal? The effects of machine learning on credit markets , author=. The Journal of Finance , volume=. 2022 , publisher=

  7. [7]

    Veldanda, Akshaj Kumar and Grob, Fabian and Thakur, Shailja and Pearce, Hammond and Tan, Benjamin and Karri, Ramesh and Garg, Siddharth , journal=. Are

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.