Pith. sign in

REVIEW 3 cited by

Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05135 v1 pith:TTIXEZYV submitted 2023-10-08 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords biasllmsresumesstatusgenderhiringraceacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) such as GPT-3.5, Bard, and Claude exhibit applicability across numerous tasks. One domain of interest is their use in algorithmic hiring, specifically in matching resumes with job categories. Yet, this introduces issues of bias on protected attributes like gender, race and maternity status. The seminal work of Bertrand & Mullainathan (2003) set the gold-standard for identifying hiring bias via field experiments where the response rate for identical resumes that differ only in protected attributes, e.g., racially suggestive names such as Emily or Lakisha, is compared. We replicate this experiment on state-of-art LLMs (GPT-3.5, Bard, Claude and Llama) to evaluate bias (or lack thereof) on gender, race, maternity status, pregnancy status, and political affiliation. We evaluate LLMs on two tasks: (1) matching resumes to job categories; and (2) summarizing resumes with employment relevant information. Overall, LLMs are robust across race and gender. They differ in their performance on pregnancy status and political affiliation. We use contrastive input decoding on open-source LLMs to uncover potential sources of bias.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights

    cs.CY 2025-08 conditional novelty 6.0 of 10

    LLMs that screen resumes systematically prefer their own generated summaries over human-written ones, with simulated shortlisting advantages of 23 to 60 percent for same-model users.

  2. Understanding Gender Bias in AI-Generated Product Descriptions

    cs.CL 2025-06 conditional novelty 6.0 of 10

    AI-generated product descriptions on eBay show systematic gender bias, including body-size exclusions, stereotyped feature emphasis, and differences in calls to action.

  3. The Biased Samaritan: LLM biases in Perceived Kindness

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across ten commercial LLMs, demographic groups other than white male middle-aged were rated as more likely to help, while the control 'person' condition aligned with that majority baseline.

Pith tools