Pith. sign in

REVIEW 3 cited by

"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12232 v2 pith:ZUAQCQM2 submitted 2024-06-18 cs.AI cs.CL

classification cs.AIcs.CL
keywords candidatesmodelsnamesrecommendationsacrossemploymentgenderhiring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Social science research has shown that candidates with names indicative of certain races or genders often face discrimination in employment practices. Similarly, Large Language Models (LLMs) have demonstrated racial and gender biases in various applications. In this study, we utilize GPT-3.5-Turbo and Llama 3-70B-Instruct to simulate hiring decisions and salary recommendations for candidates with 320 first names that strongly signal their race and gender, across over 750,000 prompts. Our empirical results indicate a preference among these models for hiring candidates with White female-sounding names over other demographic groups across 40 occupations. Additionally, even among candidates with identical qualifications, salary recommendations vary by as much as 5% between different subgroups. A comparison with real-world labor data reveals inconsistent alignment with U.S. labor market characteristics, underscoring the necessity of risk investigation of LLM-powered systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.

  2. AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights

    cs.CY 2025-08 conditional novelty 6.0 of 10

    LLMs that screen resumes systematically prefer their own generated summaries over human-written ones, with simulated shortlisting advantages of 23 to 60 percent for same-model users.

  3. Resume-ing Control: (Mis)Perceptions of Agency Around GenAI Use in Recruiting Workflows

    cs.CY 2026-04 unverdicted novelty 5.0 of 10

    Recruiters perceive themselves as retaining agency over GenAI in hiring pipelines, yet GenAI invisibly architects core evaluation inputs, producing only marginal efficiency gains at the cost of deskilling.

Pith tools