Pith. sign in

REVIEW 5 cited by

Auditing the Use of Language Models to Guide Hiring Decisions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03086 v1 pith:BJLLEXRP submitted 2024-04-03 stat.AP cs.CL

classification stat.APcs.CL
keywords auditingmodelscorrespondenceexperimentsllmsalgorithmicalgorithmsapplication
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Regulatory efforts to protect against algorithmic bias have taken on increased urgency with rapid advances in large language models (LLMs), which are machine learning models that can achieve performance rivaling human experts on a wide array of tasks. A key theme of these initiatives is algorithmic "auditing," but current regulations -- as well as the scientific literature -- provide little guidance on how to conduct these assessments. Here we propose and investigate one approach for auditing algorithms: correspondence experiments, a widely applied tool for detecting bias in human judgements. In the employment context, correspondence experiments aim to measure the extent to which race and gender impact decisions by experimentally manipulating elements of submitted application materials that suggest an applicant's demographic traits, such as their listed name. We apply this method to audit candidate assessments produced by several state-of-the-art LLMs, using a novel corpus of applications to K-12 teaching positions in a large public school district. We find evidence of moderate race and gender disparities, a pattern largely robust to varying the types of application material input to the models, as well as the framing of the task to the LLMs. We conclude by discussing some important limitations of correspondence experiments for auditing algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Chain Is Only as Strong as Its Weakest Link: A Scoping Review of System Integration Audits in AI

    cs.SE 2026-08 conditional novelty 6.0 of 10

    AI auditing that targets system integration is emerging but fragmented, and can be categorized into inter-component, system-environment, and multi-system sites serving four audit functions.

  2. Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

    cs.CY 2025-07 conditional novelty 6.0 of 10

    Models that appear demographically neutral in AI resume screening can actually be incompetent evaluators, a pattern the paper calls the Illusion of Neutrality.

  3. Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models

    cs.AI 2025-06 unverdicted novelty 6.0 of 10

    LLMs show a quality-dependent position bias, favoring the first option for high-quality choices and later options for low-quality ones, and higher-temperature sampling can reveal the underlying preference.

  4. Correlated Errors in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Large language models from different providers and architectures often make the same errors, and more accurate models are especially likely to share mistakes.

  5. Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions

    cs.LG 2025-07 reject novelty 5.0 of 10

    A benchmark of LLMs versus a proprietary hiring model on ~10,000 real candidate-job pairs reports the proprietary model wins on accuracy and fairness, while all tested LLMs show racial and intersectional bias.

Pith tools