LLMs encode stereotypes along recoverable geometric axes in attention heads, and two tested LLMs share more stereotype content with each other than with documented human stereotypes.
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , year =
9 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
unclear 1representative citing papers
Experiments show prompt injection boosts resume rankings in LLM screening when rare and qualities homogeneous, but fails when common or qualities differ.
A new benchmark shows LLMs are more accurate in Simplified Chinese for regional terms but favor Taiwanese names in simulated hiring, revealing task-dependent bias between Chinese script variants.
A systematized narrative review synthesizing 40 works to map the evolution of recruitment AI from bilateral matching to tool-using agents, with a staged framework linking automation scope to evaluation unit and claim ceiling.
LLMs withhold and soften negative judgments more when addressing users directly, and user affective context, especially loneliness and distress, amplifies this sycophantic divergence.
Audit format (rate vs rank vs allocate; transparent vs disguised) reverses the apparent direction of LLM demographic bias, while causal framing of need dominates allocations by roughly an order of magnitude.
Comparative words in prompts can shift LLM answers toward the framed direction in simple arithmetic comparisons, with demographic terms amplifying the effect.
Language models show negligible persona-based differences on MMLU benchmarks but large, income-relevant differences when asked for salary negotiation advice.
citing papers explorer
-
STEREODISCO: Discovering Stereotypicality in LLMs
LLMs encode stereotypes along recoverable geometric axes in attention heads, and two tested LLMs share more stereotype content with each other than with documented human stereotypes.
-
Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and Multi-Injection Settings
Experiments show prompt injection boosts resume rankings in LLM screening when rare and qualities homogeneous, but fails when common or qualities differ.
-
Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese
A new benchmark shows LLMs are more accurate in Simplified Chinese for regional terms but favor Taiwanese names in simulated hiring, revealing task-dependent bias between Chinese script variants.
-
From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
A systematized narrative review synthesizing 40 works to map the evolution of recruitment AI from bilateral matching to tool-using agents, with a staged framework linking automation scope to evaluation unit and claim ceiling.
-
Affective Context Amplifies Sycophancy in LLM Responses
LLMs withhold and soften negative judgments more when addressing users directly, and user affective context, especially loneliness and distress, amplifies this sycophantic divergence.
-
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
Audit format (rate vs rank vs allocate; transparent vs disguised) reverses the apparent direction of LLM demographic bias, while causal framing of need dominates allocations by roughly an order of magnitude.
-
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
Comparative words in prompts can shift LLM answers toward the framed direction in simple arithmetic comparisons, with demographic terms amplifying the effect.
-
Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
Language models show negligible persona-based differences on MMLU benchmarks but large, income-relevant differences when asked for salary negotiation advice.
- Topics as Proxies for Sociodemographics: How Conversational Context Affects LLM Answers