Pith. sign in

REVIEW 21 cited by

"Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.09219 v5 pith:AGVVVHYG submitted 2023-10-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords biaseslettersllm-generatedapplicationbiasgenderharmsprofessional
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) have recently emerged as an effective tool to assist individuals in writing various types of content, including professional documents such as recommendation letters. Though bringing convenience, this application also introduces unprecedented fairness concerns. Model-generated reference letters might be directly used by users in professional scenarios. If underlying biases exist in these model-constructed letters, using them without scrutinization could lead to direct societal harms, such as sabotaging application success rates for female applicants. In light of this pressing issue, it is imminent and necessary to comprehensively study fairness issues and associated harms in this real-world use case. In this paper, we critically examine gender biases in LLM-generated reference letters. Drawing inspiration from social science findings, we design evaluation methods to manifest biases through 2 dimensions: (1) biases in language style and (2) biases in lexical content. We further investigate the extent of bias propagation by analyzing the hallucination bias of models, a term that we define to be bias exacerbation in model-hallucinated contents. Through benchmarking evaluation on 2 popular LLMs- ChatGPT and Alpaca, we reveal significant gender biases in LLM-generated recommendation letters. Our findings not only warn against using LLMs for this application without scrutinization, but also illuminate the importance of thoroughly studying hidden biases and harms in LLM-generated professional documents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

    cs.CL 2026-08 conditional novelty 7.0 of 10

    A new tri-lingual Bavarian culture benchmark shows open-weight LLMs underperform on Bavarian and source-grounded items, and that evaluation protocol materially changes accuracy and rankings.

  2. Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Feeding a model's own explanation of its biased story output back into a rewritten prompt improves demographic parity by 2% to 20%.

  3. McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new Chinese bias benchmark with 4,077 instances and five tasks indicates larger language models are less biased than smaller ones when bias is measured through understanding tasks.

  4. From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Dual-character narrative prompts reveal gender biases in six multimodal LLMs that are largely invisible in single-character evaluations, and GENRES provides a structured benchmark to measure them.

  5. Social Group Bias in AI Finance

    econ.GN 2025-06 conditional novelty 6.0 of 10

    Open-source LLMs quote Black mortgage applicants higher interest rates than identical white applicants; a control-vector intervention reduces the gap by about a third on average, but the reduction is measured in-sample.

  6. The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

    cs.AI 2025-06 reject novelty 6.0 of 10

    MedPerturb finds that LLMs are more sensitive to gender and style changes in clinical text, while medical students are more sensitive to LLM-generated summaries and dialogues, in triage decisions.

  7. DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Caste-based stereotypes are measurably present in widely used LLMs, with the largest bias appearing when Dalits and Shudras are compared with dominant castes.

  8. Meflex: A Multi-agent Scaffolding System for Entrepreneurial Ideation Iteration via Nonlinear Business Plan Writing

    cs.HC 2026-02 conditional novelty 5.0 of 10

    A nonlinear, LLM-scaffolded business-plan writing tool with reflection and meta-reflection improves perceived usability and helps students iterate ideas in a 30-participant study.

  9. Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A name-substituted variant of the BBQ benchmark shows that LLMs retain nationality stereotypes even when explicit labels are removed, with smaller models showing more bias and lower accuracy.

  10. A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across U.S. and Chinese survey questions, DeepSeek, GPT-4o, Qwen2.5, and Llama-3.3 all show demographic overgeneralization, with no consistent home-field advantage for the Chinese model.

  11. The Biased Samaritan: LLM biases in Perceived Kindness

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across ten commercial LLMs, demographic groups other than white male middle-aged were rated as more likely to help, while the control 'person' condition aligned with that majority baseline.

  12. A Day in Their Shoes: Using LLM-Based Perspective-Taking Interactive Fiction to Reduce Stigma Toward Dirty Work

    cs.HC 2025-05 conditional novelty 5.0 of 10

    An LLM-driven interactive fiction that puts players in the role of a stigmatized worker raised self-reported understanding and empathy, but stigma reduction itself is not directly measured or controlled.

  13. Unmasking Conversational Bias in AI Multiagent Systems

    cs.CL 2025-01 conditional novelty 5.0 of 10

    In simulated echo-chamber chats, conservative-aligned LLM agents often shift to liberal-aligned messages, a drift that one-shot questionnaire tests do not detect.

  14. Implicit Priors Editing in Stable Diffusion via Targeted Token Adjustment

    cs.CV 2024-12 conditional novelty 5.0 of 10

    EMBEDIT edits a single word token embedding in Stable Diffusion to steer implicit visual priors (e.g., making 'bear' generate 'polar bear'), reporting better accuracy than cross-attention editing while using far fewer...

  15. An Empirical Investigation of Gender Stereotype Representation in Large Language Models: The Italian Case

    cs.CL 2025-07 conditional novelty 4.0 of 10

    In Italian, both ChatGPT (gpt-4o-mini) and Gemini (gemini-1.5-flash) associate higher-status professional roles with male pronouns and subordinate roles with female pronouns, e.g., 97-100% of 'she' responses pointed t...

  16. Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Chinese gender-bias corpus and shared-task benchmark show detection and classification are feasible, while automatic mitigation remains weak.

  17. Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey

    cs.CL 2025-02 conditional novelty 4.0 of 10

    A survey organizes current research on trustworthy RAG into six pillars, reliability, privacy, safety, fairness, explainability, and accountability, and maps methods, metrics, and open problems for each.

  18. Irony in Emojis: A Comparative Study of Human and LLM Interpretation

    cs.CL 2025-01 conditional novelty 4.0 of 10

    GPT-4o systematically overestimates the likelihood that emojis are used ironically, correlating only weakly with human-perceived scores from the Ciron dataset.

  19. Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives

    cs.AI 2025-01 conditional novelty 4.0 of 10

    GPT-4 personas aligned with real patient answers 54.97% on average versus 26.7% random, but only when primed with education, and the reported 88% accuracy overstates what was measured.

  20. On the Surprising Efficacy of LLMs for Penetration-Testing

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A critical review arguing that LLMs are surprisingly effective for penetration testing because the task is largely pattern-matching, while noting serious reliability, safety, and cost barriers to autonomous use.

  21. Improving LLM Group Fairness on Tabular Data via In-Context Learning

    cs.LG 2024-12

Pith tools