Pith. sign in

REVIEW 8 cited by

The Woman Worked as a Babysitter: On Biases in Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.01326 v2 pith:F5HQG4X3 submitted 2019-09-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords regardbiaseslanguagetextanalyzebiasdemographicdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a systematic study of biases in natural language generation (NLG) by analyzing text generated from prompts that contain mentions of different demographic groups. In this work, we introduce the notion of the regard towards a demographic, use the varying levels of regard towards different demographics as a defining metric for bias in NLG, and analyze the extent to which sentiment scores are a relevant proxy metric for regard. To this end, we collect strategically-generated text from language models and manually annotate the text with both sentiment and regard scores. Additionally, we build an automatic regard classifier through transfer learning, so that we can analyze biases in unseen text. Together, these methods reveal the extent of the biased nature of language model generations. Our analysis provides a study of biases in NLG, bias metrics and correlated human judgments, and empirical evidence on the usefulness of our annotated dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 10 citations worldwide. Full citation record

  1. Scaling Truth: The Confidence Paradox in AI Fact-Checking

    cs.SI 2025-09 conditional novelty 6.0 of 10

    Across LLM fact-checking, model scale correlates with an inverse pattern of accuracy and decisiveness: smaller models are overconfident and less accurate, larger models are accurate but overly cautious.

  2. Dutch CrowS-Pairs: Adapting a Challenge Dataset for Measuring Social Biases in Language Models for Dutch

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The paper presents a Dutch adaptation of the CrowS-Pairs bias benchmark and reports bias scores for seven masked and two autoregressive language models across nine demographic categories.

  3. Exploring Gender Bias Beyond Occupational Titles

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The paper presents GenderLexicon and a ClozeGender score, reporting that action verbs and object nouns carry gender bias beyond occupational stereotypes in English and Japanese language models.

  4. Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)

    cs.CY 2025-05 reject novelty 6.0 of 10

    Placing demographic audience information in system prompts rather than user prompts shifts sentiment and ranking outputs across six commercial LLMs, but the design confounds position with instruction content.

  5. Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

    cs.CY 2026-06 conditional novelty 5.0 of 10

    Across 25,000 stories from five LLMs, an LLM judge rated stories mentioning intellectual disabilities as more infantile, paternalistic, dependent, and inspirational than stories without the label.

  6. Inference Time Debiasing Concepts in Diffusion Models

    cs.GR 2025-08 reject novelty 5.0 of 10

    DeCoDi subtracts a biased-concept guidance term during diffusion inference to shift generated images away from targeted stereotypes, with evaluation on gender, ethnicity, and age.

  7. Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    The abstract claims LLMs show up to 40% coreference confidence disparities across intersectional identities, but the article body is an unrelated paper on robotic fruit handling.

  8. LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation

    cs.CL 2026-07 conditional novelty 3.0 of 10

    On the LT-EDI 2026 shared task, LoRA fine-tuning scored 80.00% for gender-inclusive rewriting, while PCA-based activation steering of Gemma-3-4B-it scored 78.12% for counter-narrative generation, with a manual analysi...

Pith tools