Pith. sign in

REVIEW 2 cited by

Stereotype and Skew: Quantifying Gender Bias in Pre-trained and Fine-tuned Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.09688 v2 pith:HMAWLT64 submitted 2021-01-24 cs.CL cs.AIcs.LGcs.NE

classification cs.CLcs.AIcs.LGcs.NE
keywords biasgenderskewstereotypemodelsfindfine-tunedlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes two intuitive metrics, skew and stereotype, that quantify and analyse the gender bias present in contextual language models when tackling the WinoBias pronoun resolution task. We find evidence that gender stereotype correlates approximately negatively with gender skew in out-of-the-box models, suggesting that there is a trade-off between these two forms of bias. We investigate two methods to mitigate bias. The first approach is an online method which is effective at removing skew at the expense of stereotype. The second, inspired by previous work on ELMo, involves the fine-tuning of BERT using an augmented gender-balanced dataset. We show that this reduces both skew and stereotype relative to its unaugmented fine-tuned counterpart. However, we find that existing gender bias benchmarks do not fully probe professional bias as pronoun resolution may be obfuscated by cross-correlations from other manifestations of gender prejudice. Our code is available online, at https://github.com/12kleingordon34/NLP_masters_project.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Language models show negligible persona-based differences on MMLU benchmarks but large, income-relevant differences when asked for salary negotiation advice.

  2. Mitigating Confounding in Speech-Based Dementia Detection through Weight Masking

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Masking weights that react to gender in a fine-tuned BERT reduces gender gaps in dementia predictions while keeping most of the detection accuracy.

Pith tools