Pith. sign in

REVIEW 4 cited by

A Survey on Gender Bias in Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.14168 v1 pith:RSBUL3H7 submitted 2021-12-28 cs.CL cs.CY

classification cs.CLcs.CY
keywords genderbiasresearchdefinitionslanguagesurveybeendeveloped
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language can be used as a means of reproducing and enforcing harmful stereotypes and biases and has been analysed as such in numerous research. In this paper, we present a survey of 304 papers on gender bias in natural language processing. We analyse definitions of gender and its categories within social sciences and connect them to formal definitions of gender bias in NLP research. We survey lexica and datasets applied in research on gender bias and then compare and contrast approaches to detecting and mitigating gender bias. We find that research on gender bias suffers from four core limitations. 1) Most research treats gender as a binary variable neglecting its fluidity and continuity. 2) Most of the work has been conducted in monolingual setups for English or other high-resource languages. 3) Despite a myriad of papers on gender bias in NLP methods, we find that most of the newly developed algorithms do not test their models for bias and disregard possible ethical considerations of their work. 4) Finally, methodologies developed in this line of research are fundamentally flawed covering very limited definitions of gender bias and lacking evaluation baselines and pipelines. We suggest recommendations towards overcoming these limitations as a guide for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Gender Inclusivity Fairness Index (GIFI): A Multilevel Framework for Evaluating Gender Diversity in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GIFI is a seven-part fairness index showing that LLMs handle 'he' and 'she' far better than neutral and neopronouns, with GPT-4o ranking highest among 22 models.

  2. GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns

    cs.CL 2025-05 conditional novelty 6.0 of 10

    GeNRe is the first French gender-neutral rewriting system to replace masculine plural member nouns with collective nouns, reaching 3.81% WER with its rule-based version.

  3. Integrating gender inclusivity into large language models via instruction tuning

    cs.CL 2025-08 reject novelty 5.0 of 10

    Abstract promises gender-inclusive Polish LLM tuning with the IPIS dataset, while the full text is an unrelated quantum transformer paper; no evidence for the declared claims is present.

  4. A Case Study of Balanced Query Recommendation on Wikipedia

    cs.IR 2025-08 conditional novelty 4.0 of 10

    BalancedQR, extended to handle multiple bias dimensions with a Pareto front, recommends less biased Wikipedia queries, and a GloVe-plus-LLM candidate generation method dominates alternatives.

Pith tools