Pith. sign in

REVIEW 1 cited by

Language Models Get a Gender Makeover: Mitigating Gender Bias with Few-Shot Data Interventions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04597 v1 pith:JBIUKVUK submitted 2023-06-07 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelsgenderpre-traineddebiasinglanguagetrainingbeenbias
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Societal biases present in pre-trained large language models are a critical issue as these models have been shown to propagate biases in countless downstream applications, rendering them unfair towards specific groups of people. Since large-scale retraining of these models from scratch is both time and compute-expensive, a variety of approaches have been previously proposed that de-bias a pre-trained model. While the majority of current state-of-the-art debiasing methods focus on changes to the training regime, in this paper, we propose data intervention strategies as a powerful yet simple technique to reduce gender bias in pre-trained models. Specifically, we empirically show that by fine-tuning a pre-trained model on only 10 de-biased (intervened) training examples, the tendency to favor any gender is significantly reduced. Since our proposed method only needs a few training examples, our few-shot debiasing approach is highly feasible and practical. Through extensive experimentation, we show that our debiasing technique performs better than competitive state-of-the-art baselines with minimal loss in language modeling ability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EQPO: Equitable Group Relative Policy Optimization for Clinical Reasoning

    cs.LG 2025-10 reject novelty 5.0 of 10

    A GRPO variant that scales advantages by group size and mean reward is claimed to reduce demographic F1 gaps in clinical VLLMs, but the abstract and body report different experiments and the body's own tables contradi...

Pith tools