REVIEW 4 cited by
Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Word embeddings are widely used in NLP for a vast range of tasks. It was shown that word embeddings derived from text corpora reflect gender biases in society. This phenomenon is pervasive and consistent across different word embedding models, causing serious concern. Several recent works tackle this problem, and propose methods for significantly reducing this gender bias in word embeddings, demonstrating convincing results. However, we argue that this removal is superficial. While the bias is indeed substantially reduced according to the provided bias definition, the actual effect is mostly hiding the bias, not removing it. The gender bias information is still reflected in the distances between "gender-neutralized" words in the debiased embeddings, and can be recovered from them. We present a series of experiments to support this claim, for two debiasing methods. We conclude that existing bias removal techniques are insufficient, and should not be trusted for providing gender-neutral modeling.
Forward citations
Cited by 4 Pith papers
-
Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Unsigned differential activations locate a few GLU-MLP neurons whose zeroing surgically destabilizes demographic bias while retaining ~99.5% of measured capabilities.
-
Mitigating Gender Bias in Contextual Word Embeddings
Regularized masked-language modeling and name-masking reduce gender bias in embeddings, but the contextual results rely heavily on evaluation metrics aligned with the training objective.
-
Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space
Word relationships in embedding space can be represented as orthogonal or linear transformations, not just translation vectors, with comparable or better analogy-solving accuracy.
-
Unlearn Dataset Bias in Natural Language Inference by Fitting the Residual
DRiFt, a residual-fitting debiasing algorithm, improves NLI model accuracy on challenge sets like HANS by training on examples a biased model cannot solve.
Discussion (0). Continue with ORCID to comment.