Pith. sign in

REVIEW 1 cited by

Gender Biases and Where to Find Them: Exploring Gender Bias in Pre-Trained Transformer-based Language Models Using Movement Pruning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.02463 v1 pith:M4HSGWU4 submitted 2022-07-06 cs.CL

classification cs.CL
keywords modelbiaspruningdebiasingframeworkgenderlanguagedemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language model debiasing has emerged as an important field of study in the NLP community. Numerous debiasing techniques were proposed, but bias ablation remains an unaddressed issue. We demonstrate a novel framework for inspecting bias in pre-trained transformer-based language models via movement pruning. Given a model and a debiasing objective, our framework finds a subset of the model containing less bias than the original model. We implement our framework by pruning the model while fine-tuning it on the debiasing objective. Optimized are only the pruning scores - parameters coupled with the model's weights that act as gates. We experiment with pruning attention heads, an important building block of transformers: we prune square blocks, as well as establish a new way of pruning the entire heads. Lastly, we demonstrate the usage of our framework using gender bias, and based on our findings, we propose an improvement to an existing debiasing method. Additionally, we re-discover a bias-performance trade-off: the better the model performs, the more bias it contains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Implicit Priors Editing in Stable Diffusion via Targeted Token Adjustment

    cs.CV 2024-12 conditional novelty 5.0 of 10

    EMBEDIT edits a single word token embedding in Stable Diffusion to steer implicit visual priors (e.g., making 'bear' generate 'polar bear'), reporting better accuracy than cross-attention editing while using far fewer...

Pith tools