Pith. sign in

Debiasing isn't enough! -- On the Effectiveness of Debiasing MLMs and their Social Biases in Downstream Tasks

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We study the relationship between task-agnostic intrinsic and task-specific extrinsic social bias evaluation measures for Masked Language Models (MLMs), and find that there exists only a weak correlation between these two types of evaluation measures. Moreover, we find that MLMs debiased using different methods still re-learn social biases during fine-tuning on downstream tasks. We identify the social biases in both training instances as well as their assigned labels as reasons for the discrepancy between intrinsic and extrinsic bias evaluation measurements. Overall, our findings highlight the limitations of existing MLM bias evaluation measures and raise concerns on the deployment of MLMs in downstream applications using those measures.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Fairness Dynamics During Training

cs.CL · 2025-06-02 · conditional · novelty 5.0

Gender bias in Pythia-6.9b grows sharply after about 80k training steps even as general performance improves, and stopping earlier could trade 1.7% LAMBADA accuracy for a large fairness gain.

citing papers explorer

Showing 1 of 1 citing paper.

  • Fairness Dynamics During Training cs.CL · 2025-06-02 · conditional · none · ref 8 · internal anchor

    Gender bias in Pythia-6.9b grows sharply after about 80k training steps even as general performance improves, and stopping earlier could trade 1.7% LAMBADA accuracy for a large fairness gain.