Pith. sign in

REVIEW 1 cited by

Your fairness may vary: Pretrained language model fairness in toxic text classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.01250 v3 pith:KJTE3IQA submitted 2021-08-03 cs.CL cs.LG

classification cs.CLcs.LG
keywords fairnesslanguagemodelspretrainedaccuracymeasuresmodeltext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The popularity of pretrained language models in natural language processing systems calls for a careful evaluation of such models in down-stream tasks, which have a higher potential for societal impact. The evaluation of such systems usually focuses on accuracy measures. Our findings in this paper call for attention to be paid to fairness measures as well. Through the analysis of more than a dozen pretrained language models of varying sizes on two toxic text classification tasks (English), we demonstrate that focusing on accuracy measures alone can lead to models with wide variation in fairness characteristics. Specifically, we observe that fairness can vary even more than accuracy with increasing training data size and different random initializations. At the same time, we find that little of the fairness variation is explained by model size, despite claims in the literature. To improve model fairness without retraining, we show that two post-processing methods developed for structured, tabular data can be successfully applied to a range of pretrained language models. Warning: This paper contains samples of offensive text.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mitigating Confounding in Speech-Based Dementia Detection through Weight Masking

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Masking weights that react to gender in a fine-tuned BERT reduces gender gaps in dementia predictions while keeping most of the detection accuracy.

Pith tools