Counterfactual Fairness in Text Classification through Robustness

Alex Beutel; Ankur Taly; Ed H. Chi; Nicole Limtiaco; Sahaj Garg; Vincent Perot

arxiv: 1809.10610 · v2 · pith:OC3EW7FEnew · submitted 2018-09-27 · 💻 cs.LG · stat.ML

Counterfactual Fairness in Text Classification through Robustness

Sahaj Garg , Vincent Perot , Nicole Limtiaco , Ankur Taly , Ed H. Chi , Alex Beutel This is my paper

classification 💻 cs.LG stat.ML

keywords fairnesscounterfactualtextclassificationtokenapproachesblindnessclassifiers

0 comments

read the original abstract

In this paper, we study counterfactual fairness in text classification, which asks the question: How would the prediction change if the sensitive attribute referenced in the example were different? Toxicity classifiers demonstrate a counterfactual fairness issue by predicting that "Some people are gay" is toxic while "Some people are straight" is nontoxic. We offer a metric, counterfactual token fairness (CTF), for measuring this particular form of fairness in text classifiers, and describe its relationship with group fairness. Further, we offer three approaches, blindness, counterfactual augmentation, and counterfactual logit pairing (CLP), for optimizing counterfactual token fairness during training, bridging the robustness and fairness literature. Empirically, we find that blindness and CLP address counterfactual token fairness. The methods do not harm classifier performance, and have varying tradeoffs with group fairness. These approaches, both for measurement and optimization, provide a new path forward for addressing fairness concerns in text classification.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Ethical and social risks of harm from Language Models
cs.CL 2021-12 accept novelty 6.0

The authors provide a detailed taxonomy of 21 risks associated with language models, covering discrimination, information leaks, misinformation, malicious applications, interaction harms, and societal impacts like job...