Pith. sign in

Demoting Racial Bias in Hate Speech Detection

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In current hate speech datasets, there exists a high correlation between annotators' perceptions of toxicity and signals of African American English (AAE). This bias in annotated training data and the tendency of machine learning models to amplify it cause AAE text to often be mislabeled as abusive/offensive/hate speech with a high false positive rate by current hate speech classifiers. In this paper, we use adversarial training to mitigate this bias, introducing a hate speech classifier that learns to detect toxic sentences while demoting confounds corresponding to AAE texts. Experimental results on a hate speech dataset and an AAE dataset suggest that our method is able to substantially reduce the false positive rate for AAE text while only minimally affecting the performance of hate speech classification.

citation-role summary

background 1

citation-polarity summary

fields

cs.CY 1

years

2025 1

verdicts

UNVERDICTED 1

roles

background 1

polarities

background 1

representative citing papers

Multimodal Political Bias Identification and Neutralization

cs.CY · 2025-06-20 · unverdicted · novelty 4.0

A proposed multimodal pipeline to identify and reduce political bias in news text and images remains unvalidated: the report presents architecture and qualitative examples, without quantitative results for most components.

citing papers explorer

Showing 1 of 1 citing paper.

  • Multimodal Political Bias Identification and Neutralization cs.CY · 2025-06-20 · unverdicted · none · ref 30 · internal anchor

    A proposed multimodal pipeline to identify and reduce political bias in news text and images remains unvalidated: the report presents architecture and qualitative examples, without quantitative results for most components.