Pith. sign in

REVIEW 2 cited by

Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.07997 v2 pith:DKRGWLM2 submitted 2021-11-15 cs.CL cs.HC

classification cs.CLcs.HC
keywords beliefslanguagetoxictoxicityannotatordetectionannotatorsanti-black
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The perceived toxicity of language can vary based on someone's identity and beliefs, but this variation is often ignored when collecting toxic language datasets, resulting in dataset and model biases. We seek to understand the who, why, and what behind biases in toxicity annotations. In two online studies with demographically and politically diverse participants, we investigate the effect of annotator identities (who) and beliefs (why), drawing from social psychology research about hate speech, free speech, racist beliefs, political leaning, and more. We disentangle what is annotated as toxic by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity. Our results show strong associations between annotator identity and beliefs and their ratings of toxicity. Notably, more conservative annotators and those who scored highly on our scale for racist beliefs were less likely to rate anti-Black language as toxic, but more likely to rate AAE as toxic. We additionally present a case study illustrating how a popular toxicity detection system's ratings inherently reflect only specific beliefs and perspectives. Our findings call for contextualizing toxicity labels in social variables, which raises immense implications for toxic language annotation and detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Keywords: Evaluating Large Language Model Classification of Nuanced Ableism

    cs.CL 2025-05 conditional novelty 7.0 of 10

    LLMs identify autism-related words but frequently misclassify nuanced ableism, over-flagging intra-community language and under-flagging harmful stereotypes.

  2. Data-Driven and Participatory Approaches toward Neuro-Inclusive AI

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Across five studies, the dissertation documents exclusion of autistic perspectives in human-robot interaction research and releases AUTALIC, a Reddit-derived benchmark for anti-autistic hate speech detection.

Pith tools