Pith. sign in

REVIEW 3 cited by

Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.00501 v1 pith:HFMRZ7QN submitted 2022-05-01 cs.HC cs.AIcs.CLcs.LG

classification cs.HCcs.AIcs.CLcs.LG
keywords raterpoolsraterstoxicitycommentsmodelsannotationsidentity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning models are commonly used to detect toxicity in online conversations. These models are trained on datasets annotated by human raters. We explore how raters' self-described identities impact how they annotate toxicity in online comments. We first define the concept of specialized rater pools: rater pools formed based on raters' self-described identities, rather than at random. We formed three such rater pools for this study--specialized rater pools of raters from the U.S. who identify as African American, LGBTQ, and those who identify as neither. Each of these rater pools annotated the same set of comments, which contains many references to these identity groups. We found that rater identity is a statistically significant factor in how raters will annotate toxicity for identity-related annotations. Using preliminary content analysis, we examined the comments with the most disagreement between rater pools and found nuanced differences in the toxicity annotations. Next, we trained models on the annotations from each of the different rater pools, and compared the scores of these models on comments from several test sets. Finally, we discuss how using raters that self-identify with the subjects of comments can create more inclusive machine learning models, and provide more nuanced ratings than those by random raters.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Machine Understanding of Scientific Language

    cs.CL 2025-06 conditional novelty 7.0 of 10

    The thesis defines and evaluates tasks and datasets for automatic fact checking, cite-worthiness, exaggeration detection, and information change measurement in science communication, culminating in SPICED, a cross-med...

  2. Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Directly predicting whether annotators will disagree on a value label outperforms inferring disagreement from per-annotator value predictions on the Touché23-ValueEval dataset.

  3. ModelCitizens: Representing Community Voices in Online Safety

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.

Pith tools