A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet most detectors still treat toxicity as an intrinsic property of isolated text. This position paper argues that toxicity detection should be evaluated as the contextual measurement of situated communicative harm, rather than as single-label text classification. Toxicity is not contained in words alone; it emerges when a communicative act is interpreted by an audience within a normative and social context. We introduce the Contextual Stress Framework (CSF), which defines toxicity as a relation between perceived norm violation and induced stress or disruption. CSF explains why text-intrinsic detectors overflag dialectal or reclaimed language, miss coded or pragmatic abuse, and remain brittle under meaning-preserving transformations. We propose CSF-Eval, an evaluation agenda that separates text risk, norm violation, disruption, uncertainty, and policy action.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ModelCitizens: Representing Community Voices in Online Safety
A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.