Twitch's AutoMod flags only about 22% of hateful comments, misses most implicit hate, and blocks a large share of non-hateful uses of sensitive words.
Automated Content Moderation Increases Adherence to Community Guidelines
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Online social media platforms use automated moderation systems to remove or reduce the visibility of rule-breaking content. While previous work has documented the importance of manual content moderation, the effects of automated content moderation remain largely unknown. Here, in a large study of Facebook comments (n=412M), we used a fuzzy regression discontinuity design to measure the impact of automated content moderation on subsequent rule-breaking behavior (number of comments hidden/deleted) and engagement (number of additional comments posted). We found that comment deletion decreased subsequent rule-breaking behavior in shorter threads (20 or fewer comments), even among other participants, suggesting that the intervention prevented conversations from derailing. Further, the effect of deletion on the affected user's subsequent rule-breaking behavior was longer-lived than its effect on reducing commenting in general, suggesting that users were deterred from rule-breaking but not from commenting. In contrast, hiding (rather than deleting) content had small and statistically insignificant effects. Our results suggest that automated content moderation increases adherence to community guidelines.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
Twitch's AutoMod flags only about 22% of hateful comments, misses most implicit hate, and blocks a large share of non-hateful uses of sensitive words.