REVIEW 1 cited by
Towards generalisable hate speech detection: a review on obstacles and solutions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Hate speech is one type of harmful online content which directly attacks or promotes hate towards a group or an individual member based on their actual or perceived aspects of identity, such as ethnicity, religion, and sexual orientation. With online hate speech on the rise, its automatic detection as a natural language processing task is gaining increasing interest. However, it is only recently that it has been shown that existing models generalise poorly to unseen data. This survey paper attempts to summarise how generalisable existing hate speech detection models are, reason why hate speech models struggle to generalise, sums up existing attempts at addressing the main obstacles, and then proposes directions of future research to improve generalisation in hate speech detection.
Forward citations
Cited by 1 Pith paper
-
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
A persona-based generation pipeline creates culturally varied content moderation test sets, but its central 'greater challenge' claim depends on unvalidated synthetic labels and unreleased data.
Discussion (0). Continue with ORCID to comment.