REVIEW 4 cited by
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Hate speech has grown significantly on social media, causing serious consequences for victims of all demographics. Despite much attention being paid to characterize and detect discriminatory speech, most work has focused on explicit or overt hate speech, failing to address a more pervasive form based on coded or indirect language. To fill this gap, this work introduces a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication. We present systematic analyses of our dataset using contemporary baselines to detect and explain implicit hate speech, and we discuss key features that challenge existing models. This dataset will continue to serve as a useful benchmark for understanding this multifaceted issue.
Forward citations
Cited by 4 Pith papers
-
A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance
A taxonomy of 14 hate speech definition components plus evidence that zero-shot LLM hate speech classification is sensitive to which definition is inserted into the prompt, with model-dependent effects.
-
Dark Personality Traits and Online Toxicity: Linking Self-Reports to Reddit Activity
Dark personality traits show weak, mostly non-significant links to toxicity in Reddit comments, while self-reported engagement with incivility strongly aligns with measurable toxic language.
-
Social Hatred: Efficient Multimodal Detection of Hatemongers
A user-level hate-monger detection model that aggregates post-level hate probabilities with ego-network features outperforms text-only and graph-only baselines across three social platforms.
-
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
On three hate speech datasets, including a new code-mixed IndoHateMix benchmark, fine-tuned LLMs such as LLaMA-3.1 beat multilingual BERT models, with the largest gains on code-mixed Indian text.
Discussion (0). Sign in to comment.