Pith. sign in

REVIEW 4 cited by

Latent Hatred: A Benchmark for Understanding Implicit Hate Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.05322 v1 pith:Z3ZTIGIR submitted 2021-09-11 cs.CL cs.SI

classification cs.CLcs.SI
keywords speechhatebenchmarkimplicitdatasetdetectunderstandingwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hate speech has grown significantly on social media, causing serious consequences for victims of all demographics. Despite much attention being paid to characterize and detect discriminatory speech, most work has focused on explicit or overt hate speech, failing to address a more pervasive form based on coded or indirect language. To fill this gap, this work introduces a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication. We present systematic analyses of our dataset using contemporary baselines to detect and explain implicit hate speech, and we discuss key features that challenge existing models. This dataset will continue to serve as a useful benchmark for understanding this multifaceted issue.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A taxonomy of 14 hate speech definition components plus evidence that zero-shot LLM hate speech classification is sensitive to which definition is inserted into the prompt, with model-dependent effects.

  2. Dark Personality Traits and Online Toxicity: Linking Self-Reports to Reddit Activity

    cs.CY 2025-12 unverdicted novelty 5.0 of 10

    Dark personality traits show weak, mostly non-significant links to toxicity in Reddit comments, while self-reported engagement with incivility strongly aligns with measurable toxic language.

  3. Social Hatred: Efficient Multimodal Detection of Hatemongers

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A user-level hate-monger detection model that aggregates post-level hate probabilities with ego-network features outperforms text-only and graph-only baselines across three social platforms.

  4. Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On three hate speech datasets, including a new code-mixed IndoHateMix benchmark, fine-tuned LLMs such as LLaMA-3.1 beat multilingual BERT models, with the largest gains on code-mixed Indian text.

Pith tools