Pith. sign in

REVIEW 6 cited by

Detecting Hate Speech with GPT-3

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.12407 v4 pith:4WDAVJQF submitted 2021-03-23 cs.CL

classification cs.CL
keywords textgpt-3hatelearningspeechcentidentifylanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sophisticated language models such as OpenAI's GPT-3 can generate hateful text that targets marginalized groups. Given this capacity, we are interested in whether large language models can be used to identify hate speech and classify text as sexist or racist. We use GPT-3 to identify sexist and racist text passages with zero-, one-, and few-shot learning. We find that with zero- and one-shot learning, GPT-3 can identify sexist or racist text with an average accuracy between 55 per cent and 67 per cent, depending on the category of text and type of learning. With few-shot learning, the model's accuracy can be as high as 85 per cent. Large language models have a role to play in hate speech detection, and with further development they could eventually be used to counter hate speech.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

    cs.CL 2026-04 conditional novelty 6.0 of 10

    LCO combines proactive LLM-generated safety constraints with evolutionary sampling to cut in-context reward hacking while preserving task performance on tweet and tool-use benchmarks.

  2. Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLMs show systematic target-dependent sentiment inconsistency that is politically biased: left and center politicians rated more positively, far-right politicians more negatively, with stronger effects in larger model...

  3. Web-Browsing LLMs Can Access Social Media Profiles and Infer User Demographics

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Web-browsing LLMs can retrieve X profile content and infer demographics with above-chance accuracy in some cases, but the study's evidence is partly confounded by training-data memorization and a heavily reduced synth...

  4. Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Zero-shot LLMs conditioned on German legal knowledge at several abstraction levels still lag far behind human legal experts at detecting punishable hate speech under Section 130 StGB.

  5. Mario at EXIST 2025: A Simple Gateway to Effective Multilingual Sexism Detection

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A LoRA fine-tuned Llama 3.1 8B with hierarchical adapter routing won all three text subtasks of EXIST 2025 sexism detection in English and Spanish.

  6. Context-Aware Content Moderation for German Newspaper Comments

    cs.CL 2025-05 conditional novelty 4.0 of 10

    LSTM and CNN models for German newspaper comment moderation improve when given article title and user history, while ChatGPT-3.5 zero-shot classification does not.

Pith tools