REVIEW 6 cited by
Detecting Hate Speech with GPT-3
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sophisticated language models such as OpenAI's GPT-3 can generate hateful text that targets marginalized groups. Given this capacity, we are interested in whether large language models can be used to identify hate speech and classify text as sexist or racist. We use GPT-3 to identify sexist and racist text passages with zero-, one-, and few-shot learning. We find that with zero- and one-shot learning, GPT-3 can identify sexist or racist text with an average accuracy between 55 per cent and 67 per cent, depending on the category of text and type of learning. With few-shot learning, the model's accuracy can be as high as 85 per cent. Large language models have a role to play in hate speech detection, and with further development they could eventually be used to counter hate speech.
Forward citations
Cited by 6 Pith papers
-
LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks
LCO combines proactive LLM-generated safety constraints with evolutionary sampling to cut in-context reward hacking while preserving task performance on tweet and tool-use benchmarks.
-
Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification
LLMs show systematic target-dependent sentiment inconsistency that is politically biased: left and center politicians rated more positively, far-right politicians more negatively, with stronger effects in larger model...
-
Web-Browsing LLMs Can Access Social Media Profiles and Infer User Demographics
Web-browsing LLMs can retrieve X profile content and infer demographics with above-chance accuracy in some cases, but the study's evidence is partly confounded by training-data memorization and a heavily reduced synth...
-
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech
Zero-shot LLMs conditioned on German legal knowledge at several abstraction levels still lag far behind human legal experts at detecting punishable hate speech under Section 130 StGB.
-
Mario at EXIST 2025: A Simple Gateway to Effective Multilingual Sexism Detection
A LoRA fine-tuned Llama 3.1 8B with hierarchical adapter routing won all three text subtasks of EXIST 2025 sexism detection in English and Spanish.
-
Context-Aware Content Moderation for German Newspaper Comments
LSTM and CNN models for German newspaper comment moderation improve when given article title and user history, while ChatGPT-3.5 zero-shot classification does not.
Discussion (0). Sign in to comment.