Harmful humorization patterns can be injected into LLMs via a few examples, producing toxic jokes that look like safe refusals and evade top detectors.
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization
Harmful humorization patterns can be injected into LLMs via a few examples, producing toxic jokes that look like safe refusals and evade top detectors.