Simple character-level perturbations, including Unicode homoglyphs and whitespace manipulations, evade two black-box hate speech classifiers for most toxic tweets, according to this study.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
All You Need is "Leet": Evading Hate-speech Detection AI
Simple character-level perturbations, including Unicode homoglyphs and whitespace manipulations, evade two black-box hate speech classifiers for most toxic tweets, according to this study.