Safety mechanisms trained in English largely fail to activate for harmful prompts in four low-resource African languages, even when models understand the meaning.
U nity AI Guard: Pioneering Toxicity Detection Across Low-Resource I ndian Languages
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Illusion of Cross-Lingual Safety in Low-Resource Languages
Safety mechanisms trained in English largely fail to activate for harmful prompts in four low-resource African languages, even when models understand the meaning.