On a bilingual NSPS benchmark, Korean language consistently lowers harmful compliance (~10pp TRS), Korean grounding often mitigates that drop, and open- vs closed-source models reverse under direct requests.
Tongue-Tied: Breaking LLM s Safety Through New Language Learning
2 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
fields
cs.CL 2years
2026 2representative citing papers
A survey that catalogs threat models, detection approaches, and mitigation strategies for toxicity in multilingual LLMs while identifying challenges such as uneven language coverage and culturally variable harm definitions.
citing papers explorer
-
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
On a bilingual NSPS benchmark, Korean language consistently lowers harmful compliance (~10pp TRS), Korean grounding often mitigates that drop, and open- vs closed-source models reverse under direct requests.
-
A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models
A survey that catalogs threat models, detection approaches, and mitigation strategies for toxicity in multilingual LLMs while identifying challenges such as uneven language coverage and culturally variable harm definitions.