REVIEW 8 cited by
All Languages Matter: On the Multilingual Safety of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Safety lies at the core of developing and deploying large language models (LLMs). However, previous safety benchmarks only concern the safety in one language, e.g. the majority language in the pretraining data such as English. In this work, we build the first multilingual safety benchmark for LLMs, XSafety, in response to the global deployment of LLMs in practice. XSafety covers 14 kinds of commonly used safety issues across 10 languages that span several language families. We utilize XSafety to empirically study the multilingual safety for 4 widely-used LLMs, including both close-API and open-source models. Experimental results show that all LLMs produce significantly more unsafe responses for non-English queries than English ones, indicating the necessity of developing safety alignment for non-English languages. In addition, we propose several simple and effective prompting methods to improve the multilingual safety of ChatGPT by evoking safety knowledge and improving cross-lingual generalization of safety alignment. Our prompting method can significantly reduce the ratio of unsafe responses from 19.1% to 9.7% for non-English queries. We release our data at https://github.com/Jarviswang94/Multilingual_safety_benchmark.
Forward citations
Cited by 8 Pith papers
-
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
ROK-FORTRESS shows Korean-language prompts increase LLM safety suppression compared with English, while Korean geopolitical grounding often reduces that suppression, indicating translation-only evaluations miss langua...
-
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
IndoSafety, a culturally grounded safety benchmark for five Indonesian language varieties, shows unsafe response rates up to 40% in regional models and demonstrates that safety tuning on formal Indonesian transfers to...
-
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
A unified 16-category refusal taxonomy with human and synthetic datasets and a low-cost classifier for auditing refusal behavior in LLMs.
-
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.
-
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
A gradient-optimized universal suffix appended to prompts reduces attack success rates in open-source LLMs, though the evaluation has significant gaps.
-
On the effective transfer of knowledge from English to Hindi Wikipedia
A retrieval, neutralization, and machine-translation pipeline can add relevant factual text to Hindi Wikipedia biography sections, but the claimed 65% and 62% gains are not fully supported by the reported evaluation.
-
Towards Robust Fact-Checking: A Multi-Agent System with Advanced Evidence Retrieval
A multi-agent LLM pipeline with credibility-filtered full-text web retrieval reports better fact-checking F1 than four baselines on small benchmark subsamples.
-
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
A new cross-lingual benchmark shows large language models comply with explicit requests to use swear words far more often in Indic languages than in English, revealing a safety alignment gap.
Discussion (0). Continue with ORCID to comment.