Pith. sign in

REVIEW 1 cited by

RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.14397 v2 pith:F3XHF2XB submitted 2024-04-22 cs.CL cs.CYcs.LG

classification cs.CLcs.CYcs.LG
keywords llmsmodelsmultilinguallanguagertp-lxtheytoxicalthough
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) and small language models (SLMs) are being adopted at remarkable speed, although their safety still remains a serious concern. With the advent of multilingual S/LLMs, the question now becomes a matter of scale: can we expand multilingual safety evaluations of these models with the same velocity at which they are deployed? To this end, we introduce RTP-LX, a human-transcreated and human-annotated corpus of toxic prompts and outputs in 28 languages. RTP-LX follows participatory design practices, and a portion of the corpus is especially designed to detect culturally-specific toxic language. We evaluate 10 S/LLMs on their ability to detect toxic content in a culturally-sensitive, multilingual scenario. We find that, although they typically score acceptably in terms of accuracy, they have low agreement with human judges when scoring holistically the toxicity of a prompt; and have difficulty discerning harm in context-dependent scenarios, particularly with subtle-yet-harmful content (e.g. microaggressions, bias). We release this dataset to contribute to further reduce harmful uses of these models and improve their safe deployment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies

    cs.CL 2024-12 conditional novelty 6.0 of 10

    M-ALERT, a 75k-prompt multilingual safety benchmark, shows that LLM safety varies substantially across five languages and across risk categories, with no model reaching the 99% safe threshold in every language.

Pith tools