Pith. sign in

REVIEW 1 cited by

Multilingual large language models leak human stereotypes across language boundaries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07141 v3 pith:743TMOOG submitted 2023-12-12 cs.CL

classification cs.CL
keywords leakagemodelsacrosslanguagemultilingualstereotypestereotypesgpt-3
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multilingual large language models have gained prominence for their proficiency in processing and generating text across languages. Like their monolingual counterparts, multilingual models are likely to pick up on stereotypes and other social biases present in their training data. In this paper, we study a phenomenon we term stereotype leakage, which refers to how training a model multilingually may lead to stereotypes expressed in one language showing up in the models' behaviour in another. We propose a measurement framework for stereotype leakage and investigate its effect across English, Russian, Chinese, and Hindi and with GPT-3.5, mT5, and mBERT. Our findings show a noticeable leakage of positive, negative, and non-polar associations across all languages. We find that of these models, GPT-3.5 exhibits the most stereotype leakage, and Hindi is the most susceptible to leakage effects. WARNING: This paper contains model outputs which could be offensive in nature.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new India-focused benchmark shows that popular LLMs exhibit measurable negative bias against marginalized Indian identities and frequently reinforce caste, religion, region, and tribe stereotypes.

Pith tools