Pith. sign in

REVIEW 4 cited by

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10378 v5 pith:JWB5XPAP submitted 2023-10-16 cs.CL cs.AIcs.HCcs.LG

classification cs.CLcs.AIcs.HCcs.LG
keywords consistencyfactualknowledgelanguagesmodelcross-linguallanguagemultilingual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the same model, we study the cross-lingual consistency (CLC) of factual knowledge in various multilingual PLMs. To this end, we propose a Ranking-based Consistency (RankC) metric to evaluate knowledge consistency across languages independently from accuracy. Using this metric, we conduct an in-depth analysis of the determining factors for CLC, both at model level and at language-pair level. Among other results, we find that increasing model size leads to higher factual probing accuracy in most languages, but does not improve cross-lingual consistency. Finally, we conduct a case study on CLC when new factual associations are inserted in the PLMs via model editing. Results on a small sample of facts inserted in English reveal a clear pattern whereby the new piece of knowledge transfers only to languages with which English has a high RankC score.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking Cross-lingual Gaps from a Statistical Viewpoint

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Cross-lingual accuracy gaps in LLMs are dominated by higher response variance in target languages, not missing knowledge; ensembling and variance-reduction prompts shrink the gap.

  2. A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A multilingual two-phase evaluation shows LLMs lean on query language for factual questions and on training-country perspective for territorial and historical disputes.

  3. Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Across six country-pair comparisons, GPT-4o-mini and GigaChat-Max side with US positions 64-81% of the time, Qwen2.5 and Llama-4 lean neutral more often, and a debias prompt shifts these numbers by only a few points.

  4. QUST_NLP at SemEval-2025 Task 7: A Three-Stage Retrieval Framework for Monolingual and Crosslingual Fact-Checked Claim Retrieval

    cs.IR 2025-06 conditional novelty 4.0 of 10

    A three-stage ensemble of retrieval models, rerankers, and weighted voting achieves strong multilingual fact-checked claim retrieval results at SemEval-2025 Task 7.

Pith tools