Pith. sign in

REVIEW 5 cited by

Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03814 v2 pith:46SHWVRQ submitted 2024-03-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsmodelslanguagesintendedlanguagelargemultiqrespond
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) need to serve everyone, including a global majority of non-English speakers. However, most LLMs today, and open LLMs in particular, are often intended for use in just English (e.g. Llama2, Mistral) or a small handful of high-resource languages (e.g. Mixtral, Qwen). Recent research shows that, despite limits in their intended use, people prompt LLMs in many different languages. Therefore, in this paper, we investigate the basic multilingual capabilities of state-of-the-art open LLMs beyond their intended use. For this purpose, we introduce MultiQ, a new silver standard benchmark for basic open-ended question answering with 27.4k test questions across a typologically diverse set of 137 languages. With MultiQ, we evaluate language fidelity, i.e. whether models respond in the prompted language, and question answering accuracy. All LLMs we test respond faithfully and/or accurately for at least some languages beyond their intended use. Most models are more accurate when they respond faithfully. However, differences across models are large, and there is a long tail of languages where models are neither accurate nor faithful. We explore differences in tokenization as a potential explanation for our findings, identifying possible correlations that warrant further investigation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual Performance Biases of Large Language Models in Education

    cs.CL 2025-04 conditional novelty 6.0 of 10

    LLMs are less reliable at tutoring, feedback, and misconception detection in lower-resource languages, and English prompts usually work as well as or better than prompts in the target language.

  2. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

    cs.CL 2024-11 conditional novelty 6.0 of 10

    INCLUDE is a multilingual benchmark of 197,243 exam questions from local sources that evaluates how well LLMs handle regional and cultural knowledge.

  3. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.

  4. Analysis of Indic Language Capabilities in LLMs

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A desk-research review finds that LLM performance is strongest for Hindi, Bengali, Marathi, Telugu, and Tamil, and recommends prioritizing these five languages for safety benchmarks.

  5. Multilingual Large Language Models: A Systematic Survey

    cs.CL 2024-11 conditional novelty 3.0 of 10

    This is a systematic review that categorizes research on multilingual LLMs into architecture, corpora, tuning, evaluation, interpretability, and applications, with a public curated paper list.

Pith tools