Pith. sign in

REVIEW 7 cited by

Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.11553 v3 pith:C7ZI7TRD submitted 2024-04-17 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagesperformancelanguagelow-resourceacrosscapabilitiesenglishllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages. This imbalance results in LLMs performing significantly better on high-resource languages like English, German, and French, while their capabilities in low-resource languages remain inadequate. Currently, there is a lack of quantitative methods to evaluate the performance of LLMs in these low-resource languages. To address this gap, we propose the Language Ranker, an intrinsic metric designed to benchmark and rank languages based on LLM performance using internal representations. By comparing the LLM's internal representation of various languages against a baseline derived from English, we can assess the model's multilingual capabilities in a robust and language-agnostic manner. Our analysis reveals that high-resource languages exhibit higher similarity scores with English, demonstrating superior performance, while low-resource languages show lower similarity scores, underscoring the effectiveness of our metric in assessing language-specific capabilities. Besides, the experiments show that there is a strong correlation between the LLM's performance in different languages and the proportion of those languages in its pre-training corpus. These insights underscore the efficacy of the Language Ranker as a tool for evaluating LLM performance across different languages, particularly those with limited resources.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla

    cs.CL 2025-09 conditional novelty 6.0 of 10

    TigerCoder is a fine-tuned Bangla code-generation LLM family that posts 0.82 Pass@1 on the new MBPP-Bangla benchmark, but its gains partly come from model selection on the test set.

  2. What Makes You CLIC: Detection of Croatian Clickbait Headlines

    cs.CL 2025-07 conditional novelty 6.0 of 10

    CLIC, a new 2,907-headline Croatian clickbait dataset, shows about 53% of sampled headlines are clickbait and fine-tuned BERTić (F1 0.78) beats zero/few-shot LLMs on detection.

  3. Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across 17 sensitive topics and six languages, large language models produce significantly different refusal rates and answer distributions depending on the language used, with the pattern varying by model.

  4. Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A new Korean benchmark, KoSEnd, shows LLMs have limited grasp of Korean sentence endings, and warning them about potentially missing endings improves their choices.

  5. Stylometry recognizes human and LLM-generated texts in short samples

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...

  6. Vuyko Mistral: Adapting LLMs for Low-Resource Dialectal Translation

    cs.CL 2025-06 reject novelty 4.0 of 10

    The authors release a Hutsul-Ukrainian corpus and show LoRA-fine-tuned 7B models beat GPT-4o on automated and LLM-based metrics, but the evaluation is contaminated by overlapping training and test sources.

  7. Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A small evaluation of Llama 3.1 shows factuality in school-level question answering degrades with decreasing language speaker count, though the statistical support is weakened by methodological issues.

Pith tools