REVIEW 7 cited by
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages. This imbalance results in LLMs performing significantly better on high-resource languages like English, German, and French, while their capabilities in low-resource languages remain inadequate. Currently, there is a lack of quantitative methods to evaluate the performance of LLMs in these low-resource languages. To address this gap, we propose the Language Ranker, an intrinsic metric designed to benchmark and rank languages based on LLM performance using internal representations. By comparing the LLM's internal representation of various languages against a baseline derived from English, we can assess the model's multilingual capabilities in a robust and language-agnostic manner. Our analysis reveals that high-resource languages exhibit higher similarity scores with English, demonstrating superior performance, while low-resource languages show lower similarity scores, underscoring the effectiveness of our metric in assessing language-specific capabilities. Besides, the experiments show that there is a strong correlation between the LLM's performance in different languages and the proportion of those languages in its pre-training corpus. These insights underscore the efficacy of the Language Ranker as a tool for evaluating LLM performance across different languages, particularly those with limited resources.
Forward citations
Cited by 7 Pith papers
-
TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla
TigerCoder is a fine-tuned Bangla code-generation LLM family that posts 0.82 Pass@1 on the new MBPP-Bangla benchmark, but its gains partly come from model selection on the test set.
-
What Makes You CLIC: Detection of Croatian Clickbait Headlines
CLIC, a new 2,907-headline Croatian clickbait dataset, shows about 53% of sampled headlines are clickbait and fine-tuned BERTić (F1 0.78) beats zero/few-shot LLMs on detection.
-
Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language Models
Across 17 sensitive topics and six languages, large language models produce significantly different refusal rates and answer distributions depending on the language used, with the pattern varying by model.
-
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
A new Korean benchmark, KoSEnd, shows LLMs have limited grasp of Korean sentence endings, and warning them about potentially missing endings improves their choices.
-
Stylometry recognizes human and LLM-generated texts in short samples
Stylometric features and tree-based classifiers separate human-written Wikipedia summaries from LLM-generated texts with high cross-validated accuracy on a new seven-class benchmark, though performance drops on other ...
-
Vuyko Mistral: Adapting LLMs for Low-Resource Dialectal Translation
The authors release a Hutsul-Ukrainian corpus and show LoRA-fine-tuned 7B models beat GPT-4o on automated and LLM-based metrics, but the evaluation is contaminated by overlapping training and test sources.
-
Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs
A small evaluation of Llama 3.1 shows factuality in school-level question answering degrades with decreasing language speaker count, though the statistical support is weakened by methodological issues.
Discussion (0). Sign in to comment.