Pith. sign in

REVIEW 8 cited by

LLaMA Beyond English: An Empirical Study on Language Capability Transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01055 v2 pith:K44MCLWE submitted 2024-01-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagetransferllamallmsnon-englishacrossbenchmarksempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent times, substantial advancements have been witnessed in large language models (LLMs), exemplified by ChatGPT, showcasing remarkable proficiency across a range of complex tasks. However, many mainstream LLMs (e.g. LLaMA) are pretrained on English-dominant corpus, which limits their performance in other non-English languages. In this paper, we focus on how to effectively transfer the capabilities of language generation and following instructions to a non-English language. To answer this question, we conduct an extensive empirical investigation based on LLaMA, accumulating over 1440 GPU hours. We analyze the impact of key factors such as vocabulary extension, further pretraining, and instruction tuning on transfer. To accurately assess the model's level of knowledge, we employ four widely used standardized testing benchmarks: C-Eval, MMLU, AGI-Eval, and GAOKAO-Bench. Furthermore, a comprehensive evaluation of the model's response quality is conducted, considering aspects such as accuracy, fluency, informativeness, logical coherence, and harmlessness, based on LLM-Eval, a benchmarks consisting instruction tasks from 17 diverse categories. Our evaluation results demonstrate that comparable performance to state-of-the-art transfer models can be achieved with less than 1% of the pretraining data, both in terms of knowledge alignment and response quality. Furthermore, the experimental outcomes across the thirteen low-resource languages also exhibit similar trends. We anticipate that the conclusions revealed by the experiments will aid the community in developing non-English LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual Self-Taught Faithfulness Evaluators

    cs.CL 2025-07 conditional novelty 6.0 of 10

    STEMF trains multilingual faithfulness evaluators from synthetic data alone, and English-only training yields the best average results across languages.

  2. Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems

    cs.IR 2025-07 conditional novelty 6.0 of 10

    Basic and multi-model LLM prompts can evaluate recommendation serendipity as well as or better than standard proxy formulas, reaching 21.5% Pearson correlation with user-study ratings.

  3. How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting

    cs.CL 2025-07 conditional novelty 5.0 of 10

    For multilingual RAG intent classification, the best translation strategy depends on the model and language; translating instructions into the user's language helps some models, while making the model answer in low-re...

  4. Teaching a Language Model to Speak the Language of Tools

    cs.IR 2025-06 conditional novelty 5.0 of 10

    LoRA fine-tuning of BgGPT models on a bilingual Bulgarian function-calling dataset yields large gains on a self-built 120-case benchmark while keeping knowledge benchmarks stable.

  5. Text2Cypher Across Languages: Evaluating and Finetuning LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new multilingual Text2Cypher benchmark shows LLMs rank English highest, Spanish next, and Turkish lowest, and multilingual finetuning narrows the language gap more than English-only finetuning.

  6. Bias Attribution in Filipino Language Models: Extending a Bias Interpretability Metric for Application on Agglutinative Languages

    cs.CL 2025-06 reject novelty 5.0 of 10

    Filipino language models attribute gender and sexuality bias to concrete nouns such as people and objects, diverging from the action-heavy bias drivers observed in English models.

  7. Beyond Text Compression: Evaluating Tokenizers Across Scales

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Tokenizer choice matters mostly for multilingual tasks, and 350M-parameter models can predict 2.7B model ranking on translation but not on English benchmarks.

  8. From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A neuron-activation-based alignment score for LLMs correlates highly with downstream multilingual performance and transferability across nine open models.

Pith tools