Pith. sign in

REVIEW 10 cited by

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.11080 v5 pith:B5H55RVH submitted 2020-03-24 cs.CL cs.LG

classification cs.CLcs.LG
keywords tasksbenchmarkmodelsacrosscross-linguallanguagesmultilingualbeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders XTREME benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KatotohananQA is a Filipino translation of TruthfulQA; seven LLMs scored 94.72% in English versus 83.87% in Filipino, with GPT-5 and GPT-5 mini showing the smallest gap.

  2. When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Projecting away the estimated Modern Standard Arabic subspace during fine-tuning improves generation across 25 Arabic dialects by up to +4.9 chrF++, evidence that subspace dominance by a high-resource variety restrict...

  3. Beyond Literal Token Overlap: Token Alignability for Multilinguality

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new metric based on subword token alignment predicts cross-lingual transfer in multilingual models better than literal token overlap, especially for different-script language pairs.

  4. Find Central Dogma Again: Leveraging Multilingual Transfer in Large Language Models

    q-bio.GN 2025-02 reject novelty 5.0 of 10

    A GPT-2 model fine-tuned on multilingual sentence similarity achieves at best 81% accuracy on classifying matching vs non-matching DNA-protein pairs, but the result is highly seed-dependent and only with an easy test set.

  5. Breaking Physical and Linguistic Borders: Multilingual Federated Prompt Tuning for Low-Resource Languages

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Federated averaging of prompt embeddings from a frozen multilingual model improves accuracy on some low-resource tasks (XNLI) but not consistently on others (MasakhaNEWS).

  6. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.

  7. Human Genome Book: Words, Sentences and Paragraphs

    q-bio.OT 2025-01 reject novelty 4.0 of 10

    A GPT-2 model trained on DNA, protein, and English is used to segment the human genome into book-like words, sentences, paragraphs, and chapters, but transfer to DNA is only hypothesized, not validated.

  8. LuxVeri at GenAI Detection Task 3: Cross-Domain Detection of AI-Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A RoBERTa ensemble with inverse perplexity weighting achieved 0.826 TPR (non-adversarial) and 0.801 TPR (adversarial) in a cross-domain AI-text detection shared task.

  9. Can linguists better understand DNA?

    cs.CL 2024-12 reject novelty 4.0 of 10

    Fine-tuning on English sentence-pair similarity appears to help GPT-2 and BERT classify DNA similarity, but the evidence is weakened by missing baselines and best-of-N seed selection.

  10. Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

    cs.CL 2025-06

Pith tools