Pith. sign in

REVIEW 2 cited by

XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.07412 v2 pith:EXSHIAUZ submitted 2021-04-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords multilingualxtremetasksxtreme-rcapabilitieschallengingevaluationhttps
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning has brought striking advances in multilingual natural language processing capabilities over the past year. For example, the latest techniques have improved the state-of-the-art performance on the XTREME multilingual benchmark by more than 13 points. While a sizeable gap to human-level performance remains, improvements have been easier to achieve in some tasks than in others. This paper analyzes the current state of cross-lingual transfer learning and summarizes some lessons learned. In order to catalyze meaningful progress, we extend XTREME to XTREME-R, which consists of an improved set of ten natural language understanding tasks, including challenging language-agnostic retrieval tasks, and covers 50 typologically diverse languages. In addition, we provide a massively multilingual diagnostic suite (MultiCheckList) and fine-grained multi-dataset evaluation capabilities through an interactive public leaderboard to gain a better understanding of such models. The leaderboard and code for XTREME-R will be made available at https://sites.research.google/xtreme and https://github.com/google-research/xtreme respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. skLEP: A Slovak General Language Understanding Benchmark

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A nine-task Slovak-language understanding benchmark with translated and newly curated datasets, plus the first broad fine-tuned model comparison for Slovak.

  2. LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multilingual visual question-answering benchmark across 11 languages and 5 social attributes, evaluated on 7 large multimodal models.

Pith tools