Pith. sign in

REVIEW 6 cited by

MLQA: Evaluating Cross-lingual Extractive Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.07475 v3 pith:2XTMC33H submitted 2019-10-16 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords mlqacross-lingualdatasetslanguagesenglishextractivelanguageother
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets. Such annotated datasets are difficult and costly to collect, and rarely exist in languages other than English, making training QA systems in other languages challenging. An alternative to building large monolingual training datasets is to develop cross-lingual systems which can transfer to a target language without requiring training data in that language. In order to develop such systems, it is crucial to invest in high quality multilingual evaluation benchmarks to measure progress. We present MLQA, a multi-way aligned extractive QA evaluation benchmark intended to spur research in this area. MLQA contains QA instances in 7 languages, namely English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese. It consists of over 12K QA instances in English and 5K in each other language, with each QA instance being parallel between 4 languages on average. MLQA is built using a novel alignment context strategy on Wikipedia articles, and serves as a cross-lingual extension to existing extractive QA datasets. We evaluate current state-of-the-art cross-lingual representations on MLQA, and also provide machine-translation-based baselines. In all cases, transfer results are shown to be significantly behind training-language performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North Africa

    cs.CL 2025-07 conditional novelty 7.0 of 10

    TyDi QA-WANA is a new 28,000-example QA benchmark covering 10 under-represented languages with long-context, information-seeking questions and baseline evaluations.

  2. Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

    cs.CL 2026-03 conditional novelty 6.0 of 10

    Language-aware query selection with a gated query bank improves multilingual, ASR-only-distilled speech LLMs on instruction following and spoken QA.

  3. FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An adaptive, per-language data filtering and deduplication pipeline produces multilingual LLM pre-training corpora that beat prior public datasets on 11 of 14 evaluated languages, and a 20TB, 1,868 language-script dat...

  4. IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems

    cs.CL 2025-06 conditional novelty 6.0 of 10

    IndicRAGSuite offers a 13-language, human-verified retrieval benchmark and two large-scale training datasets for Indian language RAG.

  5. Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    JQL trains small multilingual quality scorers from LLM judgments and human annotations, and filtering pretraining data with them improves downstream multilingual model performance over heuristic baselines.

  6. SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs

    cs.LG 2026-07 reject novelty 5.0 of 10

    SelectInfer profiles LLM neurons offline to load and compute only a subset during inference, but its accuracy claims are undermined by its own comparisons and test-data overlap.

Pith tools