Pith. sign in

MegaWika: Millions of reports and their sources across 50 diverse languages

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced source materials. We process this dataset for a myriad of applications, going beyond the initial Wikipedia citation extraction and web scraping of content, including translating non-English articles for cross-lingual applications and providing FrameNet parses for automated semantic analysis. MegaWika is the largest resource for sentence-level report generation and the only report generation dataset that is multilingual. We manually analyze the quality of this resource through a semantically stratified sample. Finally, we provide baseline results and trained models for crucial steps in automated report generation: cross-lingual question answering and citation retrieval.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning

cs.CL · 2025-09-08 · conditional · novelty 6.0

mmBERT, a 3T-token encoder-only model pretrained on over 1,800 languages with inverse mask-rate and temperature schedules, substantially outperforms prior multilingual encoders like XLM-R and approaches ModernBERT on English.

citing papers explorer

Showing 1 of 1 citing paper.

  • mmBERT: A Modern Multilingual Encoder with Annealed Language Learning cs.CL · 2025-09-08 · conditional · none · ref 4 · internal anchor

    mmBERT, a 3T-token encoder-only model pretrained on over 1,800 languages with inverse mask-rate and temperature schedules, substantially outperforms prior multilingual encoders like XLM-R and approaches ModernBERT on English.