Pith. sign in

MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline with precise character offsets of their citations in the article text. MegaWika 2 is a major upgrade from the original MegaWika, spanning six times as many articles and twice as many fully scraped citations. Both MegaWika and MegaWika 2 support report generation research ; whereas MegaWika also focused on supporting question answering and retrieval applications, MegaWika 2 is designed to support fact checking and analyses across time and language.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning

cs.CL · 2025-09-08 · conditional · novelty 6.0

mmBERT, a 3T-token encoder-only model pretrained on over 1,800 languages with inverse mask-rate and temperature schedules, substantially outperforms prior multilingual encoders like XLM-R and approaches ModernBERT on English.

citing papers explorer

Showing 1 of 1 citing paper.

  • mmBERT: A Modern Multilingual Encoder with Annealed Language Learning cs.CL · 2025-09-08 · conditional · none · ref 5 · internal anchor

    mmBERT, a 3T-token encoder-only model pretrained on over 1,800 languages with inverse mask-rate and temperature schedules, substantially outperforms prior multilingual encoders like XLM-R and approaches ModernBERT on English.