Pith. sign in

REVIEW 1 cited by

Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08918 v1 pith:CKY2ZSIE submitted 2024-10-11 cs.CY

Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing

classification cs.CY
keywords wikimediacommunitydatagreatercontentdatasetslanguagemodeling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Wikimedia content is used extensively by the AI community and within the language modeling community in particular. In this paper, we provide a review of the different ways in which Wikimedia data is curated to use in NLP tasks across pre-training, post-training, and model evaluations. We point to opportunities for greater use of Wikimedia content but also identify ways in which the language modeling community could better center the needs of Wikimedia editors. In particular, we call for incorporating additional sources of Wikimedia data, a greater focus on benchmarks for LLMs that encode Wikimedia principles, and greater multilingualism in Wikimedia-derived datasets.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

    cs.CL 2026-05 conditional novelty 7.0

    Introduces the MCN multilingual citation-needed detection corpus for 18 languages and demonstrates that fine-tuned small decoder models outperform prompted LLMs in both multilingual and cross-lingual transfer settings.